# Emergency Recovery — Restore Remote Work First

audience: AI coding agents first.

Canonical lane registry: `docs/plans/2026-08-10-master-delivery-plan.md`.

NEVER halt globally. Reorganization means preserve WIP, update durable state, select next executable lane, continue. A failed/stopped agent is not a stopped mission.

## Mission

Restore normal work through proven three-host SSH offload FIRST. NEVER make unfinished k3s, runtime hardening, Git integration, broad verification, or review ceremony block owner use.

## Live checkpoint — 2026-08-10

- Cost emergency: owner reports $600+ model budget consumed without usable delivery. Stop transcript recovery, exploration, advisory panels, broad reviews, and duplicate agents/deploys.
- #231 full-session recovery is owner-paused for cost control. Preserve transcript `f5f6f15f-b816-485f-b134-63385afdc36e.jsonl.live.jsonl`; do not analyze it during restoration.
- #170 is primary active lane. Source registry already orders Debian1/2/3; installed registry now also exposes Debian1/2/3 after an existing deployment partially progressed. `task170-ssh-offload` is clean; do not invent or duplicate source changes.
- Existing deployment processes were observed at PIDs `2973613` (`.worktrees/claudex-proxy-slice/packaging/deploy-local.sh`), `3294793` (web production dependency deploy), and `3296106` (shared-checkout `packaging/deploy-local.sh`). NEVER launch another deploy or kill by name. Determine terminal state from these exact process trees.
- #170 remains incomplete until installed `local-gate --remote-only`, whole-agent execution, Debian3-saturation spill, authoritative exit status, and absence of laptop-heavy fallback are proven.
- #179/#233/#234 WIP is preserved losslessly in `.worktrees/incidents-179-233-234`. Required `specs/0592a6e6_incident-selector-completion.md` is absent from searched worktrees and Git refs. NEVER fabricate taxonomy or historical provenance. Recover authoritative semantics from durable session/report evidence after #170; then implement, review, land, deploy, and browser-prove.
- #236 focused security review follows the exact implemented incident delta; no prior review verdict exists.
- Use `gpt-5.6-terra` medium for non-review subagents when dispatch supports it. User explicitly required native subagents for current recovery; spawn no child explorers.

Stop emergency phase only when:

1. Real build and agent work runs on Debian1/2/3, not laptop.
2. Debian3 saturation spills to another reachable host.
3. `/incidents` is complete, deployed, and browser-proven: historical database backfill reconciles every source, every incident has canonical classification, and dispatch security review is closed.
4. Factory work is visible and fails loudly.
5. Protected runtime paths no longer reference `.worktrees/`.
6. Every remaining task is running remotely, genuinely externally blocked, or explicitly owner-deferred.

## Root causes

### 1. Installed SSH dispatch is broken and artificially single-host

- Debian1, Debian2, Debian3 are reachable by SSH/rsync.
- Installed `orders.build` and `orders.e2e` contain only Debian3.
- Debian3 allows three seats.
- Full Debian3 refuses new work while Debian1/2 stay idle.
- Installed `local-gate --remote-only` currently reports dispatch unavailable.
- `local-gate --remote-doctor` times out all three host checks, reports local Git-origin resolution failure, and finds Debian3 Node/pnpm parity drift.
- Direct remote CDX has succeeded on Debian2, proving remote execution itself still works.
- Fail-closed guards correctly prevent local fallback.
- Result: all ordinary work stalls despite idle cluster capacity.

Fix installed SSH transport and three-host spill FIRST. This is dominant restoration blocker (#170).

### 2. Live runtime has split provenance

- Runtime spans deploy-clone and worktree revisions.
- `~/.claude/hooks` still resolves into `.worktrees/muxhost-final`.
- Promotion cannot establish one trustworthy landed source.
- Result: fixes cycle through merge/review churn instead of reaching runtime.

Fix after SSH offload works. Do NOT block dispatch restoration on it.

### 3. Orchestration joined independent paths

- Hardening, k3s replacement, Git integration, broad gates, and restoration became one critical path.
- Three #178 test runs targeted wrong worktree.
- Long-lived branches repeatedly absorbed moving `main`.
- Broad reviews occupied scarce seats.
- Coordination churn, not core bug complexity, consumed elapsed time and model budget.

Split by seam. Reuse existing WIP. Run deterministic remote checks in parallel. NEVER duplicate review.

## Recoverable assets

| Work | Durable state | Emergency role |
|---|---|---|
| Incidents #179 | Older buildbox candidate commits `efd273b4f`, `4584c4d2d`; newer preserved `incident179-land` / `incident-dispatch-selector-options` WIP and `specs/0592a6e6_incident-selector-completion.md` exist; report showed a later active Debian3 worker | Reconcile all deltas before landing; never assume older commits are complete |
| Incident classification #233 | Abandoned despite incomplete historical classification | Recover taxonomy and classify every historical record |
| Incidents DB backfill #234 | Abandoned; complete historical registry remains unproven | Build idempotent provenance-preserving backfill and prove history |
| Incident dispatch security review #236 / #184 | Original exact-delta review failed before reading source; no verdict exists | Recover exact delta and complete one focused review |
| SSH restoration #170 | Direct CDX remote canary passed on Debian2; installed `local-gate` doctor remains broken | Wave 1 primary seam |
| Runtime safety #178 | Report records newer candidate `54068a20` merged with `origin/main`, focused deploy 26/26 and k3s/remote/controller 32/32; still not installed; older candidate `0fc2ea38` is stale evidence | Re-establish current authoritative revision before install |
| Factory transport #181/#182 | #181 destination-symlink fix preserved but remote rerun blocked by Debian1 ENOSPC; #182 prior baseline 177 passed with four unresolved P1s | Resume after spill/capacity restoration |
| Disk admission #183 | WIP preserved; Debian2 recovered to 18,042,288 KiB free in report | Resume after spill restoration |
| Attachable sessions #41/#148/#164 | Candidate refs conflict across report (`c6bc60ddd`, `2b27753b`); remote suites mostly green; live proof incomplete | Reconcile authoritative head in Wave 3 |
| UI availability #169 | `task169-release-controller` WIP preserved; prior local red is environment-invalid due PATH-shimmed `sleep`; no production change | Remote-verify after #170 |
| JPR-06 #153 | Local-only commit `40048087a`; remote focused tests 29/29; private `.platform` blocks full typecheck | Preserve completed local-only scope; do not accidentally deploy |
| Live report #125 | Current HTML is authoritative evidence but can become stale after each transition | Refresh after every real state change |
| Live runtime | Operational but revision-split; hooks worktree-backed | Keep running; normalize in #178 lane |
| k3s #175/#96 | Incomplete; current candidate architecture conflicts with non-negotiable target in live report | Wave 3 only; MUST NOT block proven SSH |
| k3s current WIP | `factory-agents-k3s` contains preserved uncommitted `ctr` digest lookup and provisioner-test edits from interrupted activation attempt | Preserve; do not use as Wave 1 dependency |
| Durable prompt journal #235 | `/home/user/.local/bin/prompt-journal` absent although prior task #121 was marked complete | Restore installed entrypoint and verify capture/retrieval |

## Reconciliation rules

Sources consulted:

- live report `/home/user/.local/share/overdeck/reports/2026-08-08-session-report.html` refreshed `2026-08-10T09:34:26+07:00`;
- recovered briefing `/tmp/fix-rot-incident-security-and-emergency-recovery.md`;
- owner correction naming abandoned incident classification and DB backfill;
- current task registry.

Resolve conflicts before implementation:

1. Treat installed/live evidence as authority over older candidate claims.
2. Treat newer commit/branch receipts as authority only after verifying object existence and diff scope.
3. Never mark #184 review complete: capacity recovery completed, but exact incident-dispatch review did not.
4. Never treat old #179 commits as complete while classification/backfill/security review remain open.
5. Fix-rot selected short transcript `b0cb1632-132c-4757-a39a-6c85891b9499`, not full three-day session `f5f6f15f-b816-485f-b134-63385afdc36e`; its security-review finding is valid but its task inventory is incomplete.
6. Preserve every contradictory WIP ref until authoritative superset is proven.

## Emergency policy

Cut coordination and ceremony, NEVER safety floor.

MUST:

- Restore proven SSH offload before k3s.
- Use remote-only execution. NEVER fall back locally.
- Reuse completed commits and WIP.
- Preserve atomic rollback.
- Run smallest focused seam test before live promotion.
- Restart exact service only.
- Browser-prove owner-visible UI.
- Install local infrastructure live before Git ceremony.
- Record remote host and authoritative exit receipt.

MUST NOT:

- Run broad gates before owner use.
- Launch duplicate reviews.
- Perform branch-wide merge archaeology; cherry-pick clean deltas onto fresh `origin/main`.
- Start QuietContext, tmux UX, reboot proof, refactors, or report redesign during restoration.
- Run heavy build, test, browser, or agent execution locally.
- Treat wrapper exit 0 as success when remote payload failed.
- Treat launched worker as running work.

```text
DO NOT: k3s → broad gates → Git → review → deploy → remote work
TARGET: three-host SSH spill → canaries → owner work → parallel delivery → k3s activation
```

## Wave 1 — Release remote work

Timebox: 15 minutes. Run lanes A–C concurrently. Do not start model-heavy work.

### Lane A — Installed SSH transport and Debian1 restoration (#170)

1. Reproduce installed `local-gate --remote-doctor` failures with authoritative per-host reasons.
2. Fix 30-second doctor timeout behavior, local Git-origin resolution, and required Node/pnpm parity without weakening capability checks.
3. Recover Debian1 above its disk-admission floor; preserve checksum-verified WIP before cleanup.
4. Add Debian1 to installed build and e2e spill order.
5. Run one real `local-gate --remote-only` canary.
6. Run one real `ca.sh --model composer-2.5` canary.
7. Record host and authoritative exit receipt.

### Lane B — Debian2 capacity restoration

1. Verify root free space stays above 15 GiB.
2. Verify agent-seat capability and stale-seat count.
3. Add Debian2 to installed build and e2e spill order.
4. Run one build canary and one agent canary.

### Lane C — Debian3 fail-closed spill proof

1. Preserve completed Incidents output.
2. Fill or observe Debian3 at its seat limit.
3. Dispatch new work and prove placement on Debian1/2.
4. Prove refusal/spill creates no local agent, build, or browser process.

### Wave 1 gate

Proceed only after all hold:

- `ca.sh` executes remotely.
- `local-gate --remote-only` executes remotely.
- Full Debian3 spills to Debian1/2.
- No heavy workload executes locally.

This gate restores normal work. k3s is not part of it.

## Wave 2 — Ship owner-visible and runtime work

Start lanes D–G concurrently immediately after Wave 1.

### Lane D — Incidents #179/#233/#234/#236

1. Inventory and preserve every Incidents delta before changing source:
   - older commits `efd273b4f`, `4584c4d2d`;
   - `incident179-land`;
   - `incident-dispatch-selector-options`;
   - `specs/0592a6e6_incident-selector-completion.md`;
   - any later Debian3 worker commit/receipt.
2. Diff these assets against fresh `origin/main`; select one authoritative superset. Never regenerate or blindly cherry-pick stale partial commits.
3. Restore incident classification (#233):
   - define one canonical taxonomy at collector trust boundary;
   - classify every historical incident consistently;
   - expose classification in `/incidents` and intake;
   - reject unknown values rather than silently coercing.
4. Complete Incidents DB backfill (#234):
   - inventory every historical source;
   - import missing incidents idempotently;
   - preserve source provenance and timestamps;
   - deduplicate deterministically;
   - produce source/count/classification reconciliation receipts.
5. Recover exact originally submitted incident-dispatch delta and complete focused security review (#236/#184). Trace intake → persistence → worktree/systemd launch → wrapper argv → status update → backfill. Verify authorization, idempotency, injection, prompt injection, path/symlink/TOCTOU, secret exposure, and fail-closed remote execution.
6. Run only missing remote collector, UI build/typecheck/component, classification, and backfill checks.
7. Resolve every confirmed P1/P2 in authoritative Incidents delta. Do not modify unrelated fixtures.
8. Land, deploy, smoke; rollback automatically on smoke failure.
9. Browser-prove:
   - complete historical registry and reconciled counts;
   - classification for every historical and new incident;
   - required intake fields;
   - CLI/model/effort/account options;
   - filing and dispatch lifecycle;
   - remote failure ends `engine-down`, never local execution.

### Lane E — Factory transport #181/#182

1. Resume preserved WIP; retain #181 destination-symlink rejection and copied/deleted case/fullwidth VCS regressions.
2. Restrict changes to extension transport and supervised fail-loud dispatch.
3. Reconcile #182 against prior remote Factory baseline of 177 passed and close all four independently reviewed P1 blockers; never weaken lifecycle or trace gates.
4. Run focused Factory tests remotely after Debian1 ENOSPC is resolved.
5. Install live before Git.
6. Prove one visible Factory job runs remotely with durable terminal status and recovery reference.
7. Land only after installed proof.

### Lane F — Disk admission #183

1. Resume preserved WIP.
2. Run disk-admission and maintenance tests remotely.
3. Install admission logic on buildboxes.
4. Prove low-space rejection and healthy-host spill.
5. Land after live proof.

### Lane G — Runtime source #178

Never block lanes D–F on this lane.

1. Remove or reconcile obsolete merged provenance fixture.
2. Run only:
   - `tests/os/deckctl-sync.test.sh`
   - `packaging/test-deploy-local.sh`
3. Direct-land and deploy.
4. Prove every protected runtime path resolves to one immutable landed bundle.
5. Prove no protected runtime path resolves into `.worktrees/`.

## Wave 3 — Activate replacements and recovery UX

Start only after Wave 1 proves healthy SSH offload.

### Lane H — Factory on k3s #175/#96

Non-negotiable target:

- Publish one immutable GHCR image accessible by every node.
- Use ephemeral Git input/result refs and Kubernetes Secrets.
- Report status through Kubernetes API and durable Factory trace.
- Use restricted namespace/pod posture.
- Constrain image, service account, refs, volumes, Secrets, command, repository, and security context.
- NEVER use SSH workspace fan-out, hostPath, node affinity, manual per-node image distribution, runtime PATH surgery, ConfigMap workspace archives, or pod-log result transport.

Execution:

1. Preserve incompatible current `factory-agents-k3s` WIP; extract only verified reusable contracts.
2. Close known source blockers:
   - Pod proxy subresource exposure;
   - verification TOCTOU;
   - Python `assert` safety checks disabled by `PYTHONOPTIMIZE`;
   - Git config transport;
   - trusted executable resolution;
   - fallback/double-run receipt states;
   - legacy controller reconciliation;
   - exact Job UID/pod ownership.
3. Activate ordinary Jobs with immutable published image.
4. Keep proven SSH runner as automatic fallback.
5. Require Job visibility and durable receipts in `/factory`.
6. Do not redesign scheduling beyond these contracts.

### Lane I — Attach/recovery #41/#148/#164

1. Verify existing tmux/systemd candidate on real buildbox.
2. Install and prove browser attach plus Ctrl+T recovery.
3. Do not change architecture.

### Lane J — UI availability #169

1. Resume `task169-release-controller`; preserve candidate identity manifest validation and immutable lockfile staging.
2. Discard prior local test verdict as environment-invalid; PATH-shimmed `sleep` broke lock-holder timing and no production state changed.
3. Verify service watchdog and installed release switching remotely after #170.
4. Fix only observed availability defects.
5. Prove installed UI behavior.

### Lane K — Durable task capture #235

1. Restore `/home/user/.local/bin/prompt-journal`; task #121's prior completion is contradicted by missing installed binary.
2. Verify owner prompts are captured once with session identity.
3. Verify `prompt-journal turns --session <id>` retrieves complete ordered turns.
4. Install before landing and prove the live hook path.

## Owner-deferred during emergency

Do not start these until emergency exit criteria hold:

| Tasks | Reason |
|---|---|
| #176, #219–225 QuietContext | Consumes model budget; does not restore work |
| #40 reboot rescue-door proof | Disruptive; owner deferred |
| #172 repeated session kills | Requires exact future PID/timestamp evidence |
| Broad refactors/full-suite cleanup | Hardening, not restoration |
| Historical report redesign | Update status only after real transitions |

## Budget controls

- Use `gpt-5.6-terra` at medium effort for every non-review subagent.
- Use review-specific agent/model only for code review.
- Launch no exploratory agents.
- Reuse existing commits and WIP.
- Allow maximum one coding agent per unresolved seam.
- Parallelize deterministic remote tests without model calls.
- Run one security review only for Incidents local-fallback boundary.
- After two failed iterations without new evidence, stop lane and preserve WIP.
- Do not repeat transcript reconstruction, broad audits, or advisory panels.
- Update live report only after real verdict, install, incident, or blocker transition.

## Execution receipts

Every lane transition MUST record:

- owner;
- exact seam;
- command/entrypoint;
- remote host or k3s node;
- authoritative exit result;
- installed live state;
- rollback artifact;
- next bounded action or genuine blocker.

Never mark a lane complete from code, candidate commit, wrapper status, or test result alone. Complete means landed and usable now.
