move-codex-session cannot finish a move that fails after the source commit #142
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
At
1a85c77, a failure after the source database commit leaves the session complete in the destination and partly removed from the source, and rerunning the script is refused. Nothing finishes the move.mainsetssourceDeletionStartedbeforedeleteSourceDatabases, and thecatchskips the destination rollback once it is set. After the commit (line 1858) that is right, because the destination is the only complete copy. Four steps follow the commit:rewriteIndex(line 2260), the loop that removes each exclusively owned source rollout (lines 2261-2267),verifySourceRemoved(line 2268) and the lock release. If one throws, or the process is killed, the source keeps whatever it had not reached.A rerun cannot continue.
getThreadIdsreads the sourcethreadstable and fails withsource state is missing thread <id>(line 323).SKILL.mdsays to resolve the blocker the error names instead of editing the databases, but this error describes the result of the first run.Two ordinary failures reproduce it on a fixture, with no code edits:
chmod 0555on a source rollout directory makesrmSyncfail withEACCES.chflags uchgon the sourcesession_index.jsonlmakes the rename inrewriteIndexfail withEPERM.Both runs exit 1 with empty stdout. The source databases no longer hold its threads, the source keeps all seven fixture rollouts (the six that were copied are byte-identical to the destination's), and the destination is complete. After undoing the permission change, the same command exits 1 with
source state is missing thread 11111111-1111-7111-8111-111111111111.No data is lost, since the destination was verified before the commit. The user is left with the session in both homes, and the leftover rollouts can bring it back at the source: when a lookup finds a rollout whose thread has no row, Codex rebuilds the row from it (
read_repair_rollout_pathincodex-rs/rollout/src/state_db.rs, Codex 0.160.1). The window is short and the causes are uncommon: permissions, an I/O error, a kill, a power loss. A commit across attached WAL databases is not atomic (SQLite), so the leftovers can include database rows. In an earlier disk-full run (about 1 MB free) the state rows were committed away while the source kept its history, logs, goals, memories and queue rows.A failure before the commit is a separate case: the source is still intact, so the destination copy can be rolled back instead of resumed.
A fix would let a rerun detect the moved threads missing from the source
threadstable but present in the destination. It would read the thread IDs from the destination's spawn edges and compare each leftover source rollout (by hash) and row with its destination copy. Matching leftovers would be deleted, the index rewritten, and the usual JSON printed. Anything that doesn't match stays, and the error names it.Evidence: reproduced at
1a85c77for the two failures above. The Codex row rebuild is from reading Codex 0.160.1. The partial commit was seen once on the branch that became #132 and not repeated at1a85c77.