Local-disk tiering: a disposable working tree, git as the sync
Table of Contents
Design note. Written on a carleton-htc slice (job 38661192,
n0036.savio4) against main at 80581c8. Status: step (1) landed
(shared timer + snapshotter, PR #9); step (2) landed for confined
targets (PR #10) and for salloc targets, retiring the all-on-/local
layout (sucoder/local_tier.py, slurm.local_disk / --local-disk);
step (3) landed: the prelude carries a WORKSPACE (local-disk tiering)
block whenever tiering is on (MirrorManager._workspace_block).
Punchline
Put the agent's working tree, index, and every cache on the compute
node's /local/job$SLURM_JOB_ID/, keep the existing shared-filesystem
mirror as the durable repository and as the laptop's push/pull target,
and let git do the copying: a post-commit hook pushes each commit to
the shared mirror the instant it exists, and a periodic snapshot of
the dirty tree lands on a scratch ref so uncommitted work is never
more than a few minutes from durable. Slurm's epilog wipes
/local/job$ID for us, so there's nothing to clean up and nothing to
orphan.
This replaces the current --local-disk mode rather than sitting
beside it. That mode moves the whole mirror to /local/mirrors,
which is why it's orphan-prone (cli.py:1011), disabled for confined
targets (cli.py:605), and needs chmod 700 hygiene on a directory
that outlives the job (mirror.py:1741). Once commits reach the
shared mirror by themselves, the all-on-/local layout has no
remaining advantage.
What the filesystems actually do
Three tiers are visible from a compute node, and it's worth being
precise, because the folk description ("NFS doesn't lock, Lustre is
slow, /local is fast but transient") is right but incomplete.
| Path | FS | flock | Shared | Survives job | Note |
|---|---|---|---|---|---|
$HOME (/global/home) |
NFS | no | yes | yes | uv, npm, sqlite all break here |
~/mirrors -> /global/scratch/... |
Lustre | yes | yes | yes | slow metadata; locks fine |
/local/job$ID |
ext4 | yes | no | no | 1.7 TB, prolog-created, 700 |
Measured 2026-09-09 with fcntl.flock on a temp file in each
directory: $HOME fails with ENOLCK ("No locks available"), the
other two succeed. Two consequences follow.
- The mirror is on Lustre, not NFS
~/mirrorsis a symlink into/global/scratch/fsa/fc_jevons/ligon/mirrors. So "the NFS mirror" in the prompts and docs is really a Lustre mirror, and git's own locking (O_EXCLlockfiles, neverflock) would have been fine either way. What hurts on Lustre is working-tree I/O: test suites,node_modules, virtualenvs, anything with many small files.- The lock failures are a
$HOMEproblem, not a mirror problem - uv
refused to build a venv in this session because
~/.local/share/uv/.lockis on NFS;codex --versionprints the sameos error 37. Whoever setUV_TOOL_DIR=/local/job38661192/uv-toolswas dodging exactly this, at the cost of making the tool installation transient. Tiering the mirror doesn't fix this; see Open questions.
Layout
- Shared mirror (unchanged)
~/mirrors/<name>, a non-bare repo withreceive.denyCurrentBranch=updateInstead(set by the launcher atmirror.py:1764). The laptop pushes into it and pulls from it exactly as today. Nothing in_sync_remoteor_pull_from_remotechanges.- Working clone (new)
/local/job$ID/mirrors/<name>, a fullgit cloneof the shared mirror withoriginpointing back at it by path. Objects, index, working tree, and hooks all on ext4. Its$HOME-independent caches (UV_CACHE_DIR,npm_config_cache,PIP_CACHE_DIR,.gitnexus, pytest) go under/local/job$ID/cache/.- Rule
- once this mode is on, nobody edits the shared mirror's
working tree. It is a mailbox, not a desk. The README's invitation
to edit via the
-lnTRAMP alias (README.org:541) has to change accordingly; the shared tree stays clean, which is also what makes the laptop'supdateInsteadpush succeed unconditionally (today it is refused whenever the agent's tree is dirty).
A full clone rather than a git worktree: a worktree keeps its index
and HEAD under $GIT_DIR/worktrees/<name>/ on the shared
filesystem, which puts the hot path back on Lustre. Clone cost is a
few seconds for repositories of the sizes we run.
Sync on commit
A post-commit hook in the working clone runs
branch=$(git symbolic-ref --short -q HEAD) || exit 0 # detached: nothing to publish git push --quiet origin HEAD:refs/heads/"$branch"
with no force. Because the shared mirror uses updateInstead, its
checkout advances too, so the committed state is visible from the
login node and to TRAMP without any extra machinery. That answers
the "does it need to be visible between sessions" question for free.
The one race is with the laptop. _sync_remote force-pushes into
the shared mirror after pulling the agent's commits; if the laptop
lands new commits while the agent is mid-task, the agent's next hook
push is rejected as non-fast-forward. The hook must never force (it
would clobber the laptop's push); it prints the rejection loudly, the
agent runs git pull --ff-only, and commits again. Today's
divergence handling at the shared tier (mirror.py:1115) is
untouched.
Amends and rebases don't fire post-commit for every rewritten
commit, and a rewritten branch would be rejected by the hook anyway.
That's acceptable: the snapshot below covers the gap, and rewriting
published history on the mirror was already the laptop's decision to
make, not the agent's.
WIP snapshot
Every wip_snapshot_minutes (default 10) and on every deadline
warning, a snapshotter does, in the working clone,
marker=$(git rev-parse --git-dir)/sucoder-last-wip-tree # outside the tree, or add -A eats it export GIT_INDEX_FILE=$(mktemp) git read-tree HEAD git add -A # tracked + untracked; honours .gitignore tree=$(git write-tree) [ "$tree" = "$(cat "$marker" 2>/dev/null)" ] && exit 0 wip=$(git commit-tree "$tree" -p HEAD -m "WIP snapshot $(date -Is) job $SLURM_JOB_ID") ref=refs/sucoder/wip-job/<mirror>/$SLURM_JOB_ID git update-ref "$ref" "$wip" git push --quiet --force origin "$ref" echo "$tree" > "$marker"
Points that matter:
- It's a scratch ref, not a branch
_pull_from_remoteandsucoder pulllook only at the base branch, so their semantics are untouched. Forcing is safe for the same reason it's safe forrefs/sucoder/mirror-head(mirror.py:1128): the ref carries no history anyone depends on.- Unchanged trees cost nothing
- comparing tree hashes makes the
idle case a
write-treeand a string compare. - One ref per (mirror, job)
- the ref is keyed on the job id as well as
the mirror. It originally was not, and a mirror running in two jobs at
once – two clones, on two nodes, with two different dirty trees – had
one slot to store them in, so each snapshot force-pushed over the
other's (issue 19).
refs/sucoder/wip-job/<mirror>/<job>is a sibling of the oldrefs/sucoder/wip/<mirror>, never a child: git refuses to create a ref below an existing one, and--atomicdoes not rescue a delete-plus-create, so nesting under the old name would have needed a window in which the mirror held no snapshot at all. - Restore on relaunch is conditional
- rebuild the working clone, and if
a candidate snapshot exists and its parent equals the branch tip,
apply its tree to the working directory (
git read-tree -m -u <wip>followed bygit resetto unstage) so the agent finds the same dirty tree it left. If commits have landed since the snapshot, its parent no longer matches; leave it alone and name it in the handoff. - Restore never takes a live job's tree
- two jobs on one mirror are
normally on the same branch at the same commit, so "parent equals the
tip" is true for the other job's snapshot too, and the restore used
to import it and announce it as resumed work with nothing saying whose
it was. Candidates are therefore filtered by their job: this job's own
snapshot first (an idempotent re-run of prepare inside one job), then
the newest whose job
squeuereports has ended. A failedsqueueis unknown state, not "gone" – the same rule the launcher uses – so it declines. The legacy shared ref is still read, since a pre-upgrade timer on another live job keeps writing it; its job id is recovered from the snapshot's subject. Wheresqueueis not onPATHat all the restore proceeds but says so, because declining outright would break every relaunch in the environment issue 15 describes. - Ended jobs' snapshots are retired
- per-job refs would otherwise
accumulate one per allocation forever, so prepare deletes the ones whose
jobs
squeuereports have ended – keeping the newest as a fallback, since this job has not taken its own snapshot yet and will not for up towip_snapshot_minutes. A live job's snapshot is never touched, and neither is one whose job could not be checked: deleting on an unknown answer is the same mistake as restoring on one, and it deletes the only durable copy of somebody's work. When a snapshot is kept for that second reason, prepare says so on itsSUCODER:line and names how many. Retention is the only thing bounding the ref set, and a silent refusal to retire looks exactly like having nothing to retire, so an environment where the scheduler cannot be reached from inside a job would return to the unbounded accumulation of issue 14 with nothing on screen. - A clean release retires its own snapshot
sucoder releasedeletes the ref for the job it cancels, in the same round trip asscanceland only if that succeeded. This is the one rule keyed to what actually decides a snapshot's value: it exists for a job that died with uncommitted work in a clone that is now gone, and a job released on purpose is not that job – the work was committed or deliberately abandoned, whether five minutes or twelve days in. Age and count rules only approximate that, badly: a three-day-old snapshot from a crashed job is precious and a three-minute-old one from a clean exit is garbage. Only that job's ref goes, never a glob over the mirror, and never on the sibling-detach path, where the job survives for the other mirrors on it. The release names what it deleted, hash included, since the commit stays in the object store until the mirror's next prune;--keep-wipleaves the ref alone. It cannot be the only rule, because it is cleanup keyed to an exit path that often does not run:scancelby an admin, walltime, a node failure and an OOM kill all bypass it.- The sweep that needs no launch
- retention in prepare runs only
when a job is launched against that mirror, so a mirror out of use kept
its refs forever.
sucoder snapshotslists every snapshot on every configured cluster from the laptop and, with--retire, deletes those whose jobsacctreports ended more than a recovery window ago (seven days by default; a fact about how long a human takes to notice, which does not scale with job length). Ref age is only the fallback, for a job accounting cannot place, and then against the cluster's longestslurm.timeplus the window, which a live job cannot exceed; a target with no finite--timeleaves no safe bound, and those refs are kept and the reason printed. Deleting on an unanswered query is the one outcome worse than keeping too much. - Nothing runs
git gcon the mirror - deleting a ref frees nothing by
itself. Every snapshot is
commit-tree -p HEAD, so each one already orphans its predecessor whatever the ref is called; retiring the refs is what makes those orphans collectable at all. Thegc --autothatreceive-packalready runs on every hook push then reaps them on the mirror's owngc.pruneExpire, which bounds the pile at one expiry window of churn instead of forever. The default window is two weeks; a mirror whose outputs churn hard can shorten it withgit config gc.pruneExpire 2.days.agoin the mirror. That is deliberately the human's setting: an automatic gc from the launcher would take a lock on a Lustre repository at every launch, and on the workloads measured so far it would find nothing to do (10.51 MiB unreachable against 841 MiB packed, well undergc --auto's own threshold). The cheapest real fix is upstream of all of it:.gitignoregenerated outputs, so the churn never enters a snapshot. - What it doesn't cover
- ignored files (
.venv,node_modules,.gitnexus) are rebuilt, which is the point, since they're the flock-hostile things. The staged/unstaged distinction is lost on restore. Both are worth a sentence in the agent prompt.
Who runs it
The snapshotter is a loop that already exists in spirit: the deadline
timer (cli.py:1230) polls squeue every 60 s and writes warnings.
Extend that script with the snapshot step, and start it from the
right place for each launch mode:
- Confined (sbatch)
- the batch body (
mirror.py:2368) currently doescd <mirror>, starts tmux, and keeps the job alive. Add: clone to/local/job$ID/mirrors/<name>, install the hook, restore any snapshot,cdthere instead, and background the timer/snapshotter before the keeper loop. It inherits the cgroup for free. - Unconfined (salloc)
- the same script replaces what
_start_slurm_timerwrites today; the clone step goes into the launch path that computesremote_mirror_root(cli.py:605-640).
This also closes a gap I found while checking: no slurm-timer.sh
is running for this confined job. _start_slurm_timer is only
reached from _ensure_slurm_node (cli.py:1130), which the confined
path skips, so the watchdog that prompts/carleton-htc.org promises
has never fired on this target. One script for both modes fixes
that as a side effect.
What the agent is told
The prompt's "commit early and often" wants a concrete reason, so state the durability tiers:
- Committed
- durable the moment
git commitreturns; visible from the login node immediately. - Dirty, tracked or untracked
- durable to within
wip_snapshot_minutes; restored at relaunch if no commits intervened. - Ignored
- never durable; rebuilt from scratch each job.
The prelude is rendered before launch, so it cannot print the last
snapshot time; it gives the one-line git log that answers it
instead, plus the clone and mirror paths, the mailbox rule, and the
warn-file path.
Config and docs
slurm.local_disk- keep the key, change the meaning to this
mode.
trueor a path selects the/localroot; the old all-on-local layout goes away. Addslurm.wip_snapshot_minutes. PerCLAUDE.md, that touchesREADME.org:695andconfig.example.yaml:84. - Session state
session.remote_mirror_root(session.py:30) and_forget_allocation(cli.py:2850) special-case a/localroot precisely because it could be orphaned. With the shared mirror always authoritative, both simplify: the saved root is only ever the shared one, and--nodepinning stops mattering.- Caveats to rewrite
persistent-presence.org:371("orphaned on re-grab") andconnection-faq.org:306("gone with the node") both describe the old layout.- The handoff note moves
prompts/carleton-htc.organdpersistent-presence.org:466tell the agent to write~/mirrors/<mirror>/.sucoder/handoff.org, i.e. into the shared mirror's working tree. Under the mailbox rule that dirties the shared tree, and the next hook push is refused byupdateInstead("Working directory has unstaged changes") at exactly the moment you least want a confusing failure. The handoff belongs in the working clone, committed or caught by the snapshot; it reaches the shared mirror like everything else.
Open questions
- Transcript resume :: Claude Code keys transcripts by the realpath
of the cwd (this session's key is
-global-scratch-fsa-fc-jevons-ligon-mirrors-SuCoder, the resolved symlink target). A cwd of/local/job$ID/...changes every job, soclaude --resumeafter turnover would need either a stable bind path or a copy of the transcript directory. Needs a spike; a symlink alone won't do it. - Tool installs on NFS :: caches belong on
/local; tool installs (UV_TOOL_DIR, the npm prefix) must persist and so belong on Lustre or NFS, but uv's install-timeflockfails on NFS. Lustre locks fine, soUV_TOOL_DIRunder/global/scratchmay be the whole answer. Separate from this design; noted so the transient/localsetting doesn't get cargo-culted. - Snapshot retention :: last-only (force-updated ref) is simplest. A reflog on the shared mirror would keep a few generations at no cost; decide whether anyone would ever look.
- Several mirrors per job ::
/local/job$ID/mirrors/<name>already allows it; the snapshotter should iterate rather than assume one.
Spike results (2026-09-09, job 38661192, n0036.savio4)
Run by hand against this mirror (~/mirrors/SuCoder, on Lustre)
and a throwaway copy of it on /local; no sucoder code changed.
Scripts: a 9-line post-commit hook and a 20-line snapshotter,
both as sketched above.
| Step | Result | Time |
|---|---|---|
Full clone Lustre -> /local (4.1 MB .git) |
ok | 3.5 s |
Commit on spike/local-tier, hook push |
branch appeared on mirror | 1.15 s |
Push to checked-out main, mirror tree clean |
mirror HEAD + worktree advanced | -- |
| Same, mirror tree has a tracked edit | refused: "unstaged changes" | -- |
| Same, mirror tree has an untracked file | accepted | -- |
| Laptop commit lands first, agent pushes | refused non-FF; pull --ff-only fails as diverged |
-- |
| Snapshot, clean tree | no-op | 0.05 s |
| Snapshot, 1 tracked edit + 1 untracked + 1 ignored | ref pushed; 2 files in diff, ignored excluded | 1.27 s |
| Snapshot again, unchanged | no-op via marker | 0.05 s |
| Restore into fresh clone | git status and tree identical |
-- |
| Restore when a commit followed the snapshot | parent != tip, skipped | -- |
Three details the sketch got right only after review: the marker file
must live under .git/ or add -A sweeps it up and the unchanged
check never fires; the hook must exit quietly on a detached HEAD;
and the snapshotter needs a "tree equals HEAD's tree" check so a
clean tree isn't snapshotted as an empty WIP commit.
The mirror's checked-out main and working tree were never touched.
Left in place on the shared mirror for inspection: branch
spike/local-tier (one commit, this note) and
refs/sucoder/wip/SuCoder; delete both when done with them.
Everything under /local/job38661192/ goes with the job.
Next step
The spike is done; the mechanism works as designed. Implementation
order that keeps each step shippable on its own: (1) one timer/
snapshot script started from the confined batch body, which also
fixes the missing watchdog; (2) the /local clone and hook in the
launch paths, behind slurm.local_disk; (3) docs and prompt updates,
including moving the handoff into the working clone.