- Status: Accepted
- Date: 2026-08-02
oh-my-graph runs every node on the user's own logged-in claude subscription — that is the whole point of the tool (ADR 0001). Subscriptions have session limits, and a batch that is wide or long enough will hit one: the CLI's request is refused (HTTP 429 underneath) and the subprocess dies printing an error envelope whose message reads, verbatim as observed on claude 2.1.220:
You've hit your session limit · resets 5:20pm
Until now the engine treated this like any other non-zero exit: the node
FAILED, halt-on-fail cancelled its in-flight siblings, and the run reported
itself broken. Every part of that is a misreading. The node did nothing wrong
— its prompt never ran. The failure is not final — the limit resets at a
known time. And killing the siblings throws away work that was already paid
for. Repeated batch kills from exactly this shape are what motivated this
ADR. The failure-cause surfacing (#64) already carries the CLI's message into
NodeOutcome.FailureCause; this ADR makes the engine act on it.
Detection is string matching, and that is the honest option. The CLI
offers no structured signal for a session limit — unlike a --max-budget-usd
abort, which has the error_max_budget_usd envelope subtype, the limit
arrives only as prose (an is_error envelope whose result carries the
message, or a stderr tail). Matching prose is brittle: a wording change in
the CLI silently stops the match. The design accepts that with two
mitigations rather than pretending otherwise:
- The matcher lives in exactly one place —
internal/runner/sessionlimit.go— pinned by a test carrying the real message shape, so a CLI wording change is a one-line fix with a failing test to point at it, not a hunt across packages. - A missed match degrades safely. An unrecognized limit falls back to
the old behaviour — the node FAILs with the message in its detail — and
resume <run-id> --retry-failed(the same command the pause path recommends) still salvages the run. The brittleness costs convenience, never correctness.
A session limit pauses the run the way a gate does (ADR 0003), because it
is the same situation: the run cannot usefully continue right now, nothing is
broken, and a later resume should pick up exactly where it stopped.
- The runner classifies.
ClaudeCLIRunner.RunsetsNodeOutcome.SessionLimitedwhen the capturedFailureCause(envelope error report, else stderr tail) matches the limit's message shape — the same typed-flag patternBudgetExhaustedestablished, and the matcher's only call site. - The limited node is recorded nowhere. No ledger row, no snapshot
record, no terminal event, no retry attempt (an immediate re-spawn meets
the same limit). It mirrors how a gate pause records the paused gate: the
node is un-run, not FAILED — which is precisely what makes
resume --retry-failedre-launch it later, since it is in neither the completed nor the settled seed set. - The scheduler drains. On the first
SessionLimitedoutcome the run stops launching new work but lets in-flight siblings finish — never cancels them. A draining sibling may itself hit the limit; it then joins the paused set rather than failing. A limit outranks continue-on-fail pruning (the run is unfinished-and-resumable, not finished-and-failed; the retained FAIL records still tell the failures' story). A gate pause outranks a limit — the gate needs its human decision first, and the gate leg re-runs the un-recorded limited nodes anyway. Unlike a gate pause, noRecordPauseis needed: every settled node was already persisted per node, and the limited nodes' absence from the snapshot is their state. - The run exits resumable.
Scheduler.Runreturns*LimitPausedError;cmd/oh-my-graphmaps it to exit code 2 — the existing "paused and resumable" code — and prints the resume hint. The hint parses the message'sresets <time>best-effort and printsResume after 5:20pm with: oh-my-graph resume <run-id> --retry-failed; when the time cannot be extracted the hint prints without one rather than inventing a clock time from prose the CLI owns. resume --retry-failedfinishes the run. A limit-paused run may hold zero FAILED records (only PASSes plus never-run nodes), so the retry mode now also launches when there is un-recorded, launchable work — "running unfinished nodes" — while a run with genuinely nothing to launch (all passed, or blocked behind a rejected gate's standing decision) still exits 0 with no leg, and a gate-paused run is still redirected to--approve/--reject.- The event stream stays schema-stable. The leg closes with the existing
run_finishedoutcome"paused"; adetailfield — already defined on the event shape, previously unused onrun_finished— names the limited nodes and cause. Additive under docs/RUN-FEED.md's compatibility rule: no new event type, no new field, no schema bump.
Positive
- A batch that hits the limit stops cleanly with everything already earned persisted, and one copy-pasteable command (with the reset time when known) finishes it later. This converts the most common overnight batch death from "re-run and re-pay" into "resume".
- Draining preserves in-flight siblings' work exactly as a gate pause does, and siblings that also limit join the pause instead of racing to fail.
- Exit code 2 keeps its single meaning — resumable pause — so scripts watching for it need no change.
Negative / trade-offs
- Prose matching will eventually miss a reworded message. Accepted: the degradation is the pre-ADR behaviour, and the fix is one pinned pattern.
- A limit that arrives with no parseable envelope at all (rare; the observed shape is an error envelope) surfaces as a runner error and still fails the node — same safe degradation.
--retry-failedgains a second meaning ("finish unfinished work", not only "retry failures"). Accepted over a new flag: the operator story is one command that salvages a stopped run, and the exit hint promises exactly it.- Nothing in
state.jsonsays "limit-paused"; the signal lives in the exit code, the hint, and the stream'srun_finisheddetail. A consumer wanting it durable reads events.jsonl — the file that exists for history.
- A new retry cause (
retry: { on: [session_limit] }). Rejected: retrying into a standing limit spends nothing but achieves nothing; the limit resets on the clock, not on attempts. - Sleep in-process until the reset time. Rejected for the same reason a gate does not block (ADR 0003): oh-my-graph is a stateless CLI, not a daemon, and the parse of "5:20pm" is far too weak to sleep on.
- Mark the node FAILED but exit 2. Rejected: a FAIL record is a verdict
about the node, and it would be false — and
--retry-failed's clear-FAILED semantics would conflate real failures with limit victims in every surface that renders verdicts. - A new event type (
node_limited) or outcome token. Rejected: the event-type set is closed per schema version, so this forces a schema bump on every consumer for what one additivedetailfield already conveys.