Eight general-purpose starters connect a loop contract to a runtime. Three are dependency-light executables; five are copy/paste runtime templates with concrete prompts, schedules, permissions, state, and stop conditions. A nested worked implementation shows how to specialize these primitives without pretending every use case is a new runtime.
Each starter keeps the control loop visible: permissions, verification, state, and budgets remain explicit instead of disappearing inside a framework.
| Starter | Form | Best when | Trigger | State | External gate |
|---|---|---|---|---|---|
test-repair-loop.sh |
Executable Bash | A deterministic command is failing | Manual bootstrap | Markdown progress ledger | Check-command exit code |
threshold-monitor-loop.sh |
Executable Bash | A metric must stay above or below a boundary | Bounded polling cadence | Markdown sample receipts | Numeric threshold comparison |
queue-worker-loop.py |
Executable Python | Independent JSONL work items need bounded processing | Manual or scheduler | JSONL item receipts | Per-item verifier exit code |
Claude Code /loop |
Copy/paste template | You are present in an open coding session | Session interval | Session + progress file | Project check command |
| Claude desktop scheduled task | Copy/paste template | Local files need a recurring desktop task | Local schedule | Last-run marker + output | Live source checks |
| Codex automation | Copy/paste template | Background work belongs in an isolated worktree | Schedule | Worktree receipts + report | Declared repository checks |
| GitHub agentic workflow | Copy/paste template | Work starts from GitHub events or Actions schedules | Event or cron | Issue, PR, artifact, or cache | Required workflow checks |
| Shell / cron | Copy/paste template | An existing agent CLI needs minimal OS scheduling | Cron | Lock + progress file | Script exit code |
Use the runtime selection guide when persistence, file access, isolation, or permissions determine the choice. Vendor templates describe a portable operating shape, not a guarantee of current product behavior; confirm commands and scheduling semantics in the linked official docs.
This starter runs a deterministic check, sends only the failing evidence to an agent, and reruns the check until it passes, the evidence stops changing, or the budget expires.
# Claude Code
CHECK_CMD="pytest -x" AGENT_CMD="claude -p" ./test-repair-loop.sh
# Codex CLI
CHECK_CMD="npm test" AGENT_CMD="codex exec" ./test-repair-loop.shRun it inside a branch, worktree, or sandbox. The script edits nothing itself, but the delegated agent can.
| Contract part | Implementation |
|---|---|
| Objective | Make CHECK_CMD pass |
| Intake | Last EVIDENCE_LINES lines of failing output |
| Verification | CHECK_CMD exit code, judged by the script |
| State | LOOP_PROGRESS.md |
| Budget | MAX_ITERATIONS, default 5 |
| Escalation | Repeated evidence or exhausted budget returns non-zero |
This read-only starter samples a command that prints one number. It records each sample and exits on the first boundary breach. An optional agent can diagnose the evidence, but the script never remediates.
PROBE_CMD="./scripts/p95_latency_ms.sh" \
THRESHOLD=250 \
DIRECTION=max \
MAX_SAMPLES=12 \
INTERVAL_SECONDS=300 \
AGENT_CMD="codex exec" \
./threshold-monitor-loop.shSet DIRECTION=max when values above the threshold are bad, such as latency or spend. Set DIRECTION=min when values below it are bad, such as pass rate or availability.
| Contract part | Implementation |
|---|---|
| Objective | Keep a measured signal inside its declared boundary |
| Intake | Numeric final line from PROBE_CMD |
| Verification | awk compares the value with THRESHOLD |
| State | MONITOR_PROGRESS.md |
| Budget | MAX_SAMPLES and INTERVAL_SECONDS |
| Escalation | Probe error or first threshold breach |
This starter reads JSONL work items, skips IDs already marked complete, delegates one item at a time, runs an external verifier, and persists every attempt.
Minimum queue row:
{"id":"docs-101","objective":"Update the CLI install example","allowed_paths":["README.md"],"verification":"python3 scripts/check_docs.py"}Validate before invoking an agent:
python3 queue-worker-loop.py --queue tasks.jsonl --dry-runProcess a bounded batch:
python3 queue-worker-loop.py \
--queue tasks.jsonl \
--agent-command "codex exec" \
--verify-command "python3 verify_item.py {id}" \
--max-items 3 \
--max-retries 2 \
--command-timeout 900The verifier receives the current item ID through both {id} substitution and the LOOP_ITEM_ID environment variable.
| Contract part | Implementation |
|---|---|
| Objective | Each queue row's objective |
| Intake | Pending JSONL rows after completed IDs are removed |
| Verification | --verify-command exit code |
| State | QUEUE_PROGRESS.jsonl |
| Budget | --max-items, --max-retries, and --command-timeout |
| Escalation | A work item exhausts its retry budget |
A timed-out agent or verifier consumes one attempt and leaves a receipt before the item retries or escalates.
- The maker does not check its own work. External commands or thresholds decide completion.
- State lives outside the model. Markdown or JSONL receipts survive cold agent calls and reruns.
- Completed work is idempotent. Queue IDs and last-run markers prevent duplicate processing.
- Evidence stops waste. Repeated failures and threshold breaches stop or escalate instead of consuming an open-ended budget.
- Permissions remain visible. The starters expect a branch, worktree, sandbox, or read-only boundary rather than pretending isolation is automatic.
docs-drift/docs-drift-loop.py specializes the shell / cron starter for the docs-drift pattern and validated contract. The detector, acting agent, and verifier remain separate:
- the detector exits
0when docs are current,1with evidence when drift is confirmed, and any other code on detector failure; - report-only mode stores an evidence receipt and exits
2for human triage; - patch mode gives the evidence to an agent, enforces allowed and generated paths plus a file-count budget, then lets an independent command decide success;
- every outcome is appended to
.loop-state/docs-drift.jsonl, which is ignored by Git but survives model calls; - the script never commits, pushes, merges, changes a verifier, or reverts an unsafe edit behind the operator's back.
Run the detector without invoking an agent or writing state:
python3 examples/runnable/docs-drift/docs-drift-loop.py \
--discover-command "python3 scripts/check_project_consistency.py" \
--dry-runProduce a scheduled evidence-backed report:
python3 examples/runnable/docs-drift/docs-drift-loop.py \
--discover-command "python3 scripts/check_project_consistency.py" \
--report-onlyPatch confirmed drift inside a clean branch or worktree:
python3 examples/runnable/docs-drift/docs-drift-loop.py \
--discover-command "python3 scripts/check_project_consistency.py" \
--agent-command "codex exec" \
--verify-command "python3 scripts/check_project_consistency.py" \
--allowed-path README.md \
--allowed-path docs \
--allowed-path meta \
--max-attempts 2 \
--max-changed-files 8 \
--command-timeout 900Copy docs-drift/github-actions.yml into .github/workflows/ for a read-only Monday schedule that uploads the JSONL receipt. For cron, run the report-only command from the repository root and send exit code 2 to the docs owner. Add --generated-path <path> for files that must be regenerated or reviewed rather than edited directly.
Run these without an agent account:
bash -n test-repair-loop.sh threshold-monitor-loop.sh
python3 -m py_compile queue-worker-loop.py
python3 ../../scripts/check_runnable_examples.py
printf '%s\n' '{"id":"demo","objective":"Validate the queue"}' > /tmp/loop-queue.jsonl
python3 queue-worker-loop.py --queue /tmp/loop-queue.jsonl --dry-run
PROBE_CMD="printf '42\\n'" THRESHOLD=100 MAX_SAMPLES=1 ./threshold-monitor-loop.sh- Replace the verifier first; it defines what "done" means.
- Map allowed paths, tools, credentials, and network access from the chosen JSON contract.
- Keep mutable work away from the benchmark, policy, or tests that judge it.
- Choose a durable state artifact that the next cold run can read.
- Start with one item or iteration, then raise budgets only after reviewing receipts.
- Test failure, timeout, duplicate work, and human escalation before scheduling unattended runs.