-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathreplicator.service
More file actions
219 lines (207 loc) · 12.6 KB
/
Copy pathreplicator.service
File metadata and controls
219 lines (207 loc) · 12.6 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
[Unit]
Description=replicator — retrieval, fingerprinting, and temporary storage layer for the Cannabis Observer cluster
# The redis-server ordering is GONE, not loosened (CannObserv/broker#1 Phase 3).
# The change bus moved to a neutral node reached over the tailnet, so there is
# no local unit to order behind.
#
# Removing `Wants=` is the load-bearing half. It is a *dependency*, not just an
# ordering: while it stood, starting this worker pulled the retired local
# redis-server back up — observed during the cutover, after the broker's data
# had already been copied away. A resurrected empty broker that nothing points
# at is harmless; one that something later points at is not.
#
# The reasoning it replaces still holds: the broker is cluster infrastructure
# this worker only consumes, its absence must never hard-fail the start, and the
# floor check below remains the guard that actually blocks.
#
# tailscaled is the ordering that replaced it (#88): the broker is `broker` on
# the tailnet. Ordering ONLY — no Wants=/Requires=/BindsTo=. Those propagate:
# tailscaled restarting under an apt upgrade would restart the worker and spend
# one of the three starts an hour below. After= orders against tailscaled
# STARTING, and on a cold boot MagicDNS answers `broker` with no address for a
# moment after that - measured twice - so the floor check below retries an
# unreachable broker for a bounded window rather than trusting this ordering.
# A tailnet that stays down is still the worker's backoff's job.
After=network.target tailscaled.service
# Bound restart loops at BOTH timescales, then stay failed (visible in
# systemctl status + OnFailure=) instead of restarting forever.
#
# The window is long, not five minutes, because the worker's slowest failure
# mode is deliberately slow: it absorbs broker failures for
# REPLICATOR_MAX_CONSECUTIVE_CYCLE_FAILURES cycles (~10 min) before exiting, so a
# permanently unreachable Redis produces one exit every ~10 minutes. A 5-minute
# window never sees two of those and the unit would read as active (running)
# forever, with the failure buried in the journal. Six strikes catches that case
# and still catches a fast crash loop within seconds (6 x RestartSec=5 = 30s).
#
# SIX, NOT THREE, AND TWO HOURS, NOT ONE (#94). Three starts absorbed 30 minutes
# of broker outage. On 2026-09-16 the broker was away for 58 — nearly twice the
# budget — and the only reason this unit was not asked to survive the whole of it
# was an accident of timing: the network degraded from ~14:28 but the worker's
# cycles did not fail *continuously* until ~15:00. Six starts absorb an hour, and
# tests/test_deploy.py now pins that against the worst outage this cluster has
# actually had rather than against internal consistency alone.
#
# Widening tolerance is safer than it was, because the other half finally exists:
# until #94 there was no OnFailure=, so a unit that gave up was discovered by
# whoever next looked. A longer fuse on a failure nobody is told about would be
# the wrong trade; a longer fuse on one that pages is the right one.
StartLimitIntervalSec=7200
StartLimitBurst=6
# The other half of "stay failed": tell someone. The comment above claimed this
# directive existed for months while the ini file carried no OnFailure= at all,
# and on 2026-09-16 the unit sat failed for 56 minutes — the notice came from a
# sibling repo watching the broker from another VM, not from this host (#94).
#
# `%n` is the full name of the unit that failed, so one template serves any unit
# pointing here and the notification can say what broke. The handler is oneshot,
# carries no OnFailure= of its own, and exits 0 on every path: it fires because
# something already failed, so it must not be able to fail into itself.
#
# Firing is not delivery. With no REPLICATOR_NOTIFY_URL configured the handler
# writes a CRITICAL journal record and stops there — deliberately, so this lands
# without waiting on a notifier channel, and so a notifier outage (which
# correlates with the broker outages that fire this) degrades to the record
# rather than to nothing. `deploy/replicator-failure-notify@.service` carries the
# rest.
OnFailure=replicator-failure-notify@%n.service
[Service]
Type=simple
User=exedev
WorkingDirectory=/home/exedev/replicator
# Be the LAST thing this VM's OOM killer chooses, not the first (#92).
#
# co-replicator was 3.9 GB with no swap then (8 GiB + 4 G since #99), and it is
# also the dev workspace: VSCode Server, Claude Code, and any MCP server they
# start share this VM with the worker. At the default adj of 0 the worker
# scored 670, second from the top of the eligible list; -900 puts it last but
# tailscaled, whose -950 keeps the bus's only path alive longer (#112).
# Until #125 the dev tooling was not on that list at all — exe-init 8579326
# started sessions at oom_score_adj=-1000, *exempt* — so the kernel would have
# killed the worker to relieve pressure the tooling created. exe-init 14fd603
# starts them at 0, and this directive now wins a real comparison.
#
# That is not hypothetical for this cluster. CannObserv/broker#17: launching a
# SocratiCode server on the broker's VM degraded the network path for 57m48s
# with NOTHING OOM-killed — the kernel failed *atomic* allocations in
# tailscaled while every process stayed alive — and this worker did not
# reconnect on its own (#94). A capped launch is the other half of the fix and
# belongs to whoever runs the launch (docs/COMMANDS.md); this half is the unit's.
#
# -900, not -1000: the cohort's value (CannObserv/broker#25). Last to be chosen,
# not exempt — an exempt worker that leaks is unreclaimable, and the kernel
# would work through everything else on the box before touching it.
# tests/test_deploy.py pins both bounds.
#
# earlyoom is declined (#112): at -1000 it could not reach the sessions, and at
# 0 (#125) the kernel's own order already takes them first, so it would add
# only an earlier kill. tests/test_deploy.py::TestTheSessionPremise pins the 0
# live; TestTheEarlyoomDecline keeps earlyoom off the host.
OOMScoreAdjust=-900
# A reservation, never a cap (#113): keeps the worker's working set out of
# reclaim, so pressure falls on init.scope — the sessions — instead. 66 MiB
# anonymous and a 67 MiB peak, measured 2026-09-24, so 128M. Inert without the
# system.slice grant in deploy/system.slice.d/, which must cover this plus
# tailscaled's; tests/test_deploy.py checks the sum. Nothing for *atomic*
# allocations — vm.min_free_kbytes covers those.
MemoryLow=128M
# Creates /var/lib/replicator (REPLICATOR_BLOB_DIR's parent) owned by User=.
StateDirectory=replicator
# Create runtime dir (runs as root before User= takes effect with + prefix)
ExecStartPre=+/bin/bash -c 'mkdir -p /run/replicator && chown exedev:exedev /run/replicator'
# Hard guard: "Code committed to main is the deployed code" (AGENTS.md) was an
# invariant nothing enforced until #37 — a restart while the checkout sat on a
# feature branch deployed unmerged code. Refuses a non-main or detached HEAD, and
# a main carrying commits that were never pushed (#48) — unpushed is unshared, and
# unlike 'behind' that verdict does not need the cached origin/main to be fresh.
# Warns only on a stale or dirty tree. Not '-' prefixed: a guard whose refusal is
# logged and ignored is not a guard.
#
# Placed BEFORE the BUILD_ID stamp on purpose. /run/replicator/build-id outlives a
# failed start, so stamping the branch SHA first would leave the journal
# describing code that never ran — the same "looks correct, is not" failure the
# guard exists to close.
#
# REPLICATOR_ALLOW_ANY_CHECKOUT=1 in /etc/replicator/.env bypasses it. That is not
# the documented way to test a branch on this VM (run the worker directly under
# distinct REPLICATOR_CONSUMER_NAME and REPLICATOR_REPLICATE_CONSUMER_NAME, one
# per group) — it is for testing this unit itself. The
# override still reaches this step even though the EnvironmentFile= lines appear
# below it: EnvironmentFile= is a unit-level setting applied to every Exec*
# process, not a sequential instruction, so its position in the file is not an
# ordering.
#
# The checkout it inspects is WorkingDirectory= (systemd's cwd for every Exec*),
# which is the same tree the stamp below reads and ExecStart runs from — so the
# guard cannot end up verifying a different tree than the one that starts.
ExecStartPre=/bin/bash /home/exedev/replicator/scripts/check_main_checkout.sh
# Write current git SHA to a runtime env file before starting
ExecStartPre=/bin/bash -c 'echo BUILD_ID=$(git rev-parse --short HEAD) > /run/replicator/build-id'
# Load BUILD_ID, then system secrets. Order matters: later EnvironmentFiles
# override earlier ones, so BUILD_ID (read-only state from ExecStartPre) is the
# lowest-precedence layer.
#
# The repo-local .env is deliberately NOT loaded. It holds dev/agent secrets —
# GitHub PATs with org-wide write access that this worker never needs — and
# sourcing it would hand every one of them to a process whose job is fetching
# public URLs. Production configuration belongs in /etc/replicator/.env.
EnvironmentFile=-/run/replicator/build-id
EnvironmentFile=/etc/replicator/.env
# Hard guard: Replicator is the cluster's first user of claim_stale, which needs
# XAUTOCLAIM's three-element reply (Redis server >= 7.0). Soft on an unreachable
# broker — after retrying it for REPLICATOR_REDIS_FLOOR_WAIT (default 30 s), which
# is what lets it see the broker on a cold boot (#88) — and fatal on a genuine
# downgrade. Not '-' prefixed — that exit 1 must stop the start.
ExecStartPre=/bin/bash /home/exedev/replicator/scripts/check_redis_floor.sh
# Refresh the cannobserv wheelhouse. Non-fatal ('-' prefix): a transient GCS
# failure is surfaced to the journal, and an already-populated wheelhouse still
# starts. Only a genuinely missing wheel surfaces as a hard uv resolution error.
#
# What it writes to journald is plain text, not the app's JSON — it runs
# --no-project, in an environment with none of the project's dependencies
# installed, so it cannot import build_json_formatter() (#15, skills#83).
ExecStartPre=-/usr/local/bin/uv run --no-project --with 'google-cloud-storage>=2,<4' python scripts/sync_wheelhouse.py
# --frozen --no-sync: run exactly the committed lockfile; dependency sync is a
# deploy step, not a service-start side effect.
ExecStart=/usr/local/bin/uv run --frozen --no-sync python -m src.worker.main
Restart=on-failure
RestartSec=5
# SIGTERM asks the worker to stop taking new work; it then finishes the message
# in flight and acks it, so a restart does not strand an entry in the PEL. The
# grace period must exceed the blocking read window (REPLICATOR_READ_BLOCK_MS,
# 5s by default) plus the handler's own budget, or systemd SIGKILLs mid-message
# and forces a stale-claim round-trip that a clean restart has no reason to need.
#
# The handler's budget stopped being the fetch driver's own 30s in #11: a command
# may now carry its own timeout_seconds, bounded by
# REPLICATOR_MAX_FETCH_TIMEOUT_SECONDS (120s by default). That ceiling and this
# value are one decision — tests/test_deploy.py enforces the pairing, so raising
# the ceiling without widening this window fails the suite rather than producing
# a deploy that SIGKILLs every slow fetch.
#
# The retention sweep is a third term: it runs in asyncio.to_thread, which cannot
# be cancelled, so shutdown waits out whatever walk is in flight. Fast at the
# volumes the ceiling permits (REPLICATOR_BLOB_MAX_TOTAL_BYTES), but revisit this
# alongside any change that lets the tree hold far more files.
#
# Storage is the fourth (#7). Under REPLICATOR_BLOB_BACKEND=gcs a store is a
# network round trip bounded by REPLICATOR_BLOB_TIMEOUT_SECONDS (30s), also in
# asyncio.to_thread and also beyond cancellation — so it is now summed by
# tests/test_deploy.py alongside the poll, the pacing sleep and the fetch: 160s
# against this 180s.
#
# That leaves 20s of margin rather than the 50s the sweep used to have, and the
# margin is honest because **the third and fourth terms are mutually exclusive**:
# the sweep runs only on the local backend and the storage round trip only on the
# object-store one, so no single shutdown pays both. The sum deliberately
# overcounts rather than encoding that in a test nobody would revisit.
#
# The fetch term is a real bound since #104: REPLICATOR_MAX_FETCH_SECONDS (120s)
# is one deadline around the whole fetch — the destination guard's resolve on
# every hop (#100), every connect, every read. Before it the sum carried the
# per-operation timeout and the resolve as two terms, each bounding a single
# step and neither a fetch, so a redirect chain or a trickling origin ran past
# both. 160s against 180s; raising that setting means widening this.
TimeoutStopSec=180
[Install]
WantedBy=multi-user.target