Skip to content

Latest commit

 

History

History
354 lines (276 loc) · 23.6 KB

File metadata and controls

354 lines (276 loc) · 23.6 KB

Deployment

Configuration — env files, every variable, and the unit-only credentials: ENVIRONMENT.md.

Systemd Service

A systemd unit file is provided at deploy/watcher.service.

Installation

Install the Archiver service first — see § Archiver Service below. Watcher no longer holds an Archiver SDK and will boot without one (#254), but info.registry is Archiver's stream on the broker, so until Archiver is producing, watched_items cannot reconcile and no registry state arrives.

Join the tailnet first, too (#280). Notifier runs on its own VM since notifier#43 and binds its tailnet address alone — its launch scripts resolve this host's 100.x address and hand uvicorn that one --host, so it is unreachable from loopback, from exe.dev's internal 10.42.0.0/16, and from the internet. So notifier in WATCHER_NOTIFIER_BASE_URL is a MagicDNS name on the cannobserv.org.github tailnet, and a host that has not joined cannot resolve or reach it at all.

This failure is quieter than the old wrong-port one. The #277/#278 gates check that the URL and the flag are held together, not that the host answers, so a watcher off the tailnet starts clean and then fails every dispatch at call time. Join the tailnet, then confirm curl http://notifier:9000/health answers before starting the unit. This node's identity, its peers, the cold-boot race and the ACL rules: reference/tailscale.md.

Set the host's memory posture too (#307, #309). This VM has no swap and shares memory with agent sessions the OOM killer cannot pick, so the kernel reserve and the slice reservations are part of the install, not tuning done later: HOST-MEMORY.md.

# Create system env directory
sudo mkdir -p /etc/watcher
# Add production secrets (at minimum DATABASE_URL)
echo 'DATABASE_URL=postgresql+asyncpg://watcher:watcher@localhost:5432/watcher' | sudo tee /etc/watcher/.env
# The dashboard's public base — THIS host's (#296 D6). Unset only WARNs, and every
# notification loses its link; a value copied from another VM's file passes every
# check and points readers at that VM.
echo 'WATCHER_PUBLIC_BASE_URL=https://co-watcher.exe.xyz' | sudo tee -a /etc/watcher/.env
sudo chmod 640 /etc/watcher/.env
sudo chown root:exedev /etc/watcher/.env

# The notifier credential goes in its OWN file, readable by root only (#278).
# It must never join /etc/watcher/.env: scripts/load-env.sh exports that file
# into every agent shell, and this is the pair that makes a stray dispatch
# deliverable to real subscribers.
sudo install -m 600 -o root -g root /dev/null /etc/watcher/notifier.env
sudo tee /etc/watcher/notifier.env >/dev/null <<'EOF'
WATCHER_NOTIFIER_BASE_URL=http://notifier:9000
WATCHER_NOTIFIER_API_KEY=<production tenant key from notifier's scripts/seed_tenant.py>
EOF

# IMPORTANT: /etc/watcher/.env and /etc/watcher/notifier.env must both exist
# before starting the service. Without either, systemd refuses to start the
# unit (both EnvironmentFile= lines are required).

# Copy (or symlink) the unit file
sudo cp deploy/watcher.service /etc/systemd/system/watcher.service

# Reload systemd, enable, and start
sudo systemctl daemon-reload
sudo systemctl enable watcher
sudo systemctl start watcher

# needrestart lists restarts, never performs them (#331): without it a libc6
# security update restarts watcher and PostgreSQL mid-apply. Takes effect at
# the next apt run; `sudo needrestart -m u -r l -b` prints "Disabling Ubuntu
# mode" once it is read, and restarts nothing. It does not reach maintainer
# scripts — postgresql-16's own upgrade still restarts its cluster. Measured on
# a scratch cluster (#338, 2026-09-30): requests ride through it — 5xx only
# while the cluster is down, none after, thanks to the checkout ping (#335) —
# but the worker stops and restarts (#340): *Managing the Service*.
sudo install -D -m 644 deploy/needrestart.conf.d/watcher.conf \
     /etc/needrestart/conf.d/watcher.conf

Managing the Service

# Restart after code changes
sudo systemctl restart watcher

# After any PostgreSQL cluster restart (apt or manual) the embedded worker
# stops (#338) and the supervisor restarts it with backoff (#340): an ERROR
# "procrastinate worker stopped unexpectedly" or "… died" (one per attempt
# while the cluster is down), then "Starting worker on all queues". The stop
# waits for running jobs (aborting any left at 30s, #334) and an unregister
# (7s, 30s measured 2026-09-30), and /ready's "queue" stays true meanwhile —
# trust the journal. No "Starting worker" minutes after the cluster is back →
# restart watcher.
sudo journalctl -u watcher --since -10min | grep -E 'supervisor|Starting worker'
curl -s localhost:8000/ready   # {"status":"ready","db":true,"queue":true}

# A job a dead worker left in `doing` (SIGKILL, OOM, a restart above) is retried
# by the 5-minute sweep once its worker is 120s silent (#334): WARNING "stalled
# job retried"; ERROR "stalled job failed" carries its `reason` (attempts_cap,
# abort_requested); ERROR "… not recovered" repeats per sweep until its
# conflict clears.
sudo journalctl -u watcher --since -1h | grep 'stalled job'

# Check status
sudo systemctl status watcher

# Follow logs
sudo journalctl -u watcher -f

# Follow WARNING+ only (app records are JSON: timestamp/level/logger/message — #238).
# `-R` + `fromjson?` is load-bearing: the ExecStartPre wheelhouse sync writes a
# plain-text line on every start (#247), and bare `jq` aborts on it with a parse
# error — killing the follow and hiding every record after it. Non-JSON lines are
# tagged PLAIN rather than dropped, so a failed sync still surfaces here.
# Needs jq (present on this VM at /bin/jq); the grep below is the no-jq fallback.
sudo journalctl -u watcher -f -o cat | jq -R 'fromjson? // {level: "PLAIN", message: .} | select(.level | IN("WARNING","ERROR","CRITICAL","PLAIN"))'
sudo journalctl -u watcher -f -o cat | grep -E '"level": "(WARNING|ERROR|CRITICAL)"|^error: '

# Reload after editing deploy/watcher.service
sudo systemctl daemon-reload && sudo systemctl restart watcher

Every line the application writes is JSON, uvicorn's included: the unit's ExecStart passes --log-config src/core/log_config.json, which routes uvicorn's own uvicorn/uvicorn.access/uvicorn.error loggers (propagate=False, plain-text handlers by default) through the app's build_json_formatter() (#244), and strips uvicorn's ANSI color_message duplicate from the payload (#246). Drop the flag and the jq filter above silently skips the access/boot lines, which come back as plain text. Same flag is baked into scripts/dev_server.sh.

The unit's ExecStartPre steps are the exception (#247): they run before the project is importable, so their output is plain text by design — the wheelhouse sync writes wheelhouse in sync: … on every start and error: could not sync gs://… when it fails, and the BUILD_ID stamp writes plain text on failure. Any pipeline that json.loads every MESSAGE must tolerate them; reading journald's own fields (_SYSTEMD_UNIT, SYSLOG_IDENTIFIER, MESSAGE) is unaffected. See CONVENTIONS.md → ExecStartPre output is plain text.

Host memory posture

Split out to HOST-MEMORY.md — the #307 reservation, the #309 slice drop-ins and dependency-chain table, the kernel reserve, why earlyoom is declined (#323), and the pinned SocratiCode install.

Database Migrations

Split out to MIGRATIONS.md — the manual alembic upgrade head step, the two-role grant model (#259), the autogenerate scratch-database rule, and the information-schema drop (#234).

Archiver Service

The Archiver is a sibling service on its own VM (co-registrar, archiver.service on port 8000 there, per archiver's own AGENTS.md) that owns the canonical InfoItem / InfoSource / SourceRevision / RepSpec registry.

Watcher makes no HTTP calls to it. The SDK, its API key, and the lifespan pre-warm were removed in #254 together with the last outbound call (get_info_item on WatchedItem create); registry state now arrives as info.registry announcements, which Watcher reconciles into watched_items. watcher.service therefore boots regardless of Archiver's state — but with the broker down or the producer not running, the registry simply never converges, and the log line to look for is the info.registry consumer failing to start.

See the Archiver repo's docs/DEPLOYMENT.md for the full Archiver install on its VM (key generation, env-var registration, systemd unit). After installing archiver.service there, restart watcher.service here:

sudo systemctl restart watcher

Archiver Sync

Every detected change enqueues a pending_archiver_sync row, and a full fetch that renews the blob reference behind the latest revision upserts one (#293); the drain_pending_archiver_sync periodic task (src/workers/source_revisions_drain.py) publishes each as source_revision_observed on content.revisions, on a fixed 1-minute cadence — a hardcoded Procrastinate cron, not an env var. Nothing is POSTed to Archiver any more (#253); Archiver consumes the stream and decides what to persist. A broker outage self-heals within a minute of Redis returning, and the outbox holds the backlog meanwhile.

With WATCHER_BUS_REDIS_URL unset the drain skips loudly and rows accumulate rather than draining — the same "no bus, no publish" posture as the fetch-policy producer, and the reason an unset broker URL shows up as a growing backlog.

The drain runs on the embedded worker inside the single uvicorn process; there is no separate unit to start or monitor.

There is no scratch cache any more (#253). Watcher used to write its own copy of the extracted bytes, report that path as content_cache_uri, sweep it, and PATCH null — three moving parts doing nothing Replicator's blob_uri does. The copy, the sweeper, and the WATCHER_CACHE_* variables are gone; a stuck outbox now shows up only as backlog rows, never as disk growth.

A dead-lettered row is the one thing the backlog query can miss. A row whose payload cannot be built (a missing wire-required field) is stamped dead_lettered_at and stops being selected — deliberately, so it cannot spin forever. It will sit in the table indefinitely, so count it separately rather than reading a flat backlog number as "the drain is fine".

To spot one:

source scripts/load-env.sh
# Backlog size, age of the oldest undrained row, and the last failure reason.
# psql needs a driverless URL — strip the SQLAlchemy "+asyncpg" dialect suffix.
psql "${DATABASE_URL/+asyncpg/}" -c \
  "SELECT count(*), min(created_at) AS oldest, max(attempts) AS max_attempts FROM pending_archiver_sync;"
psql "${DATABASE_URL/+asyncpg/}" -c \
  "SELECT id, attempts, last_error FROM pending_archiver_sync ORDER BY created_at LIMIT 5;"
# Rows that will never drain on their own — these need an operator, not patience.
psql "${DATABASE_URL/+asyncpg/}" -c \
  "SELECT id, dead_lettered_at, last_error FROM pending_archiver_sync
   WHERE dead_lettered_at IS NOT NULL ORDER BY dead_lettered_at DESC LIMIT 10;"

A non-empty backlog with an oldest row older than a few minutes means the drain is failing, not merely busy — check journalctl -u watcher -f for drain attempt failed and confirm Archiver is reachable.

BUILD_ID

The systemd unit automatically sets BUILD_ID to the current git short SHA before each start via ExecStartPre. This value is used for:

  • Static asset cache-busting (?v=<sha> query params)
  • Page footer version display
  • /health endpoint build field

Manual Override

To pin a specific build ID, set it in .env:

echo 'BUILD_ID=abc1234' >> .env

.env is loaded after /etc/watcher/.env and /run/watcher/build-id (both use the - prefix to be optional), so values in .env take precedence.

Fallback

When BUILD_ID is not set (e.g., running locally without the systemd unit), the application defaults to "dev".

Disk Cleanup Timer

A weekly cleanup timer removes stale caches and rotates journal logs to keep the 25 GB VM disk healthy.

Files:

  • deploy/watcher-cleanup.service — oneshot service, runs as exedev
  • deploy/watcher-cleanup.timer — fires Sun 03:00 UTC ± 30 min
  • deploy/watcher-cleanup.sudoers — two targeted NOPASSWD rules (apt-get clean, journalctl --vacuum-time=14d)
  • scripts/cleanup.sh — the cleanup script; logs to /var/log/watcher/cleanup-<timestamp>.log (keeps 10)

Installing the timer

# Sudoers rules
sudo cp deploy/watcher-cleanup.sudoers /etc/sudoers.d/watcher-cleanup
sudo chmod 440 /etc/sudoers.d/watcher-cleanup
sudo chown root:root /etc/sudoers.d/watcher-cleanup
sudo visudo -c  # verify no syntax errors

# Make the script executable
chmod +x scripts/cleanup.sh

# Install and enable the systemd units
sudo cp deploy/watcher-cleanup.service /etc/systemd/system/watcher-cleanup.service
sudo cp deploy/watcher-cleanup.timer   /etc/systemd/system/watcher-cleanup.timer
sudo systemctl daemon-reload
sudo systemctl enable --now watcher-cleanup.timer

Managing the Timer

# Confirm it is scheduled
systemctl list-timers watcher-cleanup.timer

# Manual run (for testing)
sudo systemctl start watcher-cleanup.service

# Follow output
sudo journalctl -u watcher-cleanup.service -f

# View logs
ls -lt /var/log/watcher/cleanup-*.log
cat "$(ls -t /var/log/watcher/cleanup-*.log | head -1)"

# Reload after editing the timer
sudo systemctl daemon-reload && sudo systemctl restart watcher-cleanup.timer

What it cleans

Target Action
VS Code server installs >30 days old rm -rf
~/.npm/_npx, ~/.npm/_cacache rm -rf only above 2 GB — below that the cache is kept, because the plugin's MCP server cold-installs through it at session start (#314)
uv build cache uv cache prune
APT package cache apt-get clean
Journal logs >14 days journalctl --vacuum-time=14d
Playwright cache audit only — logs size, warns if >2 GB, never deletes

Job history — not this timer

Procrastinate's finished jobs are pruned in-process, by the hourly prune_job_history periodic task (src/workers/retention.py, #296 D7): succeeded jobs after 7 days; failed, cancelled and aborted after 30. Unfinished jobs are never touched, events cascade with their job, and watcher_app already holds the DELETE it needs. Before #296 nothing pruned them, and the two tables were 96 % of the database — 679 k jobs back to March.

A backlog that size is pruned once from a shell, not by the task's first run. Every delete_old_jobs call sorts every event in the table (its age filter sits outside the sort), and the task runs inside the service's worker, one job at a time — so the first run over 679 k jobs would have been one long transaction there. The script steps the horizon down a week at a time, from one step below the oldest job, one statement per slice, with progress: each step repeats the sort, but bounds the rows each transaction deletes. It deletes irreversibly, so take a dump first:

source scripts/load-env.sh
WATCHER_ALLOW_PRODUCTION_DB=1 uv run python -m src.ops.prune_job_history --dry-run
WATCHER_ALLOW_PRODUCTION_DB=1 uv run python -m src.ops.prune_job_history

If a backlog ever comes back — a restore of the pre-prune safety dump, say — prune it this way before the service starts. Procrastinate defers a missed periodic tick up to ten minutes late at boot, so a start soon after :23 runs the task at once: the #296 restart at 23:26 did exactly that, harmlessly, because the backlog was already gone.

The disk does not shrink. Deleted rows free space for Postgres to reuse, not back to the filesystem: the two tables' files keep their size until a dump-and-restore (the #296 move does one) or a VACUUM FULL. A pg_dump is compact either way — 1.7 MB after the prune.

Database Backup Timer

watcher-backup.timer dumps the database to gs://co-gcs-watcher-backup nightly (03:17 UTC, Persistent=true) through watcher-backup.service — create-only, holding no database credential, run as its own dynamic user with no capabilities (#297), and checking in to a co-status dead-man monitor on every run (#296 D8/D9, #330). Provisioning (the role script and the three files), install, why its sandbox has this shape, and the restore runbook with its go/no-go gates: RECOVERY.md. Enable the timer only after a hand-started run has succeeded.

Cannobserv wheelhouse

Cannobserv wheelhouse (#220). co-core + co-core-aio (the shared cannabis-observer substrate) resolve from a local wheelhouse mirrored from the private GCS index gs://co-gcs-pypi, via [tool.uv] find-links = ["./.wheelhouse"] — not git sources. Populate it before any uv command (find-links makes every uv invocation require the dir; .wheelhouse/.gitkeep is tracked so a fresh clone has it):

Auth is ADC: on the VM/deploy the co-pypi-reader SA key at GOOGLE_APPLICATION_CREDENTIALS (in /etc/watcher/.env); in CI, keyless via Workload Identity Federation (.github/workflows/ci.yml). The identity needs only roles/storage.objectViewer. Reproducibility is uv.lock (pins the exact version), not wheelhouse contents. Upgrade: re-sync, then uv lock --upgrade-package co-core (bump the floor if the minor moved). Currently pinned: v0.19.7, exact through #325's shadow (lockstep with CannObserv/processor; floors were >=0.19.6,<0.20 — 0.19.6 carries SourceRevisionObservedEvent.blob_fingerprint (cannobserv#493), the raw-bytes digest #329 echoes onto content.revisions, and its Emit twin's bare-hex validator; 0.19.7 is docs plus the destination_refused failure token, which watcher journals generically because it enumerates no reasons. Diffed wheel-to-wheel against 0.19.4, every module watcher imports that moved moved additively: changes.py, envelope.py and streams.py gain the persist pair (ContentPersistCommand / BlobPersistedEvent / PersistFailedEvent, content.persist, their envelope keys) and blobstore.py gains FingerprintMismatch and the signed-URL bounds; watcher publishes and consumes none of them. Previously v0.19.4 — floors >=0.19.4,<0.20: 0.19.4 carries cannobserv#486: canonical_text / canonical_text_fingerprint / processor_version / spec_schema_version, the lifted media-type dispatch that src/core/media_type.py now re-exports, and the content.process / content.derived contract that #324–#326 adopt. The 0.13 → 0.19 bump (#324) was diffed wheel-to-wheel on every module watcher imports: changes.py, streams.py, envelope.py and extract/__init__.py change additively, html.py is #457's default_factory idiom only, base.py / pdf.py / csv_excel.py / bus/exceptions.py are byte-identical, and co_core_aio/bus.py adds claim_stale_page and widens dead_letter, neither of which watcher's consumers call differently; every Breaking entry 0.14.0–0.19.4 is on co/v1, transcript or ext/ surfaces this repo does not import. Previously v0.13.2 — floors >=0.13.1,<0.14: 0.13.1 adds group_name, the derived consumer-group naming the #285 rename adopts, and 0.13.2 changes only co_v1 adapters; 0.10.0 widens WatchStatusEmit.health with "unknown" (#328), the value the #264 status producer emits for a pre-first-check item; 0.9.3 carried RegistryAnnouncementState / WatchStatusState and the INFO_REGISTRY / INFO_WATCH_STATUS streams the registry consumer reads, #324 and #254; 0.8.1 already carried the spec_fingerprint derivation the revisions producer imports unconditionally, #309; co-core-aio carries the bus extra for the fetch-policy producer and the tail reader — #245). The 0.10 → 0.13 bump (#285) is additive on every surface Watcher touches, established by diffing the wheels rather than by reading release notes: across co_core / co_core_aio 0.10.0 → 0.13.2 the only module Watcher imports that changed at all is pure/adapters/bus/streams.py, and it changed additively (StreamKind, stream_kind, group_name; every stream constant byte-identical). pure/models/changes.py, pure/adapters/bus/envelope.py, pure/adapters/bus/exceptions.py, pure/extract/*, pure/util/hashing.py and all of co_core_aio/bus.py are byte-identical; everything else that moved is pure/adapters/co_v1/* and pure/models/task_performer.py, which this repo does not import. Re-run that diff on the next multi-minor bump — a floor spanning three MINORs is not covered by any one release's notes. The 0.9 bump is additive on every surface Watcher touches — the content.fetch / content.blobs / content.revisions paths are untouched, and the ChangeEventPayload union simply widens; the MINOR is driven by WordPressConfig / relationship.py / cdio changes this repo has no exposure to. The 0.8 bump (#300, adopted in #252) is breaking in both directions: info_source_id is required on all three content contracts and BlobAvailableEvent.command_id stopped being optional, so a 0.7.7 fact no longer decodes here — the deploy-ordering rule that follows is in docs/CONTENT-PIPELINE.md → "info_source_id on the wire". co-core carries the extract extra (co-core[extract]) — the heavy HTML/PDF/CSV parsers behind the extractors constructed in src/core/registry.py. The content-acquisition pipeline (fetch → extract → fingerprint) is now co-core's (co_core.pure.extract.*, co_core.effects.fetch, co_core_aio.fetch), adopted in #236; watcher no longer fetches at all (#241 step 5) — src/core/fetch.py is gone; the watcher/0.1.0 User-Agent now lives in src/core/fetch_commands.py beside its only consumer and rides out on every content.fetch command's headers to preserve fingerprint byte-continuity. The systemd unit refreshes the wheelhouse via a non-fatal ExecStartPre so restarts self-heal (its output is plain text, not the app's JSON — see Logging).

notifier-client

notifier-client (#284). The notifier SDK is a pinned git-tag source in [tool.uv.sources] — https://github.com/CannObserv/notifier.git, subdirectory = "clients/python", tag = "vX.Y.Z". The repo is public, so it needs no credential: no SSH key, no host alias, no CI URL rewrite (the ssh://git@github-notifier/ alias it replaced existed only on the retired shared VM). The tag is the pin — uv enforces no version floor against a git source, so "notifier-client" stays bare in [project.dependencies]; a floor there would document intent and enforce nothing. Upgrade: change the tag, uv sync (refreshes uv.lock, freezing the commit), read notifier's clients/python/CHANGELOG.md. A Removed dependency can no longer take a module this repo imports: tests/test_dependency_imports.py requires every third-party import declared here, bar an allowlist that names each exception's provider — croniter, via procrastinate (#294). To stay put, do nothing. tests/test_dependency_sources.py fails a git source that is not HTTPS (the scheme only — it cannot tell a public repo from a private one) or not pinned by tag, and any workflow under .github/workflows/ that rewrites a git URL (insteadOf in any capitalisation — git config keys are case-insensitive). Currently pinned: v0.3.1 — public API unchanged from 0.2.1; it dropped python-dateutil from the SDK's dependencies, which nothing here imports (it still resolves via procrastinate and co-core's dateparser). Release procedure and the trigger for graduating to a published package: notifier's docs/RELEASING.md.