# Install dependencies (creates .venv automatically)
uv syncTwo env files, loaded in order:
# Production secrets (DATABASE_URL) — persistent, survives repo resets
/etc/watcher/.env
# Dev/agent secrets (GH_TOKEN, TEST_DATABASE_URL) — repo root, git-ignored
.env
# Load both for shell commands
source scripts/load-env.shThe systemd service loads only /etc/watcher/.env, plus its unit-only
notifier.env — never the repo .env, whose agent tokens are no production
configuration (#296 D5; deploy/watcher.service, ENVIRONMENT.md).
A production setting put in the repo .env is one the service never sees.
scripts/load-env.sh is sourced, not executed — the exports have to land in your
shell. It parses each file rather than sourcing it, so a secrets file is never run, and
skips a malformed line instead of aborting. It replaced
export $(cat … | xargs), which printed the whole environment (secrets included) when
both files were absent, died under set -e on any comment line, and word-split values
containing spaces. Paths are overridable via WATCHER_SYSTEM_ENV_FILE /
WATCHER_PROJECT_ENV_FILE; guarded by tests/scripts/test_load_env.py.
# Full ship gate: ruff check, ruff format --check, pytest (non-integration)
bash scripts/pre-ship.shThis is watcher's thin wrapper — it loads the env files above, then delegates to the
vendored gate in shipping-work-python-fastapi. Run it from the repo root; it exits
non-zero on any failure and 2 on tooling/infra problems (including an uninitialized
skills-vendor/ submodule). See SKILLS.md for why the gate itself is not
forked.
The watcher service runs via systemd. Always use systemctl — never start uvicorn manually on port 8000.
# Restart after code changes (migrations are NOT auto-run)
sudo systemctl restart watcher
# Check status
sudo systemctl status watcher
# Follow logs
sudo journalctl -u watcher -f
# Reload systemd after editing deploy/watcher.service
sudo systemctl daemon-reload && sudo systemctl restart watcher# Dev server (port 8001) — the ONLY sanctioned launch path (#233).
# Targets TEST_DATABASE_URL (or WATCHER_DEV_DATABASE_URL), migrates it,
# applies the procrastinate schema (#341), and refuses any DB whose name lacks
# a _test/_dev suffix. Never hand-run uvicorn with the prod env loaded:
# /etc/watcher/.env points DATABASE_URL at production, and the embedded worker
# would consume the prod task queue.
bash scripts/dev_server.sh
# Knobs: WATCHER_DEV_PORT (default 8001; 8000 refused — it belongs to
# systemd), WATCHER_DEV_DATABASE_URL, WATCHER_DEV_SKIP_MIGRATE=1,
# WATCHER_DEV_BUS_REDIS_URL, WATCHER_DEV_NOTIFIER_BASE_URL +
# WATCHER_DEV_NOTIFIER_API_KEY (both or neither)Never launch uvicorn by hand with the prod env loaded — /etc/watcher/.env
points DATABASE_URL at production, and a hand-run "dev" server would share
the prod DB, run a second Procrastinate worker on the prod queue, and split
the rate-limiter budget (#233). The script targets TEST_DATABASE_URL (or
WATCHER_DEV_DATABASE_URL), migrates it, and refuses anything whose DB name
lacks a _test/_dev suffix. The same rule is enforced in-app by
src/core/db_safety.py; only deploy/watcher.service opts into prod via
WATCHER_ALLOW_PRODUCTION_DB=1 (in the unit, never an env file).
No migration creates procrastinate's tables, so the script applies them too
(#341): always after the TEST_DATABASE_URL branch's public-schema reset, and
on a persistent WATCHER_DEV_DATABASE_URL only when procrastinate_jobs is
missing — schema --apply is not idempotent. Without them the embedded
worker's register_worker fails at boot.
The database is not the only production resource an env file hands out. The
script clears an inherited WATCHER_BUS_REDIS_URL (#262) and an inherited
WATCHER_NOTIFIER_BASE_URL/WATCHER_NOTIFIER_API_KEY (#277) unless a scratch replacement is
named, because a dev server runs the embedded worker against a real check
pipeline — so an inherited notifier key delivers real notifications to real
subscribers as the production tenant, and succeeds, leaving no error behind.
Each resource has a unit-only opt-in flag; a URL held without its flag aborts
startup rather than going quiet. See
ENVIRONMENT.md → Environment Variables.
Both sanctioned launch paths (scripts/dev_server.sh and the systemd
ExecStart) pass --log-config src/core/log_config.json so uvicorn's own
access/error lines are JSON like the app's (#244). Any ad-hoc uvicorn command
needs the same flag, or it emits plain text alongside the JSON records.
After committing to main, restart the service to pick up changes:
sudo systemctl restart watcher
# Verify
curl -s http://localhost:8000/health | python3 -m json.toolRun a worktree build on a different port to avoid conflicting with the service:
cd .worktrees/<branch>
bash scripts/dev_server.sh # same guard rails as the repo-root dev server# Run all tests (excludes integration)
uv run pytest
# Run with coverage
uv run pytest --cov
# Run a specific file
uv run pytest tests/path/to/test_file.py --no-cov
# Run integration tests (hits live external services)
uv run pytest -m integration# Check
uv run ruff check .
# Fix auto-fixable issues
uv run ruff check --fix .
# Format (CI runs the --check form)
uv run ruff format --check .
uv run ruff format .ruff format also formats the python fences in Markdown (ruff ≥0.16, #315):
every Markdown file is in scope except docs/plans/ and skills/ — the
reasons sit beside extend-exclude in pyproject.toml.
# PostgreSQL setup (first time)
sudo apt-get install -y postgresql postgresql-client
sudo systemctl start postgresql
sudo -u postgres psql -c "CREATE USER watcher WITH PASSWORD 'watcher';"
sudo -u postgres psql -c "CREATE DATABASE watcher OWNER watcher;"
sudo -u postgres psql -c "CREATE DATABASE watcher_test OWNER watcher;"
# Then split the roles: scripts/setup-db-roles.sql (#259) — see
# docs/MIGRATIONS.md -> "Migration role and application role". Grants are not
# schema state, so no migration recreates them on a fresh host.
# Watcher migrations. Alembic connects with WATCHER_MIGRATION_DATABASE_URL when
# set, else DATABASE_URL (#259). alembic.ini carries no URL, so a shell that
# skipped load-env.sh fails instead of reaching for a default.
source scripts/load-env.sh
uv run alembic upgrade head
uv run alembic current
# Archiver service migrations live in the sibling Archiver repo.
# See its docs/COMMANDS.md.Autogenerate wants a scratch database, not production. --autogenerate
diffs the models against whatever database it connects to, so pointing it at
DATABASE_URL means diffing production. Build a throwaway instead — which is
what CREATEDB on the migration role (#259) is for:
source scripts/load-env.sh
MIG="${WATCHER_MIGRATION_DATABASE_URL:-$DATABASE_URL}" # the owning role
SCRATCH="${MIG%/*}/watcher_autogen_dev" # same credentials, new database
psql "${MIG/+asyncpg/}" -c "CREATE DATABASE watcher_autogen_dev"
WATCHER_MIGRATION_DATABASE_URL="$SCRATCH" uv run alembic upgrade head
WATCHER_MIGRATION_DATABASE_URL="$SCRATCH" uv run alembic revision \
--autogenerate -m "description of change"
WATCHER_MIGRATION_DATABASE_URL="$SCRATCH" uv run alembic check # drift, locally
psql "${MIG/+asyncpg/}" -c "DROP DATABASE watcher_autogen_dev"watcher_test is not a substitute: pytest builds it with
Base.metadata.create_all, so its alembic_version never matches its tables
and upgrade head fails part-way through the chain. Alembic warns on stderr
when the two URLs name different databases — that line is expected here, and is
the signal to take seriously anywhere else.
# Apply procrastinate schema (first time, after DB setup; dev_server.sh does
# this itself for the dev database, #341)
source scripts/load-env.sh
uv run procrastinate --app=src.workers.app schema --apply
# Run worker standalone (alternative to embedded mode in FastAPI)
uv run procrastinate --app=src.workers.app worker
# The worker also runs embedded in FastAPI via lifespan — no separate process needed for devOperator entry points in src/ops/, each with its runbook:
# The backup's read-only database role — once per cluster, before its first run
# (docs/RECOVERY.md → Install and first run). Idempotent; its report is the check:
sudo -u postgres psql -d watcher < scripts/setup-backup-role.sql
# Nightly dump to GCS — the timer runs it; one run by hand (docs/RECOVERY.md):
sudo systemctl start watcher-backup.service
# Restore — always name the host that shipped the dump (docs/RECOVERY.md → Restore).
# This host ships under `co-watcher/`; `watcher/` is the retired VM's timeline, whose
# newest object post-dates the #296 cutover — `--latest --prefix watcher` is the wrong DB:
sudo bash -c "set -a; . /etc/watcher/backup.env; set +a; export GOOGLE_APPLICATION_CREDENTIALS=/etc/watcher/co-watcher-backup.json; .venv/bin/python -m src.ops.restore --list"
# One-off job-history backlog prune; the hourly task holds it after
# (docs/DEPLOYMENT.md → Job history). The opt-in covers the dry run too:
source scripts/load-env.sh
WATCHER_ALLOW_PRODUCTION_DB=1 uv run python -m src.ops.prune_job_history --dry-run# VM setup (tailwindcss CLI — global npm). Already installed on co-watcher.
# Keep the pin: output.css carries its builder's version in the banner, and a
# newer CLI rebuilds it differently, which check-css.sh reads as stale.
sudo npm install -g @tailwindcss/cli@4.2.4
# Build Tailwind CSS
bash scripts/build-css.sh
# Watch mode (auto-rebuild on changes)
bash scripts/build-css.sh --watch# Init after cloning
git submodule update --init --recursive
# Force-refresh vendor skills
git submodule update --remote --merge skills-vendor/gregoryfoster-skills skills-vendor/obra-superpowerstests/conftest.py builds watcher's tables in TEST_DATABASE_URL with
Base.metadata.create_all and drops them at session end; per-test isolation is
db_session's savepoint rollback. Tests that bypass db_session and write
directly via engine.connect() leak rows into subsequent sessions — don't do
that.
No sibling checkout (#311). The suite used to subprocess-run the Archiver
checkout's alembic, located by ARCHIVER_REPO_PATH, to build an information
schema its factories wrote rows into. Those rows only minted two ULIDs:
production has had no information schema since #271 (see docs/MIGRATIONS.md
→ "Drop the dead information schema"), and the WatchedItem links carry no FK.
make_watched_item now mints both links, test_engine drops a leftover
information schema from a pre-#311 session, and
tests/test_archiver_isolation.py fails if the test database carries one
again. No migration may create or drop it either — tests/test_migration_chain.py.
pytest-xdist is unsupported. The fixture writes to a single
TEST_DATABASE_URL; multiple xdist workers would race on creating and
dropping the watcher tables. Tracked in #150 — worker-id-suffixed databases
are the rework when xdist is actually adopted.
GitHub Actions (.github/workflows/ci.yml) runs on push/PR to main: a
lint job (ruff check + ruff format --check), a test job
(pytest -m "not integration" against a postgres:16 service), and a
migrations job (independent migration-chain smoke-check, #234 — alembic upgrade head from an empty postgres:16 then alembic check for drift). No
job checks out a sibling repo — #254 removed the archiver-client path dep
that lint and migrations needed, and #311 the test job's archiver alembic
run. All three jobs authenticate to GCS keyless via WIF
(vars.GCP_WIF_PROVIDER → co-pypi-reader SA) and sync the wheelhouse before
uv sync; notifier-client installs as written, from
its public HTTPS tag source — no URL rewrite (#284). The migrations job needs
nothing from archiver either — the #234 squash collapsed the pre-existing
chain into a self-contained genesis baseline (2addddea0b03) that references
no information schema, so upgrade head from empty is fully standalone (no
archiver seeding, no cross-service ordering).
Squash cutover: already-migrated DBs need a one-time alembic stamp 2addddea0b03 --purge before their next upgrade — see docs/MIGRATIONS.md
→ "Migration baseline (squash)". Integration tests hit live external services
and are excluded in CI. One-time GCP grant (operator, for WIF) — bind watcher's repo
to the read-only SA; the org-scoped github-ci provider needs no change:
gcloud iam service-accounts add-iam-policy-binding \
co-pypi-reader@co-gcs.iam.gserviceaccount.com --project=co-gcs \
--role=roles/iam.workloadIdentityUser \
--member="principalSet://iam.googleapis.com/projects/912903030445/locations/global/workloadIdentityPools/github/attribute.repository/CannObserv/watcher"