Skip to content

Latest commit

 

History

History
329 lines (257 loc) · 13.2 KB

File metadata and controls

329 lines (257 loc) · 13.2 KB

Common Commands

Setup

# Install dependencies (creates .venv automatically)
uv sync

Environment

Two env files, loaded in order:

# Production secrets (DATABASE_URL) — persistent, survives repo resets
/etc/watcher/.env

# Dev/agent secrets (GH_TOKEN, TEST_DATABASE_URL) — repo root, git-ignored
.env

# Load both for shell commands
source scripts/load-env.sh

The systemd service loads only /etc/watcher/.env, plus its unit-only notifier.env — never the repo .env, whose agent tokens are no production configuration (#296 D5; deploy/watcher.service, ENVIRONMENT.md). A production setting put in the repo .env is one the service never sees.

scripts/load-env.sh is sourced, not executed — the exports have to land in your shell. It parses each file rather than sourcing it, so a secrets file is never run, and skips a malformed line instead of aborting. It replaced export $(cat … | xargs), which printed the whole environment (secrets included) when both files were absent, died under set -e on any comment line, and word-split values containing spaces. Paths are overridable via WATCHER_SYSTEM_ENV_FILE / WATCHER_PROJECT_ENV_FILE; guarded by tests/scripts/test_load_env.py.

Shipping

# Full ship gate: ruff check, ruff format --check, pytest (non-integration)
bash scripts/pre-ship.sh

This is watcher's thin wrapper — it loads the env files above, then delegates to the vendored gate in shipping-work-python-fastapi. Run it from the repo root; it exits non-zero on any failure and 2 on tooling/infra problems (including an uninitialized skills-vendor/ submodule). See SKILLS.md for why the gate itself is not forked.

Service Management

The watcher service runs via systemd. Always use systemctl — never start uvicorn manually on port 8000.

# Restart after code changes (migrations are NOT auto-run)
sudo systemctl restart watcher

# Check status
sudo systemctl status watcher

# Follow logs
sudo journalctl -u watcher -f

# Reload systemd after editing deploy/watcher.service
sudo systemctl daemon-reload && sudo systemctl restart watcher

Development

# Dev server (port 8001) — the ONLY sanctioned launch path (#233).
# Targets TEST_DATABASE_URL (or WATCHER_DEV_DATABASE_URL), migrates it,
# applies the procrastinate schema (#341), and refuses any DB whose name lacks
# a _test/_dev suffix. Never hand-run uvicorn with the prod env loaded:
# /etc/watcher/.env points DATABASE_URL at production, and the embedded worker
# would consume the prod task queue.
bash scripts/dev_server.sh

# Knobs: WATCHER_DEV_PORT (default 8001; 8000 refused — it belongs to
# systemd), WATCHER_DEV_DATABASE_URL, WATCHER_DEV_SKIP_MIGRATE=1,
# WATCHER_DEV_BUS_REDIS_URL, WATCHER_DEV_NOTIFIER_BASE_URL +
# WATCHER_DEV_NOTIFIER_API_KEY (both or neither)

Never launch uvicorn by hand with the prod env loaded — /etc/watcher/.env points DATABASE_URL at production, and a hand-run "dev" server would share the prod DB, run a second Procrastinate worker on the prod queue, and split the rate-limiter budget (#233). The script targets TEST_DATABASE_URL (or WATCHER_DEV_DATABASE_URL), migrates it, and refuses anything whose DB name lacks a _test/_dev suffix. The same rule is enforced in-app by src/core/db_safety.py; only deploy/watcher.service opts into prod via WATCHER_ALLOW_PRODUCTION_DB=1 (in the unit, never an env file).

No migration creates procrastinate's tables, so the script applies them too (#341): always after the TEST_DATABASE_URL branch's public-schema reset, and on a persistent WATCHER_DEV_DATABASE_URL only when procrastinate_jobs is missing — schema --apply is not idempotent. Without them the embedded worker's register_worker fails at boot.

The database is not the only production resource an env file hands out. The script clears an inherited WATCHER_BUS_REDIS_URL (#262) and an inherited WATCHER_NOTIFIER_BASE_URL/WATCHER_NOTIFIER_API_KEY (#277) unless a scratch replacement is named, because a dev server runs the embedded worker against a real check pipeline — so an inherited notifier key delivers real notifications to real subscribers as the production tenant, and succeeds, leaving no error behind. Each resource has a unit-only opt-in flag; a URL held without its flag aborts startup rather than going quiet. See ENVIRONMENT.md → Environment Variables.

Both sanctioned launch paths (scripts/dev_server.sh and the systemd ExecStart) pass --log-config src/core/log_config.json so uvicorn's own access/error lines are JSON like the app's (#244). Any ad-hoc uvicorn command needs the same flag, or it emits plain text alongside the JSON records.

Testing code changes against the live site

After committing to main, restart the service to pick up changes:

sudo systemctl restart watcher
# Verify
curl -s http://localhost:8000/health | python3 -m json.tool

Worktree testing

Run a worktree build on a different port to avoid conflicting with the service:

cd .worktrees/<branch>
bash scripts/dev_server.sh   # same guard rails as the repo-root dev server

Testing

# Run all tests (excludes integration)
uv run pytest

# Run with coverage
uv run pytest --cov

# Run a specific file
uv run pytest tests/path/to/test_file.py --no-cov

# Run integration tests (hits live external services)
uv run pytest -m integration

Linting

# Check
uv run ruff check .

# Fix auto-fixable issues
uv run ruff check --fix .

# Format (CI runs the --check form)
uv run ruff format --check .
uv run ruff format .

ruff format also formats the python fences in Markdown (ruff ≥0.16, #315): every Markdown file is in scope except docs/plans/ and skills/ — the reasons sit beside extend-exclude in pyproject.toml.

Database

# PostgreSQL setup (first time)
sudo apt-get install -y postgresql postgresql-client
sudo systemctl start postgresql
sudo -u postgres psql -c "CREATE USER watcher WITH PASSWORD 'watcher';"
sudo -u postgres psql -c "CREATE DATABASE watcher OWNER watcher;"
sudo -u postgres psql -c "CREATE DATABASE watcher_test OWNER watcher;"
# Then split the roles: scripts/setup-db-roles.sql (#259) — see
# docs/MIGRATIONS.md -> "Migration role and application role". Grants are not
# schema state, so no migration recreates them on a fresh host.

# Watcher migrations. Alembic connects with WATCHER_MIGRATION_DATABASE_URL when
# set, else DATABASE_URL (#259). alembic.ini carries no URL, so a shell that
# skipped load-env.sh fails instead of reaching for a default.
source scripts/load-env.sh
uv run alembic upgrade head
uv run alembic current

# Archiver service migrations live in the sibling Archiver repo.
# See its docs/COMMANDS.md.

Autogenerate wants a scratch database, not production. --autogenerate diffs the models against whatever database it connects to, so pointing it at DATABASE_URL means diffing production. Build a throwaway instead — which is what CREATEDB on the migration role (#259) is for:

source scripts/load-env.sh
MIG="${WATCHER_MIGRATION_DATABASE_URL:-$DATABASE_URL}"   # the owning role
SCRATCH="${MIG%/*}/watcher_autogen_dev"                  # same credentials, new database

psql "${MIG/+asyncpg/}" -c "CREATE DATABASE watcher_autogen_dev"
WATCHER_MIGRATION_DATABASE_URL="$SCRATCH" uv run alembic upgrade head
WATCHER_MIGRATION_DATABASE_URL="$SCRATCH" uv run alembic revision \
  --autogenerate -m "description of change"
WATCHER_MIGRATION_DATABASE_URL="$SCRATCH" uv run alembic check   # drift, locally
psql "${MIG/+asyncpg/}" -c "DROP DATABASE watcher_autogen_dev"

watcher_test is not a substitute: pytest builds it with Base.metadata.create_all, so its alembic_version never matches its tables and upgrade head fails part-way through the chain. Alembic warns on stderr when the two URLs name different databases — that line is expected here, and is the signal to take seriously anywhere else.

Task Queue (Procrastinate)

# Apply procrastinate schema (first time, after DB setup; dev_server.sh does
# this itself for the dev database, #341)
source scripts/load-env.sh
uv run procrastinate --app=src.workers.app schema --apply

# Run worker standalone (alternative to embedded mode in FastAPI)
uv run procrastinate --app=src.workers.app worker

# The worker also runs embedded in FastAPI via lifespan — no separate process needed for dev

Backup, restore and job history (#296)

Operator entry points in src/ops/, each with its runbook:

# The backup's read-only database role — once per cluster, before its first run
# (docs/RECOVERY.md → Install and first run). Idempotent; its report is the check:
sudo -u postgres psql -d watcher < scripts/setup-backup-role.sql

# Nightly dump to GCS — the timer runs it; one run by hand (docs/RECOVERY.md):
sudo systemctl start watcher-backup.service

# Restore — always name the host that shipped the dump (docs/RECOVERY.md → Restore).
# This host ships under `co-watcher/`; `watcher/` is the retired VM's timeline, whose
# newest object post-dates the #296 cutover — `--latest --prefix watcher` is the wrong DB:
sudo bash -c "set -a; . /etc/watcher/backup.env; set +a; export GOOGLE_APPLICATION_CREDENTIALS=/etc/watcher/co-watcher-backup.json; .venv/bin/python -m src.ops.restore --list"

# One-off job-history backlog prune; the hourly task holds it after
# (docs/DEPLOYMENT.md → Job history). The opt-in covers the dry run too:
source scripts/load-env.sh
WATCHER_ALLOW_PRODUCTION_DB=1 uv run python -m src.ops.prune_job_history --dry-run

Tailwind CSS

# VM setup (tailwindcss CLI — global npm). Already installed on co-watcher.
# Keep the pin: output.css carries its builder's version in the banner, and a
# newer CLI rebuilds it differently, which check-css.sh reads as stale.
sudo npm install -g @tailwindcss/cli@4.2.4

# Build Tailwind CSS
bash scripts/build-css.sh

# Watch mode (auto-rebuild on changes)
bash scripts/build-css.sh --watch

Git Submodules

# Init after cloning
git submodule update --init --recursive

# Force-refresh vendor skills
git submodule update --remote --merge skills-vendor/gregoryfoster-skills skills-vendor/obra-superpowers

The test database

tests/conftest.py builds watcher's tables in TEST_DATABASE_URL with Base.metadata.create_all and drops them at session end; per-test isolation is db_session's savepoint rollback. Tests that bypass db_session and write directly via engine.connect() leak rows into subsequent sessions — don't do that.

No sibling checkout (#311). The suite used to subprocess-run the Archiver checkout's alembic, located by ARCHIVER_REPO_PATH, to build an information schema its factories wrote rows into. Those rows only minted two ULIDs: production has had no information schema since #271 (see docs/MIGRATIONS.md → "Drop the dead information schema"), and the WatchedItem links carry no FK. make_watched_item now mints both links, test_engine drops a leftover information schema from a pre-#311 session, and tests/test_archiver_isolation.py fails if the test database carries one again. No migration may create or drop it either — tests/test_migration_chain.py.

pytest-xdist is unsupported. The fixture writes to a single TEST_DATABASE_URL; multiple xdist workers would race on creating and dropping the watcher tables. Tracked in #150 — worker-id-suffixed databases are the rework when xdist is actually adopted.

CI (#220)

GitHub Actions (.github/workflows/ci.yml) runs on push/PR to main: a lint job (ruff check + ruff format --check), a test job (pytest -m "not integration" against a postgres:16 service), and a migrations job (independent migration-chain smoke-check, #234 — alembic upgrade head from an empty postgres:16 then alembic check for drift). No job checks out a sibling repo — #254 removed the archiver-client path dep that lint and migrations needed, and #311 the test job's archiver alembic run. All three jobs authenticate to GCS keyless via WIF (vars.GCP_WIF_PROVIDER → co-pypi-reader SA) and sync the wheelhouse before uv sync; notifier-client installs as written, from its public HTTPS tag source — no URL rewrite (#284). The migrations job needs nothing from archiver either — the #234 squash collapsed the pre-existing chain into a self-contained genesis baseline (2addddea0b03) that references no information schema, so upgrade head from empty is fully standalone (no archiver seeding, no cross-service ordering). Squash cutover: already-migrated DBs need a one-time alembic stamp 2addddea0b03 --purge before their next upgrade — see docs/MIGRATIONS.md → "Migration baseline (squash)". Integration tests hit live external services and are excluded in CI. One-time GCP grant (operator, for WIF) — bind watcher's repo to the read-only SA; the org-scoped github-ci provider needs no change:

gcloud iam service-accounts add-iam-policy-binding \
  co-pypi-reader@co-gcs.iam.gserviceaccount.com --project=co-gcs \
  --role=roles/iam.workloadIdentityUser \
  --member="principalSet://iam.googleapis.com/projects/912903030445/locations/global/workloadIdentityPools/github/attribute.repository/CannObserv/watcher"