This file is the canonical, tool-agnostic source of guidance for AI agents working in this repository. Tool-specific entry files (CLAUDE.md, GEMINI.md, .github/copilot-instructions.md) import or reference this file so there is exactly one place to edit project-wide rules.
AGENTS.md(this file) — repo-wide guidance: commands, architecture, conventions, agent roles, pipeline contracts.CLAUDE.md— Claude Code entry shim:@-imports this file and adds the Claude Code Skills (Superpowers) section.GEMINI.md— Gemini CLI entry shim:@-imports this file..github/copilot-instructions.md— GitHub Copilot guidance; points readers back here for the authoritative version.- App-local rules — each app's
AGENTS.mdcarries stage-specific contracts, module notes, and admin conventions:apps/alerts/AGENTS.mdapps/checkers/AGENTS.mdapps/intelligence/AGENTS.mdapps/notify/AGENTS.mdapps/orchestration/AGENTS.mdbin/AGENTS.md
docs/— long-form docs served via GitHub Pages (Architecture.md,Security.md,Installation.md, plan documents underplans/).
Django-based server monitoring and alerting system with a strict 4-stage orchestration pipeline: alerts → checkers → intelligence → notify. The orchestrator controls all stage transitions; stages never call downstream stages directly.
# Install dependencies
uv sync --extra dev
# Run tests
uv run pytest # All tests
uv run pytest apps/checkers/_tests/ # Single app
uv run pytest apps/checkers/_tests/checkers/test_cpu.py -v # Single file
# Code quality
uv run black . # Format
uv run ruff check . --fix # Lint + fix imports
uv run mypy . # Type check (optional)
# Pre-commit hooks
uv run pre-commit install
uv run pre-commit run --all-files
# Django
uv run python manage.py migrate
uv run python manage.py runserver
uv run python manage.py check # Django system checks
# Health checks
uv run python manage.py check_health # Run all checks
uv run python manage.py check_health --list
uv run python manage.py run_check cpu # Single checker
# System preflight
uv run python manage.py preflight # Dashboard + all checks
uv run python manage.py preflight --json # JSON output for CI
# Pipeline testing
uv run python manage.py run_pipeline --sample
uv run python manage.py run_pipeline --sample --dry-run
# Security
uv run pip-audit --strict --desc # Dependency CVE scan
uv run bandit -r apps/ config/ -c pyproject.tomlapps.alerts.ingest() → apps.checkers.run() → apps.intelligence.analyze() → apps.notify.dispatch()
Each stage emits monitoring signals (pipeline.stage.started, pipeline.stage.succeeded, pipeline.stage.failed) tagged with correlation IDs (trace_id, run_id).
Core rule: one orchestrator, one trace. Only the orchestrator (apps.orchestration) is allowed to move work from one stage to the next. Every pipeline run gets a correlation ID that must be attached to logs, monitoring events/spans, DB records / audit trail, and outbound notifications. Given a notification, you can jump back to the exact incident + checker output + analysis + retries/errors.
Hard boundary rule: stage code may call internal helpers in its own app, but must not call the next app directly. Only the orchestrator advances the pipeline.
Diagnostic I/O clarification: stages may call external systems (HTTP APIs, monitoring vendors) when needed to produce their own stage output — e.g. apps.checkers fetching StatusCake/uptime data or recent PagerDuty incident history. These calls must be treated as inputs only (no cross-stage advancement, no direct notifications). Always enforce timeouts/retries and redact secrets.
Apps under apps/ should follow this layout. A few legacy views.py modules — apps/alerts/views.py, apps/notify/views.py, apps/orchestration/views.py — are pending migration to the views/ package form; all new apps must use the package layout from day one.
views/— a package (not a monolithicviews.py), organized by endpoint (e.g.views/webhook.py,views/health.py)._tests/— a package mirroring the source structure (e.g._tests/views/test_webhook.py).AGENTS.md— app-specific AI agent guidance.admin.py— extensive admin for operations.
| App | Purpose | Key Models |
|---|---|---|
alerts |
Webhook ingestion (8 drivers) | Alert, Incident, AlertHistory |
checkers |
Health checks (CPU, memory, disk, disk_macos, disk_linux, disk_common, disk_inodes, network, process, raid, disk_temp, cpu_temp, io_strain, listening_ports) | CheckRun |
intelligence |
AI analysis via provider pattern | Uses StageExecution |
notify |
Notification delivery (Email, Slack, PagerDuty, Generic) | NotificationChannel |
orchestration |
Pipeline state machine, retry logic | PipelineRun, StageExecution, PipelineDefinition |
- Driver/Provider Pattern — all integrations inherit from abstract base classes (e.g.
BaseDriver,BaseChecker,BaseProvider). - DTOs — normalized data objects between stages (
ParsedPayload,CheckResult,AnalysisResult). - Correlation IDs — every pipeline run has
trace_idandrun_idfor tracing. - Stage configuration —
PipelineDefinitionrows are the routing table:matchconditions select a lane, its orderedstageslist (a subset of["check", "analyze", "notify"]) selects which downstream stages run, and its singlechannelFK is the notify target. Unmatched traffic fails non-retryably asno_route— there is no implicit fallback, only the seededcatch-allrow.NotificationChannel.is_activeandIntelligenceProvider.is_activefor DB-level enable/disable.
- ingest:
apps/alerts/AGENTS.md - diagnose:
apps/checkers/AGENTS.md - analyze:
apps/intelligence/AGENTS.md - communicate:
apps/notify/AGENTS.md - orchestration rules / state machine / node handlers:
apps/orchestration/AGENTS.md
Use the smallest agent that can complete the job safely and correctly.
- Plan — architecture, approach, multi-step work breakdown.
- Coder — implement code changes in specific files/directories.
- Debug — diagnose and fix failing tests/errors/logs.
- Review — quality, security, correctness, style, performance, edge cases.
- Docs — update docs/READMEs/usage guides for implemented changes.
For non-trivial changes, start with Plan, then hand the plan to Coder, then use Review and Debug as needed.
Purpose: research and outline multi-step plans for complex monitoring workflows and architectural changes.
When to use:
- Adding drivers: designing new inbound alert drivers (e.g. adding Grafana webhooks to
apps/alerts/drivers/). - New checkers: architecting new system checkers (e.g. adding a Kubernetes pod status checker to
apps/checkers/). - Intelligence: planning LLM prompt strategies for incident analysis in
apps/intelligence/. - Communication: adding new notification drivers (e.g. PagerDuty or MS Teams) to
apps/notify/drivers/.
Plan deliverable (handoff contract):
- Files to add/change (paths and brief purpose)
- Public interfaces (classes/functions, method signatures)
- Config/settings/env vars (and defaults)
- Error handling + edge cases
- Tests to add/update
- Acceptance criteria (what "done" means)
Purpose: implement specific logic and code changes, following the project's conventions.
When to use:
- Implementing a driver/checker/provider described in a plan
- Refactoring a module or adding a small feature with clear scope
- Adding tests and wiring configuration
Coder deliverable:
- Code changes in the specified folders/files
- Minimal, well-scoped diffs
- Tests updated/added for the new behavior
- Notes on how to run/verify locally
Purpose: troubleshoot errors, failing tests, runtime exceptions, incorrect behavior, and deployment issues.
When to use: failing CI/test output, stack traces, migrations failing, driver payload parsing issues, unexpected alerts / duplicated incidents / timeouts.
Debug deliverable: root cause explanation, minimal fix, regression test (when reasonable), verification steps.
Purpose: improve correctness, readability, security, performance, and consistency without changing intended behavior.
When to use: before merging a PR, after a large Coder change, when adding anything security-sensitive (webhooks, tokens, external APIs).
Review checklist highlights:
- Input validation for external payloads
- Idempotency for inbound alerts (avoid duplicate incidents)
- Timeouts/retries/backoff for outbound calls
- Avoid logging secrets and full payloads containing credentials
- Clear exception handling with actionable logs
Purpose: keep documentation in sync with behavior and configuration. Used when adding new drivers/checkers/providers, new env vars or settings, or new management commands/runbooks.
pipeline.stage.startedpipeline.stage.succeededpipeline.stage.failed(withretryable=true/false)- Duration metric (stage timing)
- Counters for retries and failures
Minimum tags/fields on every signal: trace_id / run_id, incident_id, stage (alerts|checkers|intelligence|notify), source (grafana/alertmanager/custom), alert_fingerprint, environment, attempt.
Artifacts to attach (or store refs to):
- Normalised inbound payload ref (never raw secrets)
- Checker output ref
- Intelligence output ref (prompt/response refs, redacted)
- Notification delivery refs (provider message IDs, response codes)
Rule: never log secrets; payloads and prompts should be stored as redacted refs and only selectively attached.
- The orchestrator decides whether a failure is retryable.
- Prefer stage-local retries with backoff for transient I/O (HTTP timeouts, provider 5xx).
- Prefer idempotency keys for outbound notify to prevent duplicate messages.
- If
apps.intelligencefails, the pipeline may still notify with a "no AI analysis available" fallback (configurable), but must record that downgrade in monitoring + audit trail.
start pipeline span (trace_id)
run alerts.ingest() → record + emit signals
run checkers.run() → record + emit signals
run intelligence.analyze() → record + emit signals
run notify.dispatch() → record + emit signals
close pipeline span
- Absolute imports always.
from apps.alerts.models import Incident— never relative. - App layout is required. Every app under
apps/<app_name>/must includeviews/(package),_tests/(mirrors source layout),AGENTS.md, and a substantiveadmin.py. - Django Admin is an operations surface. Admin should make it easy to manage models and trace pipeline behavior via
Incident,trace_id/run_id, and orchestration links. App-specific admin expectations live in each app'sAGENTS.md. The customMonitoringAdminSite(config/admin.py) drives the ops console: its dashboard (config/dashboard.py→get_dashboard_context()/build_readiness(), templatetemplates/admin/dashboard.html) shows a readiness panel (channels / LLM provider / preflight / inbox / nodes, green-amber-red) plus activity cards, and itsget_app_list()override regroups every registered model into operator-facing sections viaSECTION_MAP(Operations / Configuration / History & Audit; unmapped models fall into "Other" — a completeness test guards the map). Seedocs/plans/2026-08-11-admin-ops-console-design.md. - Driver / Provider pattern. New checkers, drivers, and providers must inherit from the project's abstract base classes (
BaseDriver,BaseChecker,BaseProvider, etc.). - 100% branch coverage on changed code. Verify with
uv run coverage run -m pytest && uv run coverage report. - Line length: 100 characters (Black + Ruff configured in
pyproject.toml). - Always use absolute paths. Resolve all file/directory paths to absolute form using
pathlib.Path.resolve()before use. Never pass user-supplied relative paths to file operations, subprocess calls, or provider methods. Validate that resolved paths fall within allowed directories to prevent path traversal. - Always use full executable paths for subprocess. Resolve via
shutil.which("toolname")and pass the absolute result asargv[0]— never a bare name like["less", "-FRX"]. Bare-name PATH lookups at exec time let an attacker-controlled PATH steer the call. Pair with# nosec B603 # nosemgrepon thesubprocess.Popenline so bandit and Semgrep's dynamic-argv detectors accept the (resolved) call as intentional. - Be safe with external I/O. Always set timeouts; handle retries; redact secrets from logs.
- Prefer small, testable units. Parse and validate payloads separately from side effects (DB writes, network calls).
- Reference existing code. Point agents to existing directories (e.g.
apps/checkers/) so new code matches the established pattern. - Package management via
uv.- Runtime deps:
uv add <package> - Dev tooling:
uv sync --extra dev - Django commands:
uv run python manage.py <command>
- Runtime deps:
- Environment variables. Copy
.env.sampleto.envfor local development. Main settings live inconfig/settings.py.
These rules exist because the observability effort (#154/#155/#156, reverted) reinvented a node→hub mechanism the project already had. See the post-mortem: docs/plans/2026-05-31-observability-overbuild-postmortem.md.
- Inventory before building. Before adding any new mechanism — endpoint, model, transport, management command — search the codebase for existing capability that already covers the need, and name it in the design. If it exists, reuse or extend it; do not build a parallel one. (The node→hub channel already existed:
push_to_hub+ the inboundClusterDriver+ the/alerts/webhook/cluster/path.) - Respect the established pattern. New integrations use the existing Driver/Provider base classes and the existing webhook ingestion path — not a parallel endpoint, auth scheme, or registry. If a design finds itself fighting the established pattern, stop and reconsider; that friction is a signal the approach is wrong.
- App vs. utility test. Cross-cutting concerns (logging, correlation IDs, formatting) belong in shared configuration/utilities consumed through standard interfaces (e.g. stdlib
logging), not in a Django app that other apps import. If an "app" ends up imported across the pipeline, it is a utility miscategorized as an app — that coupling is what turns a small feature into a 50-file revert. - Solve the topology you have. Do not build mesh, multi-hop, dedup, loop-prevention, or scaling machinery without a concrete, current requirement named in the design. Single-hop fan-in has no cycles to break and no duplicates to suppress. Relatedly, treat plan size as a smell, not rigor: an implementation plan that dwarfs the code it plans is a scope warning — re-scope before building.
The repo standardises on:
- Formatting: Black (configured in
pyproject.toml) - Linting / import sorting: Ruff (configured in
pyproject.toml) - Testing: pytest + pytest-django (configured in
pyproject.toml) - Optional typing: mypy + django-stubs
- Security: pip-audit (deps), bandit (code)
CI runs these in GitHub Actions (.github/workflows/ci.yml). Any PR should keep the following green:
uv run black . --checkuv run ruff check .uv run pytestuv run pip-audit --strict --desc- 100% branch coverage on changed lines —
uv run coverage run -m pytest && uv run coverage report
All markdown files under docs/ are served via GitHub Pages (Jekyll + Just the Docs).
Plan documents (docs/plans/) require Jekyll front matter:
---
title: "Plan Title Here"
parent: Plans
---If a plan contains Jinja2/template syntax ({% %}, {{ }}), wrap the entire content (after the front matter) in {% raw %}...{% endraw %} to prevent Jekyll from interpreting it as Liquid tags. GitHub Pages uses Jekyll 3.x, which does not support render_with_liquid: false.
Top-level docs under docs/ use title-case filenames (e.g. Architecture.md) and include:
---
title: Page Title
layout: default
nav_order: N
---Historical record: plan documents under docs/plans/ are immutable historical records. Do not modify them to clean up stale references; they should describe the state of the world at the time they were written.
A change is typically "done" when:
- Code follows the existing base class / module patterns.
- Config changes are wired correctly (settings/env).
- Tests achieve 100% branch coverage on changed code.
- Basic verification steps are provided (how to run / check locally).
- Docs are updated if behavior or config changed.
- Security tooling is clean (
pip-audit,bandit). - All CI checks are green on the PR.
| Agent | Use case | Example prompt |
|---|---|---|
| Plan | Multi-step planning & architecture | "Plan how to add a Disk Space checker to apps/checkers/" |
| Coder | Implementing specific logic | "Create the Slack notification driver in apps/notify/drivers/slack.py" |
| Debug | Troubleshooting errors | "Fix the circular import between apps.alerts and apps.checkers" |
| Review | Quality & security pass | "Review this webhook driver for validation & idempotency" |
| Docs | Documentation updates | "Document env vars + setup steps for the new driver" |