Existing observability platforms (Datadog, Honeycomb, Grafana) are built for humans staring at dashboards. AI agents — like Claude Code, Devin, or custom autonomous agents — need to monitor, debug, and react to production systems programmatically. They don't use browsers. They need structured, queryable, CLI-first access to traces, metrics, and logs.
There is no observability platform designed for machine consumption as a first-class interface.
- Ingest OpenTelemetry traces/logs and OpenMetrics/Prometheus metrics via standard protocols — no custom SDKs required.
- Provide a CLI as the primary interface that AI agents (and power-user humans) use to query, monitor, and alert on telemetry data.
- Return structured output (JSON, tables) that agents can parse and reason over without scraping HTML or interpreting screenshots.
- Support natural-language and structured queries so agents can ask "what's slow?" or run precise PromQL/trace filters.
- Optimize for agent workflows: correlation, root-cause suggestions, anomaly detection, and watch/subscribe patterns.
- Close the agent reliability loop: turn production traces into classified failures, high-signal golden cases, production signals, and experiment comparisons.
- Building a full GUI dashboard (out of scope for v1; a minimal web UI for humans is a future consideration).
- Replacing Prometheus or Jaeger — we sit on top of standard protocols, not beside them.
- Multi-tenancy or enterprise RBAC in v1.
┌─────────────────────────────────────────────────────────┐
│ Data Sources │
│ (any app instrumented with OTel SDK or Prometheus) │
└──────────┬──────────────────────┬───────────────────────┘
│ OTLP (gRPC/HTTP) │ Prometheus remote-write
▼ ▼
┌─────────────────────────────────────────────────────────┐
│ Ingestion Layer │
│ │
│ ┌──────────────┐ ┌───────────────┐ ┌──────────────┐ │
│ │ OTLP Receiver│ │ Prom Remote │ │ Log Receiver │ │
│ │ (traces+logs)│ │ Write Receiver│ │ (OTLP logs) │ │
│ └──────┬───────┘ └──────┬────────┘ └──────┬───────┘ │
│ └─────────┬───────┘───────────────────┘ │
│ ▼ │
│ ┌─────────────────┐ │
│ │ Pipeline │ │
│ │ (normalize, │ │
│ │ enrich, │ │
│ │ route) │ │
│ └────────┬────────┘ │
└──────────────────┼──────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────┐
│ Storage Layer │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Trace Store │ │ Metric Store │ │ Log Store │ │
│ │ (ClickHouse │ │ (ClickHouse │ │ (ClickHouse │ │
│ │ or DuckDB) │ │ or DuckDB) │ │ or DuckDB) │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└──────────────────┬──────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Query Engine │
│ │
│ - PromQL-compatible metric queries │
│ - Trace search (by service, span, duration, status) │
│ - Log filtering (structured field queries) │
│ - Cross-signal correlation (trace → metrics → logs) │
│ - Anomaly detection (baseline comparisons) │
└──────────────────┬──────────────────────────────────────┘
│
┌────────┴────────┐
▼ ▼
┌──────────────┐ ┌────────────────┐
│ CLI │ │ API Server │
│ (primary) │ │ (gRPC + REST) │
└──────────────┘ └────────────────┘
Accepts telemetry via standard protocols. No proprietary agents.
| Signal | Protocol | Port (default) |
|---|---|---|
| Traces | OTLP gRPC / OTLP HTTP | 4317 / 4318 |
| Metrics | OTLP gRPC / Prometheus remote-write | 4317 / 9090 |
| Logs | OTLP gRPC / OTLP HTTP | 4317 / 4318 |
The ingestion layer normalizes everything into a unified internal model before storage. We use the OpenTelemetry Collector as a reference but implement a purpose-built receiver to keep the binary small and focused.
v1 (single-node): DuckDB as an embedded columnar database. Zero external dependencies, single-file storage, and analytical query performance far beyond SQLite for time-series and aggregation workloads. DuckDB's columnar engine gives us ClickHouse-like query patterns (fast GROUP BY, window functions, range scans) locally without running a separate server.
v2 (scale): ClickHouse for distributed columnar storage. The migration path is smooth since both are columnar and share similar SQL dialects — queries written for DuckDB translate to ClickHouse with minimal changes.
Key schema concepts:
- Traces: stored as spans with parent references, indexed on service, operation, duration, status, and attributes.
- Metrics: time-series with labels, stored in a columnar layout optimized for range queries.
- Logs: structured records with indexed fields and full-text search on body.
The query engine supports two modes:
Structured queries — a filter DSL that maps cleanly to CLI flags:
tael query traces --service=api-gateway --min-duration=500ms --status=error --last=1h
PromQL-compatible metric queries:
tael query metrics 'rate(http_requests_total{status="500"}[5m])'
Cross-signal correlation — the killer feature. Given a trace ID, pull every span, every log tagged with that trace_id, and all metrics from the touched services inside the trace's time window:
tael correlate --trace <trace-id>
Future work: metric-driven correlation (--metric http_latency_p99 --threshold '>2s') to walk the other direction.
The CLI is the primary interface. It must be excellent for both AI agents and human power users.
- Every command returns structured JSON by default (
--format=json). Human-readable tables via--format=table. - Consistent flag patterns across all subcommands.
- Streaming support for watch/tail operations.
- Exit codes that encode error categories (not just 0/1).
- Built-in
--explainflag that adds plain-English context to results (useful for agents reasoning about output).
tael ingest status # health of ingestion pipelines
tael query traces [filters] # search/filter traces
tael query metrics [promql] # query metrics
tael query logs [filters] # search/filter logs
tael get trace <trace-id> # full trace waterfall as structured data
tael get metric <name> # describe a metric (type, labels, recent values)
tael correlate # cross-signal correlation
tael watch <query> # stream matching results in real-time
tael diff <query> --baseline=<range> # compare current vs baseline period
tael anomalies [--service=X] # surface anomalies detected over recent window
tael services # list known services and their health
tael topology # service dependency map from trace data
tael summarize --last=1h # agent-friendly summary of system health
tael summarize: returns a structured health summary an agent can use to decide what to investigate further. Includes: top errors, latency regressions, anomalous metrics, and recent deploys correlated with changes.tael anomalies: surfaces statistically significant deviations without requiring the agent to define thresholds.tael correlate: eliminates manual cross-signal pivoting — the agent says "this metric spiked, what's related?" and gets traces + logs back.tael watch: polls the summary endpoint on an interval and prints signed deltas per tick (span count, error count, error rate, p95, log errors, metric volume). The--exit-on=<condition>flag lets an agent subscribe to a query and exit once a threshold is crossed ("watch this deploy and tell me if error rate exceeds 1%").tael diff: compare a time range against a baseline. Agents use this to answer "is this deploy worse than the last one?"
gRPC and REST endpoints that mirror the CLI surface 1:1. The CLI is a thin client over this API. This means any agent that prefers HTTP can use the API directly.
All signals share a common envelope:
{
"timestamp": "2026-04-09T12:00:00Z",
"service": "api-gateway",
"environment": "production",
"signal": "trace|metric|log",
"attributes": { ... },
"resource": { ... },
"payload": { ... } // signal-specific
}
This allows cross-signal queries without joins.
{
"trace_id": "abc123",
"span_id": "def456",
"parent_span_id": "ghi789",
"operation": "HTTP GET /users",
"duration_ms": 142,
"status": "ok|error",
"events": [ ... ],
"links": [ ... ]
}
{
"name": "http_requests_total",
"type": "counter|gauge|histogram|summary",
"value": 42,
"labels": { "method": "GET", "status": "200" }
}
{
"severity": "ERROR",
"body": "connection refused to downstream",
"trace_id": "abc123",
"span_id": "def456",
"fields": { "retry_count": 3 }
}
| Component | Choice | Rationale |
|---|---|---|
| Language | Rust | Fast, single binary, memory-safe, no GC pauses |
| Storage (v1) | DuckDB (via duckdb-rs) | Embedded columnar DB, analytical perf locally |
| Storage (v2) | ClickHouse | Columnar, fast aggregations, proven at scale |
| CLI framework | clap | Standard Rust CLI library, excellent completions |
| API transport | tonic + axum | tonic for gRPC, axum for REST, both async on tokio |
| Serialization | serde + prost | serde for JSON/YAML, prost for protobuf decoding |
| Config | YAML | Familiar, easy for agents to read/write |
| OTel parsing | opentelemetry-rust | Official Rust SDK for OTLP decoding |
| Async runtime | tokio | Industry-standard async runtime for Rust |
tael server start --storage=sqlite --data-dir=./data
One process handles ingestion, storage, query, and API. Good for local dev, single-team use, or an agent monitoring its own infra.
Superseded: distribution is built on the tael-backend engine rather than ClickHouse — trace-id-sharded fan-out queries (TAEL_QUERY_SHARDS), WAL shipping to standbys, and gossip-based leader election. See docs/tael-server-scaling-ha.md for the current architecture and remaining phases.
The CLI is the immediate interface, but the natural evolution is an MCP server that exposes observability tools directly to agents:
{
"tools": [
{ "name": "query_traces", "description": "Search distributed traces", ... },
{ "name": "get_anomalies", "description": "Surface anomalous metrics", ... },
{ "name": "correlate", "description": "Cross-signal correlation", ... }
]
}This lets agents like Claude Code call observability tools without shelling out.
- Project scaffolding (Rust workspace, CI, linting)
- OTLP gRPC receiver for traces
- DuckDB trace storage with basic schema
-
tael query traceswith service/duration/status filters -
tael get trace <id>with structured JSON output
- OTLP metrics receiver
- Prometheus remote-write receiver
- DuckDB metric storage
- OTLP log receiver + storage
-
tael query metricswith PromQL subset -
tael query logswith field filters
-
tael summarize— system health digest -
tael anomalies— baseline-vs-current regression detection -
tael correlate— cross-signal correlation by trace ID -
tael watch— polling summary deltas -
tael eval— trace-native eval collection, scoring, reporting, and live progress (design) -
tael diff— baseline comparison -
tael topology— service dependency graph
-
tael issue— classify production stumbles into recurring failure patterns, backed by structured trace comments -
tael signal— define and trend long-running agent behaviors such as ignored tool errors, bad refusals, context loss, and user frustration -
tael eval case add --from-trace— promote a production trace into a golden regression case provenance record -
tael eval suite inspect— identify missing provenance, missing expected behavior, duplicate failure modes, and critical-path coverage -
tael experiment compare— compare production variants by error rate, latency, and optional issue/signal rate - self-diagnostic comment convention — allow agents to report suspected missing context, capability gaps, broken tools, or task failure without treating those reports as trusted facts
- Dedicated issue/signal tables and distributed query support if comment-backed conventions become limiting
-
ClickHouse storage backend— superseded: the purpose-built tael-backend engine (WAL + LSM hot tier + Parquet cold tier) replaced this plan; see docs/tael-backend-design.md and docs/tael-server-scaling-ha.md - MCP server integration (
tael mcp serve) - Retention policies and downsampling (
tael-server/src/retention.rs, 5m rollups in the cold tier) - Auth (API keys) (
tael-server/src/auth.rs, roles + authz middleware) - Packaging (Homebrew, Docker) (
packaging/homebrew/,Dockerfile,install.sh)
- Auth should be zero-friction for single-agent/local use and scale to multi-agent environments without rearchitecting.
- Agents are first-class principals — not users impersonating humans.
Single-node (v1): No auth required by default. The server binds to 127.0.0.1 only. If you can reach the socket, you're in. This matches the local-dev mental model — no API key ceremony to get started.
Enable auth explicitly with tael server start --auth=required.
Multi-agent: API key per agent identity. Keys are created via the CLI:
tael auth create-key --name="claude-code-prod" --role=reader
tael auth create-key --name="deploy-bot" --role=reader
tael auth create-key --name="admin" --role=admin
| Role | Capabilities |
|---|---|
reader |
Query traces, metrics, logs. Read anomalies, topology, summaries. |
writer |
Everything in reader + push telemetry via OTLP/remote-write. |
admin |
Everything in writer + manage keys, retention, server config. |
Most agents are reader — they consume observability data. The writer role exists for agents that also instrument and report their own telemetry. admin is for operators.
Keys are prefixed for easy identification: tael_r_<random> (reader), tael_w_<random> (writer), tael_a_<random> (admin). Passed via --api-key flag or TAEL_API_KEY env var.
Future: optional service-level scoping so an agent can only see telemetry from specific services. Not needed for v1 — most deployments are single-team.
| Signal | Raw Retention | Downsampled Retention | Rationale |
|---|---|---|---|
| Traces | 7 days | — | Large, high-cardinality; 7d covers most investigations |
| Metrics | 30 days (raw) | 1 year (5m rollups) | Agents need recent precision + long-term trends |
| Logs | 14 days | — | Middle ground; logs are verbose but searchable |
- Retention is enforced by a background cleanup job that runs hourly.
- Configurable per signal via
tael server start --retention-traces=7d --retention-metrics=30d --retention-logs=14d. - Also configurable in the YAML config file:
retention:
traces: 7d
metrics:
raw: 30d
downsampled: 365d
downsample_interval: 5m
logs: 14dRaw metric data points are rolled up into 5-minute aggregates (min, max, avg, sum, count) after the raw retention window. This preserves trend visibility for capacity planning and long-term comparisons while keeping storage bounded.
Assuming a mid-size deployment (~50 services, moderate traffic):
- Traces: ~2-5 GB/day → ~15-35 GB at 7d retention
- Metrics: ~500 MB/day raw → ~15 GB at 30d + ~5 GB/year downsampled
- Logs: ~1-3 GB/day → ~15-40 GB at 14d retention
Total: ~50-90 GB for a single-node DuckDB deployment. Well within local disk for most machines.
Naming: resolved —tael(trace agent event log). Short, unique, no conflicts.Storage default: resolved — the tael-backend engine (WAL + LSM hot tier + Parquet cold tier) is now the default storage backend; DuckDB is an opt-in Cargo feature. The single-writer throughput concern is addressed by the WAL/ingest-buffer design.Agent auth model: resolved — see Auth section below.Retention: resolved — see Retention section below.Natural language query layer: resolved — leave it to the calling agent. Tael returns structured data; the agent is already an LLM that can formulate queries and interpret results. Embedding a query translator adds complexity and a model dependency we don't need.