Open-source, self-hosted ML experiment tracking platform -- a drop-in replacement for Weights & Biases.
Track experiments, run hyperparameter sweeps, trace LLM calls, version artifacts, manage models, and share reports. Same wandb.init() / wandb.log() API you already know.
Captured against openrun.gladia.io at 1440×900 (2× retina). Full set in docs/screenshots.
Welcome page listing organizations, projects, and a copy-paste SDK quick-start with framework-specific snippets (PyTorch, Lightning, Transformers, Keras, …).
Searchable, filterable run list with config columns, status pills, tag filters, and column controls.
Live metric panels with Step/Wall/Rel x-axes, Line/Area chart modes, Auto/Outlier smoothing, and Full Fidelity toggle.
Multi-run overlay view — pick runs from the table, get side-by-side metric charts plus chart-type swap (Line, Bar, Scatter, Box Plot, Grid, Column View) and PNG/SVG/PDF export.
Versioned datasets, models, and reports with type pills (MODEL/REPORT) and per-version timestamps.
Cross-project model registry with stage badges (DEVELOPMENT/STAGING/PRODUCTION), aliases (latest, …), and grid/table view toggle.
Empty-state views below — captured against a project with no traces / sweeps / reports yet, so you can see the empty UX before you log your first one. Rich-data captures will replace these once a public demo project ships.
| Traces | Sweeps | Reports |
|---|---|---|
![]() |
![]() |
![]() |
Email/password login plus per-environment SDK key management.
| Login | API keys |
|---|---|
![]() |
![]() |
- W&B-compatible SDK -- switch by changing one import (
import openrunner as wandb) - Experiment tracking -- metrics, hyperparameters, system stats (CPU/GPU/memory)
- Run dashboard -- real-time charts, search/filter/sort, 10K+ run performance
- Run comparison -- overlay charts, bar/scatter/box plots, config diff
- Custom dashboards -- drag-and-drop chart panels with persistent layouts
- Full-fidelity export -- CSV, JSON, Parquet (no sampling limits)
- Grid, random, and Bayesian search strategies
- Hyperband early termination -- stop underperforming runs automatically
- W&B-compatible sweep config -- same YAML format, same API
- Parallel coordinates visualization -- explore parameter-metric relationships
@openrunner.tracedecorator -- capture inputs, outputs, timing, and errors- OpenAI auto-patching --
openrunner.trace.patch_openai()instruments all completions - Async support -- trace sync and async functions with automatic span nesting
- Waterfall visualization -- inspect call chains, latency, and token usage
openrunner.launch()-- submit jobs to remote workers with Docker imagesopenrunner.launch.from_run()-- re-launch from a previous run's config- Job lifecycle --
.wait(),.cancel(),.refresh()for monitoring - Worker queue -- Redis-backed job dispatch with configurable concurrency
- Artifact versioning -- datasets, models, checkpoints with SHA-256 content-addressed dedup
- Model registry with aliases --
latest,production,stagingfor deployment workflows openrunner.link_artifact()-- upload and set aliases in one callname:aliassyntax --openrunner.use_artifact("my-model:production")- Lineage graph -- track producer/consumer relationships across runs
openrunner.alert()-- send alerts from training code (INFO/WARN/ERROR)- Slack webhooks -- receive alerts in your Slack channel
- Email notifications -- delivery via Resend API
- Console fallback -- alerts logged when no external channel is configured
- Shareable reports -- snapshots with metric charts and run data
- Report anonymization -- strip identifiers for blind paper reviews
- Real-time collaborative editing -- WebSocket-based concurrent editing
- Image -- numpy arrays, PIL Images, file paths with captions
- Table -- structured columnar data
- Audio -- WAV serialization from numpy arrays or file paths
- Video -- file paths or numpy array sequences
- Histogram -- distribution visualization from numeric arrays
- Plotly -- interactive Plotly figures (JSON serialization)
- PlotlyChart -- enhanced Plotly with static PNG fallback
- MatplotlibFigure -- capture matplotlib figures as PNG
- PointCloud3D -- 3D point cloud visualization
- BoundingBoxes2D -- bounding box overlay on images
- SAML/OIDC SSO -- Okta, Azure AD, OneLogin, and any SAML 2.0 / OpenID Connect IdP
- OAuth -- Google and GitHub social login
- Audit logs -- compliance-ready event trail with filtering by action, user, resource, date
- Organization management -- teams, roles (admin/member), invitations
openrunner.Api()-- read-only access to projects, runs, artifacts- Filter and sort --
api.runs("project", filters={"state": "finished"}, order="-summary.accuracy") - History export --
run.history(pandas=True)returns a DataFrame
- PyTorch -- gradient logging
- HuggingFace Transformers --
OpenRunnerCallbackfor Trainer - PyTorch Lightning --
OpenRunnerLoggerfor pl.Trainer - Keras -- callback for training loops
- XGBoost -- callback for boosting rounds
- Scikit-learn -- experiment logging wrapper
- FastAI -- Learner callback
- LangChain -- chain/agent tracing callback
- Self-hosted -- Docker Compose or Kubernetes, your data stays yours
- Offline mode -- JSONL storage, idempotent sync when back online
- Kubernetes Helm chart -- production-ready deployment with HA
- Email delivery via Resend -- transactional emails without SMTP setup
flowchart LR
subgraph clients["Clients"]
SDK["Python SDK<br/>(W&B-compatible)"]
CLI["CLI"]
BROWSER["Browser"]
MCP["MCP server<br/>(Claude Code / agents)"]
PROD["Production inference<br/>(drift events)"]
end
subgraph platform["Self-hosted platform (Docker Compose / Kubernetes)"]
WEB["Web UI<br/>React SPA · nginx"]
API["API<br/>FastAPI · gunicorn workers"]
subgraph tasks["Background tasks (asyncio)"]
CRASH["crash detector"]
GC["artifact GC"]
ARCHIVER["drift S3 archiver"]
CONSUMER["drift ingest consumer"]
end
PG[("PostgreSQL + TimescaleDB<br/>runs · orgs · projects · registry<br/>dataset catalog · lineage · governance<br/>metric_points_ts (partitioned)<br/>drift events (hypertable)")]
REDIS[("Redis<br/>cache · rate limiting · launch queue<br/>drift ingest buffer (capped stream)")]
end
subgraph storage["Object storage"]
S3[("S3 / MinIO (default bucket)<br/>+ per-org vault connections<br/>(artifacts · datasets · exports · drift)")]
DS[("Datasets<br/>WebDataset tar shards + JSON index<br/>parquet · managed uploads")]
LANCE[("LanceDB indexes<br/>dataset FTS / vectors")]
end
BROWSER --> WEB
WEB -->|"REST /api/v1"| API
SDK -->|"REST /api/v1"| API
CLI --> API
MCP --> API
PROD -->|"drift ingest (429 backpressure)"| API
API --> PG
API -->|"XADD · no pg conn on hot path"| REDIS
CONSUMER -->|"bulk INSERT batches"| PG
REDIS --> CONSUMER
API -->|"uploads · presigned URLs · Range GETs"| S3
tasks --> PG
ARCHIVER -->|"parquet day partitions"| S3
S3 --- DS
S3 --- LANCE
BI["BI / DuckDB / warehouse"] -->|"read_parquet · CSV export"| S3
Key flows:
- Experiment tracking — the SDK posts runs, metrics, and logs to the API; metrics land in a time-partitioned Postgres table and stream to the UI charts.
- Datasets — the catalog (metadata, tags, cards, splits, lineage, governance scoring) lives in Postgres; bytes live on S3 under per-org storage connections (credentials encrypted at rest, uploads flow through the API so every file is tracked with sha256/uploader). Audio ships as WebDataset tar shards + a JSON byte-offset index, previewed via ranged reads — no full-file downloads. Search reads LanceDB FTS/vector indexes stored next to the data; runs link datasets for end-to-end lineage (dataset → run → model).
- Artifacts & models — versioned files on S3, content-addressed dedup; Postgres keeps versions, aliases (
production,latest), lineage, and model cards. Artifacts double as git/HF-LFS repos (realgit clone) and can be pushed/pulled through the built-in Docker Registry v2 broker (short-lived tokens; upstream PAT never touches the client). - Governed compute — cloud GPUs (RunPod, Vast.ai, Lambda, OVH) are provisioned server-side with the org's encrypted provider key; serverless fleets (Modal, Forge/Shadow-PaaS) run functions on demand via the stored token instead of a pod. Self-hosted machines join either as pull-based runners (poll jobs over outbound HTTPS, no inbound SSH) or SSH hosts. Cost guards accrue spend and auto-terminate idle/over-budget pods; serverless deployments reconcile live replicas toward a desired count.
- Inference drift monitoring — production services bulk-report per-inference metrics. Ingest is buffered through a capped Redis stream (no Postgres connection on the hot path; above high water producers get
429 + Retry-After); a consumer group drains it with bulk inserts. A TimescaleDB hypertable holds a compressed hot window, and a background archiver exports older days to S3 as BI-ready parquet (daily rollups keep long-range charts alive). - Background tasks — plain asyncio tasks in the API lifespan (no broker to operate): crash detection, artifact GC, GPU cost sweep + reconcile, deployment controller, drift ingest consumers, drift archival (serialized across workers by a Postgres advisory lock).
Design rule: Postgres stays small and fast (bounded tables, partitions, retention), S3 is the permanent byte store, Redis absorbs bursts and is never a source of truth.
Every moving part, by layer. Counts are approximate and grow — this is the map, not a cap.
Backend API — FastAPI on a gunicorn/uvicorn async worker pool.
| Group | What's inside |
|---|---|
REST routers (app/api/v1, ~70+) |
auth · sso (OIDC/SAML) · scim (2.0) · users · sessions · api-keys · organizations · projects · roles · runs · metrics · streams (SSE) · sweeps · jobs · launch · runners · compute · traces · feedback · evaluations · eval_automations · guardrails · monitors · prompts · artifacts · registry · repo (git/LFS) · media · datasets · governance · catalog · annotations/annotation_specs · collections · leaderboards · hub · hf (HuggingFace sync) · deployments · gpu · drift/inference_drift · traffic · efficiency · alerts · notifications · automations · reports · dashboards · chart_presets · table_views · pinned_runs · tags · search · papers · blogs · decisions · research/crews · patent_lab · ai_sessions · collab · export · retention · storage_connections · audit_logs · security_events · admin · resolve · health · otel (OTLP ingest) |
| Root-mounted | Docker Registry v2 broker · SCIM 2.0 provisioning |
Services (app/services, ~80) |
run · metric · artifact · dataset_version · dataset_index · storage/storage_resolver/storage_connection · gpu_orchestration/gpu_provision/gpu_cost/gpu_prices/gpu_reconcile · deployment/deployment_controller/deployment_router · runner · job · launch · sweep · trace · evaluation/online_eval/llm_judge · guardrails · prompt · drift/inference_drift(+queue,+archive,+nlq) · traffic · efficiency · report/paper_agents/blog_agents · research_agents · patent_agent/patent_report · registry_proxy/registry_tokens/container_registry · git_http · hf_sync · email · notification · alert · automation · organization/project/role/api_key/auth · sso/scim/security/audit · retention · summary · markdown_render · llm_pricing · … |
| Middleware | security headers (HSTS/CSP) · body-size guard · Redis sliding-window rate limiter (buckets: default/auth/api_write/dataset_upload, env-tunable) · CORS |
| Lifespan background tasks | crash detector · artifact GC · GPU cost sweeper · GPU reconciler (auto-terminate) · deployment controller · drift ingest consumer · drift → S3 parquet archiver |
| Data layer | ~70 SQLAlchemy async ORM models · ~40 Pydantic schema modules · 90+ Alembic migrations · TimescaleDB hypertables · partitioned metric table |
| Plugins & governance | pluggable startup/shutdown registry (reference: multi-tenant isolation plugin) · audio-schema + metrics-catalog governance |
| Auth & security | JWT access/refresh (token-versioned) · scoped or_* API keys (org/project) · OIDC + SAML SSO · SCIM 2.0 · fine-grained RBAC · TOTP MFA · Argon2 hashing · Docker secrets · audit + security-event trails |
Python SDK (pip install openrunner-sdk) — W&B-compatible, ~100 modules.
| Group | What's inside |
|---|---|
| Core (~50) | run · sender · buffer · wal · offline (crash-safe non-blocking metric pipeline) · api_client · config/summary · settings · environment/git_info · system_metrics · media (19 types) · plot · tensorboard · sweep · launch · trace · evaluation/scorers/dataset · prompt · guardrails · datasets · artifact · cache · model/model_card/handover · reload · migrate (W&B import) · sync · query_api · research/research_policy/machine · session (AI-session capture) · redact/pii/_secret_rules/_safepath · wer/transcript_formatter · cost · feedback · autoresearch · patent_scan |
| Framework integrations (~31) | pytorch · lightning · tensorflow · keras · jax · fastai · sklearn · xgboost · catboost · lightgbm · huggingface · trl · accelerate · diffusers · optuna · hydra · ultralytics · sb3 · gymnasium · ignite · langchain · llamaindex · openai_tracer · anthropic_tracer · openai_finetune · whisper · tts · voice_agent · forced_alignment · gladia · skypilot |
GPU orchestration (gpu/) |
sync facade · runner (SSH pods + run_forge_sync serverless) · api_integration · cost_guard · price_compare · workspace_sync · provider_base/gpu_types · providers: runpod · vastai · lambda_cloud · modal · ovhcloud · forge |
| Interfaces | MCP server (~50 tools: runs/metrics · sessions · papers/research/crews · decisions · datasets · models/artifacts · handover · patents · GPU/jobs/runners · drift/monitoring · reproducibility) · CLI (cli.py, groups: artifact datasets sweep launch/jobs gpu registry deploy drift machine runner session catalog handover mcp + login/init/sync/push/pull/migrate/patent-scan/install/update) · install_commands (Claude Code / Codex / Qwen / OpenCode slash commands + MCP) · wandb_compat drop-in |
Web frontend (src/web) — React 19 + React Router 7, Vite, TanStack Query/Table, ECharts + Vega-Lite; served in prod as a prebuilt static nginx image.
- ~40 pages / ~50 routes: dashboard · runs + run detail · run compare (+ weave) · artifacts + detail · model registry · datasets + detail (governance/lineage) · collections · sweeps · jobs · traces · evaluations · prompts (+ playground) · guardrails · monitors · drift · deployments · GPU (+ settings) · reports (+ shared/embed) · dashboards · papers · blog · research + research-crew · patent-lab · handover · hub · leaderboards · catalog · storage browser · org settings/roles/audit-log · api-keys · 2FA · invitations · docs · auth (login/register/callback/CLI-auth/email flows).
- ~36 hooks, 5 contexts (auth/tenant/theme/toast/project), ~100 feature components across artifacts/charts/comparison/crew/dashboard/datasets/gpu/hf/hub/lineage/markdown/media/reports/run/tables/tags, 21 base UI components + theme tokens, SSE for live metrics.
JS/TS SDK (sdk-js) — npm package (ESM + CJS) for browser/Node tracking.
Storage & data — PostgreSQL 16 + TimescaleDB · Redis (cache · rate limit · launch queue · drift stream) · MinIO/S3 default bucket + per-org storage connections (S3 / GCS / Azure, encrypted) · LanceDB FTS/vector indexes · WebDataset tar shards + JSON offset index · parquet exports.
Ops & deploy — docker-compose.yml (api · web · postgres · redis · minio) · production Helm chart (helm/, ~15 templates: deployments, services, ingress, HPA, secrets, stateful backends) · host nginx (two vhosts: app + registry.* broker, TLS/HSTS, SSE passthrough, large-upload streaming) · 5 GitHub Actions: ci (backend + SDK tests + coverage) · security-scan (pip-audit/bandit) · smoke (compose E2E) · publish-sdk · docs (mkdocs).
OpenRunner is MIT-licensed and free forever. There is no paid tier today. Every feature in this repo — SAML/OIDC SSO, full-fidelity export, sweeps, registry, LLM tracing — is in the OSS build with no enterprise gate. We commit to never paywalling existing OSS features; new paid SaaS features may exist in the future, but the OSS feature set will not regress.
Full comparison vs ClearML, W&B Pro, and Valohai (with sources): docs.openrunner.io/pricing.
git clone https://github.com/jqueguiner/openrunner.git
cd openrunner
docker compose up -dAll 5 services start automatically: PostgreSQL, Redis, MinIO, API server, frontend.
Open http://localhost:3000 and create an account.
Hit a snag on first run? See Troubleshooting first-run for the five issues most fresh-clone users encounter (port mismatch, empty
SECRET_KEY, MinIO endpoint and credential drift).
pip install openrunner-sdkexport OPENRUNNER_API_KEY="or_your_key"
export OPENRUNNER_BASE_URL="http://localhost:8000"Get your API key from the dashboard under Settings > API Keys.
import openrunner
openrunner.init(project="my-project", config={"lr": 0.001, "epochs": 10})
for epoch in range(10):
loss = train(epoch)
acc = evaluate()
openrunner.log({"loss": loss, "accuracy": acc, "epoch": epoch})
openrunner.finish()import openrunner as wandb
wandb.init(project="my-project")
wandb.log({"loss": 0.5})
wandb.finish()For production deployments, use file-based secrets instead of storing credentials in .env. Docker Compose mounts secret files read-only at /run/secrets/ inside the container. The app reads *_FILE env vars and uses the file content as the value.
# Generate strong secret files (won't overwrite existing)
./scripts/generate-secrets.sh
# Start with file-based secrets
docker compose up -dmkdir -p secrets && chmod 700 secrets
# Generate each secret
openssl rand -hex 32 > secrets/secret_key
openssl rand -hex 32 > secrets/jwt_access_secret
openssl rand -hex 32 > secrets/jwt_refresh_secret
# Database URL (edit for your setup)
echo "postgresql+asyncpg://user:pass@host:5432/openrunner" > secrets/database_url
# Lock permissions
chmod 600 secrets/*Any config field can be loaded from a file. Create the file, add it to the secrets: section of docker-compose.yml, and set the corresponding *_FILE env var:
# docker-compose.yml
services:
api:
environment:
RESEND_API_KEY_FILE: /run/secrets/resend_api_key
secrets:
- resend_api_key
secrets:
resend_api_key:
file: ./secrets/resend_api_keySupported *_FILE fields: SECRET_KEY, JWT_ACCESS_SECRET, JWT_REFRESH_SECRET, DATABASE_URL, RESEND_API_KEY, OPENAI_API_KEY, OPENROUTER_API_KEY, MINIO_ACCESS_KEY, MINIO_SECRET_KEY, SCIM_TOKEN, SMTP_PASSWORD, LDAP_BIND_PASSWORD, GOOGLE_CLIENT_SECRET, GITHUB_CLIENT_SECRET, OIDC_CLIENT_SECRET, AZURE_STORAGE_KEY.
- Docker Compose mounts
./secrets/read-only at/secrets/inside the API container *_FILEenv vars (e.g.SECRET_KEY_FILE=/secrets/secret_key) tell the app where to look_resolve_secret_file()inconfig.pyreads the file content and uses it as the field value- File content takes precedence over the plain env var
- Missing files are silently ignored — the app falls back to env vars / defaults
- The
secrets/directory is gitignored (only*.examplefiles are tracked)
No data loss — the app reads file-based secrets first and falls back to .env. To migrate incrementally:
- Run
./scripts/generate-secrets.sh(or copy values from.envintosecrets/files) - Remove the migrated keys from
.env - Restart:
docker compose up -d
Super admins can restrict registration to specific email domains (e.g. only @yourcompany.com).
# Set yourself as super admin in .env
SUPER_ADMIN_EMAILS=["admin@yourcompany.com"]
# Then via the API:
curl -X PUT /api/v1/admin/allowed-domains \
-H "Authorization: Bearer $TOKEN" \
-d '{"allowed_email_domains": ["yourcompany.com"]}'Empty list = open registration (default). Enforced on all auth entry points: register, OAuth, SSO, LDAP, SCIM, and password reset.
OpenRunner ships a self-hostable LLM observability surface — traces, monitors, guardrails, and human feedback — alongside experiment tracking. MIT-licensed, no per-trace metering. Self-discover the building blocks below; the unified /observability route lands with the N1 milestone (see GLA-104).
| Capability | Surface | Code |
|---|---|---|
| Trace decorator + native API | @openrunner.trace, span nesting, async support |
sdk/openrunner/trace.py |
| LangChain auto-instrumentation | OpenRunnerCallbackHandler for LLMChain, agents, tools |
sdk/openrunner/integration/langchain.py |
| OpenAI / Anthropic auto-patching | openrunner.trace.patch_openai() and Anthropic tracer |
openai_tracer.py, anthropic_tracer.py |
| OTEL ingest | OTLP HTTP endpoint for non-Python SDKs | src/api/app/api/v1/otel.py |
| Monitors | Project-scoped CRUD + scorer registry over trace/run streams | src/api/app/api/v1/monitors.py, src/api/app/services/monitors.py |
| Eval automations | Score-drift alerts, per-automation history feed, project aggregate (Wave B) | src/api/app/api/v1/eval_automations.py, src/api/app/services/eval_automations.py |
| Guardrails | PII / policy checks on prompts and completions | src/api/app/api/v1/guardrails.py, sdk/openrunner/guardrails.py |
| Annotation specs | Reviewer-defined label schemas for traces | src/api/app/api/v1/annotation_specs.py |
| Human feedback | Thumbs / scalar / categorical feedback on traces and runs | src/api/app/api/v1/feedback.py, sdk/openrunner/feedback.py |
examples/llm_observability_langchain.py—@traceon a LangChain chain end-to-end against a freshdocker-compose upstack (ships with GLA-602).
TODO: screenshot of
/observabilityroute once GLA-605 (route shell) and GLA-606 (widgets) land.
import openrunner
sweep_config = {
"method": "bayes",
"name": "lr-sweep",
"metric": {"name": "val_loss", "goal": "minimize"},
"parameters": {
"lr": {"min": 0.0001, "max": 0.1, "distribution": "log_uniform_values"},
"batch_size": {"values": [16, 32, 64]},
"epochs": {"value": 10},
},
"early_terminate": {"type": "hyperband", "min_iter": 3, "eta": 3},
}
sweep_id = openrunner.sweep(sweep_config, project="my-project")
def train_fn():
run = openrunner.init()
lr = openrunner.config.lr
for epoch in range(openrunner.config.epochs):
loss = train(lr=lr, epoch=epoch)
openrunner.log({"val_loss": loss})
openrunner.finish()
openrunner.agent(sweep_id, function=train_fn)Sweep methods: grid, random, bayes. Early termination via Hyperband stops underperforming runs.
import openrunner
# Decorator-based tracing
@openrunner.trace
def summarize(text: str) -> str:
response = openai.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": f"Summarize: {text}"}],
)
return response.choices[0].message.content
# Or auto-patch all OpenAI calls
openrunner.trace.patch_openai()
openrunner.init(project="llm-app")
result = summarize("Long article text...")
openrunner.finish()Traces capture inputs, outputs, duration, errors, and token usage. View them in the Trace Waterfall UI.
import openrunner
# Submit a remote training job
job = openrunner.launch(
project="my-project",
image="pytorch/pytorch:2.1.0-cuda12.1-cudnn8-runtime",
command="python train.py --lr 0.001 --epochs 50",
name="gpu-training-run",
)
# Wait for completion
final_state = job.wait(poll_interval=10.0)
print(f"Job finished with state: {final_state}")
# Re-launch from a previous run's config
job2 = openrunner.launch.from_run(
run_id="abc12345",
image="pytorch/pytorch:2.1.0-cuda12.1-cudnn8-runtime",
command="python train.py",
)import openrunner
openrunner.init(project="my-project")
# Log a model artifact with aliases
artifact = openrunner.Artifact(name="resnet50", type="model")
artifact.add_file("model.pth")
openrunner.link_artifact(artifact, aliases=["production", "v2.1"])
openrunner.finish()
# Later -- download by alias
openrunner.init(project="my-project")
model_dir = openrunner.use_artifact("resnet50:production")
# model_dir is a Path to the cached local directoryBuilt-in aliases: latest (auto-set), plus custom aliases like production, staging, best.
import openrunner
api = openrunner.Api()
# List projects
projects = api.projects()
# Query runs with filters and ordering
runs = api.runs("my-project", filters={"state": "finished"}, order="-summary.accuracy")
# Access run details
for run in runs:
print(f"{run.name}: accuracy={run.summary.get('accuracy')}")
# Get full metric history
run = api.run("my-project/abc12345")
history = run.history() # List of dicts
df = run.history(pandas=True) # pandas DataFrame
# Get artifact by alias
artifact = api.artifact("resnet50:production")import openrunner
openrunner.init(project="my-project")
# Send alerts from training code
openrunner.alert("Training complete", text="Final accuracy: 0.95", level="INFO")
openrunner.alert("Loss spike", text="Loss jumped from 0.1 to 5.0", level="WARN")
openrunner.alert("OOM Error", text="GPU out of memory at batch 1024", level="ERROR")
openrunner.finish()Configure Slack webhooks via the SLACK_WEBHOOK_URL environment variable. Email alerts use Resend (RESEND_API_KEY).
openrunner/
src/
api/ # FastAPI backend (Python) — routers · services · models · migrations
web/ # React 19 frontend (TypeScript, Vite) — served as static nginx in prod
sdk/ # Python SDK (pip install openrunner-sdk) — SDK · CLI · MCP · integrations · gpu
sdk-js/ # JavaScript/TypeScript SDK (npm)
helm/ # Kubernetes Helm chart
nginx/ # Production reverse-proxy config (app + registry vhosts)
scripts/ # entrypoint · smoke/e2e · secret generation · workers
catalog/ # org data-catalog package
examples/ # Demo scripts (MNIST, artifacts, LLM evals)
docs/ # MkDocs source (SDK ref · integrations · self-hosting)
docker-compose.yml
See Component inventory above for the exhaustive per-layer breakdown.
| Component | Technology |
|---|---|
| Backend | Python, FastAPI, SQLAlchemy (async), asyncpg |
| Frontend | React 19, TypeScript, Vite, ECharts, TanStack Table |
| Database | PostgreSQL 16 (partitioned metrics table) |
| Object storage | MinIO (S3-compatible) |
| Cache/pubsub | Redis |
| SDK | Python, httpx, Click CLI |
| Resend API (no SMTP needed) | |
| SSO | SAML 2.0, OpenID Connect |
| Deployment | Docker Compose, Kubernetes Helm chart |
The full REST surface (70+ router groups) is listed in the Component inventory.
Full API reference: openrunner-sdk on PyPI
import openrunner
# Initialize a run
run = openrunner.init(
project="my-project",
name="experiment-1",
config={"lr": 0.001},
tags=["baseline"],
notes="First experiment",
group="ablation-study",
resume=True, # or "must" to require an existing run
)
# Log metrics (non-blocking)
openrunner.log({"loss": 0.5, "accuracy": 0.85})
# Log media
openrunner.log({"samples": openrunner.Image(img_array, caption="epoch 5")})
openrunner.log({"audio": openrunner.Audio(waveform, sample_rate=16000)})
openrunner.log({"fig": openrunner.Plotly(plotly_fig)})
openrunner.log({"detections": openrunner.BoundingBoxes2D(image, boxes)})
# Log tables
table = openrunner.Table(
columns=["input", "predicted", "actual"],
data=[["img_01", 7, 7], ["img_02", 3, 5]],
)
openrunner.log({"results": table})
# Log and manage artifacts
artifact = openrunner.Artifact(name="my-model", type="model")
artifact.add_file("model.pth")
run.log_artifact(artifact)
# Config & Summary
openrunner.config["batch_size"] = 64
openrunner.config.optimizer.lr # dot notation access
openrunner.summary["best_accuracy"] = 0.95
# Alerts
openrunner.alert("Training done", level="INFO")
# Finish
openrunner.finish()# HuggingFace Transformers
from openrunner.integration.huggingface import OpenRunnerCallback
trainer = Trainer(callbacks=[OpenRunnerCallback()])
# PyTorch Lightning
from openrunner.integration.lightning import OpenRunnerLogger
trainer = pl.Trainer(logger=OpenRunnerLogger(project="my-project"))
# PyTorch
from openrunner.integration.pytorch import log_gradients
log_gradients(model)
# Keras
from openrunner.integration.keras import OpenRunnerCallback
model.fit(x, y, callbacks=[OpenRunnerCallback()])
# XGBoost
from openrunner.integration.xgboost import OpenRunnerCallback
xgb.train(params, dtrain, callbacks=[OpenRunnerCallback()])
# Scikit-learn
from openrunner.integration.sklearn import log_model
log_model(clf, X_test, y_test)
# FastAI
from openrunner.integration.fastai import OpenRunnerCallback
learn = cnn_learner(dls, resnet34, cbs=[OpenRunnerCallback()])
# LangChain
from openrunner.integration.langchain import OpenRunnerCallback
chain.invoke({"input": "hello"}, config={"callbacks": [OpenRunnerCallback()]})Install framework extras:
pip install openrunner-sdk[pytorch]
pip install openrunner-sdk[huggingface]
pip install openrunner-sdk[lightning]
pip install openrunner-sdk[keras]
pip install openrunner-sdk[xgboost]
pip install openrunner-sdk[sklearn]
pip install openrunner-sdk[fastai]
pip install openrunner-sdk[gpu] # NVIDIA GPU monitoringexport OPENRUNNER_MODE=offline
python train.py
openrunner sync # Upload when back onlineOffline runs are stored as JSONL files. Config is saved at init() time (not finish()), so it survives crashes.
openrunner login # Store API key
openrunner sync # Upload offline runs
openrunner ls # List projects and runsA production-ready Helm chart is included for Kubernetes clusters:
helm install openrunner ./helm/openrunner \
--namespace openrunner --create-namespace \
--set api.replicas=3 \
--set image.tag=0.2.0 \
--set env.DATABASE_URL="postgresql+asyncpg://user:pass@db:5432/openrunner" \
--set env.SECRET_KEY="your-production-secret"The chart deploys: API server, web frontend, PostgreSQL, Redis, MinIO -- all with resource limits, health checks, and PVC storage.
For external databases in production, disable the bundled PostgreSQL:
# values-production.yaml
postgresql:
enabled: false
env:
DATABASE_URL: "postgresql+asyncpg://user:pass@your-rds-host:5432/openrunner"See helm/openrunner/ for full chart documentation and values.yaml.
| Variable | Description | Default |
|---|---|---|
OPENRUNNER_API_KEY |
API key for server authentication | (required) |
OPENRUNNER_BASE_URL |
Server URL | http://localhost:8000 |
OPENRUNNER_PROJECT |
Default project name | (none) |
OPENRUNNER_MODE |
online or offline |
online |
OPENRUNNER_SYSTEM_METRICS |
Enable GPU/CPU/memory monitoring | true |
W&B environment variables (WANDB_API_KEY, WANDB_BASE_URL) are supported as fallback for migration.
| Variable | Description | Default |
|---|---|---|
DATABASE_URL |
PostgreSQL connection string | postgresql+asyncpg://openrunner:openrunner@postgres:5432/openrunner |
DB_USE_PGBOUNCER |
Enable PgBouncer compatibility mode | false |
REDIS_URL |
Redis connection string | redis://redis:6379/0 |
MINIO_INTERNAL_ENDPOINT |
MinIO/S3 endpoint reached from inside the docker/cluster network — used for API & worker reads/writes. Legacy MINIO_ENDPOINT is still accepted. |
minio:9000 |
MINIO_PUBLIC_ENDPOINT |
MinIO/S3 endpoint baked into presigned URLs returned to the browser/SDK. Must resolve from outside the cluster. Legacy MINIO_EXTERNAL_ENDPOINT is still accepted. |
localhost:9000 |
MINIO_ACCESS_KEY |
MinIO/S3 access key | minioadmin |
MINIO_SECRET_KEY |
MinIO/S3 secret key | minioadmin |
MINIO_SECURE |
Use HTTPS for MinIO | false |
MINIO_BUCKET |
Object storage bucket name | openrunner |
SECRET_KEY |
Application secret key | change-me-in-production |
DEBUG |
Enable debug mode | false |
API_HOST |
API bind address | 0.0.0.0 |
API_PORT |
API bind port | 8000 |
MAX_UPLOAD_SIZE |
Max artifact upload size accepted by the in-stack nginx /api/ proxy. Sized for typical ML model artifacts; lower it for shared deployments. Restart the web service to apply. |
5g |
JWT_ACCESS_SECRET |
JWT access token signing key | change-me-access-secret |
JWT_REFRESH_SECRET |
JWT refresh token signing key | change-me-refresh-secret |
JWT_ACCESS_EXPIRE_MINUTES |
Access token lifetime | 15 |
JWT_REFRESH_EXPIRE_DAYS |
Refresh token lifetime | 30 |
FRONTEND_URL |
Frontend URL for CORS and redirects | http://localhost:3000 |
About
FRONTEND_URLand CORS. The standard self-hosted topology serves the SPA and the API on the same origin behind nginx (http://localhost:3000indocker compose, your domain in production), so same-origin requests bypass CORS entirely and the default works out of the box. OverrideFRONTEND_URLonly when the SPA runs on a different origin from the API — e.g. a Vite dev server onhttp://localhost:5173hitting an API onhttp://localhost:8000— and point it at the SPA's actual origin so the browser is allowed to call the API.
| Variable | Description | Default |
|---|---|---|
GOOGLE_CLIENT_ID |
Google OAuth client ID | (empty) |
GOOGLE_CLIENT_SECRET |
Google OAuth client secret | (empty) |
GOOGLE_REDIRECT_URI |
Google OAuth callback URL | http://localhost:8000/api/v1/auth/google/callback |
GITHUB_CLIENT_ID |
GitHub OAuth client ID | (empty) |
GITHUB_CLIENT_SECRET |
GitHub OAuth client secret | (empty) |
GITHUB_REDIRECT_URI |
GitHub OAuth callback URL | http://localhost:8000/api/v1/auth/github/callback |
OIDC_CLIENT_ID |
OpenID Connect client ID | (empty) |
OIDC_CLIENT_SECRET |
OpenID Connect client secret | (empty) |
OIDC_DISCOVERY_URL |
OIDC discovery endpoint (.well-known/openid-configuration) |
(empty) |
OIDC_PROVIDER_NAME |
Display name for OIDC button | SSO |
OIDC_REDIRECT_URI |
OIDC callback URL | http://localhost:8000/api/v1/auth/oidc/callback |
SAML_ENTITY_ID |
SAML Service Provider entity ID | (empty) |
SAML_SSO_URL |
SAML IdP single sign-on URL | (empty) |
SAML_CERTIFICATE |
SAML IdP certificate (PEM) | (empty) |
SAML_SP_CERT |
SAML SP certificate (PEM) | (empty) |
SAML_SP_KEY |
SAML SP private key (PEM) | (empty) |
| Variable | Description | Default |
|---|---|---|
SLACK_WEBHOOK_URL |
Slack incoming webhook for alerts | (none) |
RESEND_API_KEY |
Resend API key for email delivery | (none) |
RESEND_FROM_EMAIL |
Sender email address | noreply@gladia.io |
For production (e.g., openrun.gladia.io):
# Use external PostgreSQL
cp docker-compose.prod.yml docker-compose.override.yml
# Edit .env with your database credentials
# Set up HTTPS
# Add DNS A record pointing to your server
# Use certbot or your reverse proxy for SSLSee docker-compose.prod.yml for external database configuration.
- Docker and Docker Compose (or Kubernetes with Helm)
- 2GB RAM minimum (4GB+ recommended for production)
- 10GB disk (grows with data)
- Python 3.10+ for the SDK
MIT licensed. PRs welcome.
# Development setup
git clone https://github.com/jqueguiner/openrunner.git
cd openrunner
pip install pre-commit && pre-commit install # secret-leakage guardrail (≤60s)
make install # Install backend deps
cd src/web && npm install # Install frontend deps
docker compose up -d # Start infra (DB, Redis, MinIO)
make dev # Start API dev server
cd src/web && npm run dev # Start frontend dev serverThe pre-commit install step is required — it wires up the
secret-leakage hook (gitleaks + OpenRunner ruleset) so a stray .env or a
real SECRET_KEY cannot land in a commit. The same scanner runs in CI as
the secret-scan job, so PRs that bypass the local hook still fail there.
See SECURITY.md for details.
The Backend Tests and SDK Tests CI jobs run pytest with coverage and
publish coverage.xml artifacts. Reproduce locally with the same commands CI
uses:
# Backend (FastAPI app) — ~45% line coverage
cd src/api && python -m pytest tests/ -q --cov=app --cov-report=term-missing
# SDK (openrunner) — ~12% line coverage (server/integration paths excluded)
cd sdk && python -m pytest tests/ -q --cov=openrunner --cov-report=term-missingThe badge above is a snapshot of the latest measured line coverage; the
authoritative numbers come from the coverage.xml artifacts on each CI run.









