Skip to content

Repository files navigation

WMATA Open Data Demo Apps

A starter kit for the AU Kogod MS Analytics program. Built to demonstrate what's possible with modern agentic coding tools — Claude Code, OpenAI Codex, GitHub Copilot, and Perplexity Computer — when paired with a real public dataset (the DC Metro WMATA API). A working Streamlit dashboard, plus the research, tests, design rubric, and methodology notes that show how it was built — not just the finished product.

Unofficial dashboard. Not affiliated with, endorsed by, or sponsored by WMATA. Transit data is provided by the WMATA API as-is; accuracy is not guaranteed.


What's in the Box

A ready-to-run Streamlit dashboard with four tabs:

  • 🚇 Rail — live arrivals at any of 102 stations, line status, alerts, elevator outages
  • 🚌 Bus — geographic stop discovery near a Metro station, live bus arrivals, fleet positions, alerts
  • 📊 System — full system view (all 6 lines, all alerts, 30+ elevator outages grouped by station)
  • 📈 History(Phase 3) historical snapshots captured into DuckDB; row-count metrics, capture-run health, 6-hour arrival-time chart for the selected station

Plus a thin Python API client (wmata/client.py), a DuckDB capture pipeline (wmata/storage.py + wmata/capture.py), a 47-test suite, a documented test methodology, a UI/UX standards document with scoring rubric, and four reusable Claude skills.

Stack: Python 3.11 · Streamlit · DuckDB · Plotly · Pandas · uv

Accessibility: every visible text/background pair meets WCAG AA (4.5:1) and most reach AAA (7:1). See docs/UI_UX_STANDARDS.md.


Quick Start

# 1. Install uv (one-time, if needed)
curl -LsSf https://astral.sh/uv/install.sh | sh

# 2. Clone and install
git clone <your-fork-url>
cd "WMATA Open Data Demo Apps"
uv sync

# 3. Get a free WMATA API key (~ 2 minutes)
#    → register at https://developer.wmata.com
#    → subscribe to the "Default Tier" product
#    → copy your Primary key

# 4. Configure your key
cp .env.example .env
# Edit .env and paste your key after WMATA_API_KEY=

# 5. Run the app
.venv/bin/streamlit run app/dashboard.py
# Opens at http://localhost:8501

# 6. (Optional) run the test suite
.venv/bin/python tests/run_tests.py

Demo mode runs without a key — sample data renders so you can see the UI before signing up.


Project Structure

.
├── README.md                       ← you are here
├── pyproject.toml · uv.lock        ← reproducible install
├── .env.example                    ← template; copy to .env
├── .gitignore                      ← excludes .env, venv, caches, OS files
│
├── app/
│   └── dashboard.py                ← Streamlit web app (Rail / Bus / System tabs)
│
├── wmata/                          ← reusable API client library
│   ├── __init__.py
│   ├── client.py                   ← rail + bus endpoint wrappers
│   ├── storage.py                  ← Phase 3 — DuckDB schema + helpers
│   └── capture.py                  ← Phase 3 — capture loop
│
├── cli/                            ← Phase 2 — CLI tool (placeholder)
│
├── scripts/
│   └── capture.py                  ← Phase 3 — capture CLI (--once / --interval)
│
├── data/
│   ├── cache/                      ← runtime cache (gitignored)
│   ├── gtfs/                       ← GTFS files (gitignored)
│   └── wmata.duckdb                ← captured historical data (gitignored)
│
├── tests/
│   ├── run_tests.py                ← L1/L2 test runner (35 tests)
│   ├── test_methodology.md         ← three-level test pyramid design
│   └── test_results.md             ← results log per iteration
│
├── Research/                       ← background reading
│   ├── wmata_api_overview.md
│   ├── wmata_license_compliance.md
│   ├── bi_analytics_best_practices.md
│   ├── use_cases_and_personas.md
│   └── project_plan_ideas.md
│
├── docs/
│   ├── STUDENT_GUIDE.md            ← walkthrough + working with AI dev tools
│   └── DEPLOYMENT.md               ← hosting options for sharing your version
│
├── .claude/                        ← Claude Code skills + launch config
│   ├── launch.json
│   └── skills/                     ← reusable agent skills (see docs)
│
├── action_log.md                   ← chronological build log
└── lessons_learned.md              ← reflective notes session-by-session

Documentation

Doc When to read it
docs/STUDENT_GUIDE.md Setup walkthrough + advice for working with this code in Claude Code, OpenAI Codex, GitHub Copilot, and Perplexity Computer
docs/DEPLOYMENT.md Run locally for personal use, or deploy to free-tier hosting platforms
docs/R_VERSION_PLAN.md Plan for porting this dashboard to R (Shiny + golem, or Quarto Dashboards) using agentic AI tools
docs/PHASE3_PLAN.md Design rationale for historical capture: schema, capture-rate math, risks
docs/UI_UX_STANDARDS.md Color palette, type scale, spacing tokens, elevation, sources cited (WCAG, Material 3, Refactoring UI, NN/g)
docs/UI_UX_RUBRIC.md 10-dimension scoring rubric (1–5) for grading any page
docs/UI_UX_SCORECARD.md Baseline + post-overhaul scores per tab
tests/test_methodology.md The L1/L2/L3 test pyramid and iteration template
Research/wmata_api_overview.md API endpoints, rate limits, data quirks
Research/wmata_license_compliance.md What you may and may not do with WMATA data
action_log.md What was built when, in what order, and why
lessons_learned.md What worked, what didn't, what to do differently

Roadmap

Phase Focus Status
1 Streamlit dashboard (Rail · Bus · System tabs) ✅ Complete
2 CLI tool (wmata rail predictions, wmata status) 🚧 Planned
3 Historical data capture + DuckDB analytics ✅ Complete
4 On-time performance + headway analysis 🚧 Planned
5 GTFS-RT crowding + equity analysis 🚧 Planned

Phase 3 quickstart:

python scripts/capture.py --once          # single snapshot
python scripts/capture.py --interval 120  # capture every 2 minutes

After capture has run, the dashboard's 📈 History tab shows row counts, capture-run health, and an arrival-time chart for the selected station. See docs/PHASE3_PLAN.md for the full design.


How the History Feature Works

The History tab is built on a four-layer capture-and-query pipeline:

WMATA API  ──poll──▶  scripts/capture.py  ──insert──▶  data/wmata.duckdb  ──query──▶  📈 History tab
                       (one_capture or
                        capture_loop)

Capture (wmata/capture.py) A capture iteration polls predictions for a configurable list of stations (default: Metro Center, Gallery Pl, Foggy Bottom), all rail and bus alerts, and all elevator outages. Each row is timestamped with captured_at. A summary row is written to capture_runs after every iteration so health is queryable at a glance.

Storage (wmata/storage.py) A single DuckDB file at data/wmata.duckdb (gitignored). Six tables:

Table Grain
predictions one row per train per capture
bus_arrivals one row per bus prediction per capture
rail_incidents one row per active rail alert per capture
bus_incidents one row per active bus alert per capture
elevator_outages one row per outage per capture
capture_runs one row per capture iteration (health, counts, errors)

The schema is append-only with timestamps — no dedup at write time. Phase 4 analytics aggregate at query time. Predictions also store a precomputed min_int (with ARR/BRD/--- collapsed to NULL) so headway and on-time queries don't have to re-parse the string Min field.

Query helpers in wmata/storage.py:

  • get_capture_stats() — row counts per table, time span, last-run summary
  • get_arrival_history(station, hours) — recent numeric predictions for a station
  • get_recent_runs(limit) — last N capture iterations for the health table

Visualization (app/dashboard.py — History tab)

  • Five KPI tiles for row counts per table
  • Time-span caption (earliest → latest capture, total capture-runs count)
  • Plotly scatter chart of min_int predictions over the last 6 hours, colored by line, hover on destination
  • Recent capture runs table for at-a-glance pipeline health
  • Empty-state messaging that tells the user exactly what command to run if no DB exists yet

Capture-rate budget Default config (3 stations × 120s interval × 6 calls per run) = ~4,320 calls/day, well under the WMATA free-tier 50,000 calls/day limit. See docs/PHASE3_PLAN.md for the math and tuning advice.

Running it for real

# In a tmux/screen session, or as a launchd/cron job:
python scripts/capture.py --interval 120 --stations A01,B01,C04,A03

# Verify the DB is growing:
.venv/bin/python -c "from wmata.storage import connect, get_capture_stats; \
  import json; print(json.dumps(get_capture_stats(connect()), default=str, indent=2))"

Students are encouraged to pick a phase, fork, and extend.


Suggested Next Features

The roadmap is a guide, not a prescription. Pick anything below — they're sized for one to two work sessions each.

Phase 4 — Analytics on the History DB (natural next step)

The capture pipeline already populates the schema; these queries write themselves on top of it.

  • On-time performance per line / station / hour. For each numeric prediction sample at a station, the change in min_int between captures tells you whether the next train arrived faster or slower than predicted. Group by line and hour-of-day to surface patterns.
  • Headway distribution. From predictions filtered to a single line/destination, derive inter-train arrival gaps (MIN(min_int) per group) over rolling windows. Plot as a violin or box chart per hour-of-day.
  • Alert duration analysis. SELECT incident_id, MIN(captured_at), MAX(captured_at), MAX(captured_at) - MIN(captured_at) AS duration FROM rail_incidents GROUP BY incident_id gives you alert lifespans. Plot the distribution and find the worst-affected lines.
  • Elevator MTTR. Same idea for elevator_outages keyed on (station_code, unit_name). Surface the longest-running outages and the stations with the worst record.

Phase 5 — New data sources

  • GTFS-RT vehicle positions. WMATA publishes a Protocol Buffer feed of every train's lat/lon. Adds map visualization.
  • GTFS schedules (static). The full timetable as zipped CSVs. Combined with predictions, lets you compute "scheduled vs. actual" — the gold-standard on-time metric.
  • Equity overlay. Join census-tract demographics to station catchment areas and surface ridership/service gaps.

Smaller cross-cutting improvements

  • Map view of nearby bus stops. Folium or pydeck pinning stops within 500m of the selected station.
  • Station search box. With 102 stations, a typeahead is faster than the dropdown for power users.
  • Loading skeletons. Replace the generic Streamlit spinner with content-shaped placeholders (would push UI/UX rubric scores from 4 → 5 on dimension #7).
  • CLI (Phase 2 — still open). Typer + Rich, with wmata rail predictions A01, wmata status, wmata capture --start.
  • Cron / launchd recipe for the capture script so students can leave it running for a semester.

Patterns to follow

Build it in R instead?

Some students will prefer R over Python for the analytics work that makes this dashboard interesting. docs/R_VERSION_PLAN.md lays out a complete porting plan: Shiny + golem + httr2 + duckdb + renv, with a file-by-file mapping from the Python source, sketch implementations of the API client and the Rail tab as a Shiny module, and an estimated 20-hour effort budget for a student working with an AI coding assistant. Quarto Dashboards are also covered as a simpler-but-less-interactive Path B.

The same DuckDB file works from both languages — you can run the Python capture loop and read it from your R Shiny app simultaneously.


License & Credits

Code: MIT License (see LICENSE if present, or add one for your fork). Data: Provided by WMATA under their Transit Data Terms of Use. This is an unofficial application; WMATA name, logo, and branding are the property of WMATA.

Built as a teaching example for the AU Kogod School of Business MS Analytics program.

About

Starter kit for the AU Kogod MS Analytics program — a Streamlit dashboard, DuckDB capture pipeline, and design rubric for the WMATA DC Metro public API. Demonstrates modern agentic coding workflows with Claude Code, Codex, GitHub Copilot, and Perplexity Computer.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages