Releases: TIGER-AI-Lab/ClawBench
Release list
v0.9.2
What's Changed
- feat(harness): add WebBrain extension agent by @alectimison-maker in #273
- docs: update docs after WebBrain by @Perry2004 in #287
- docs: slim the README, fix broken links and corpus counts by @reacher-z in #283
- fix(judge): {"match": "false"} is scored as a PASS — parse verdicts as tri-state (closes #295) by @reacher-z in #305
- ci: fail static-check on broken markdown links (fixes a broken README image + 2 dead script refs) by @reacher-z in #291
- docs: surface Awesome Works above the fold (it sits at 85% depth, below the FAQ) by @reacher-z in #289
- build: v0.9.2 release by @Perry2004 in #306
New Contributors
- @alectimison-maker made their first contribution in #273
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.9.1
What's Changed
- Fixes an
x11vncstartup crash when runningclawbench-runwith--human. by @ramanbansal1 in #279 - build: v0.9.1 release by @Perry2004 in #280
New Contributors
- @ramanbansal1 made their first contribution in #279
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.9.0
What's Changed
- docs: collapse readme curated list section by @Perry2004 in #274
- docs(README): COLM 2026 WAB acceptance news + canonical TIGER-AI-Lab links by @reacher-z in #275
- docs: update changelog by @Perry2004 in #276
- Feat/104 support browserbase by @Perry2004 in #277
- build: v0.9.0 release by @Perry2004 in #278
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.8.0
What's Changed
- Feat/support remote browser by @Perry2004 in #235
- docs(models): ship default judge in models.example.yaml (fixes #240) by @reacher-z in #246
- Feat/248 try api before task by @Perry2004 in #249
- feat(prorl): run ClawBench as a ProRL-Agent-Server (Polar) RL environment (closes #219) by @reacher-z in #251
- fix(prorl): green up CI after #251 (pyright typing + Windows bash skip) by @reacher-z in #252
- feat(judge): support Gemini (google-generative-ai) as an LLM judge (closes #247) by @reacher-z in #254
- fix(security): reject path-traversal in extra_info paths by @reacher-z in #238
- docs(README): add one-command corpus run to Human Quick Start by @reacher-z in #260
- docs: update changelog by @Perry2004 in #261
- Track benchmark outreach and improve citation discoverability by @reacher-z in #267
- Show verified merged ClawBench featured lists by @reacher-z in #268
- docs: refresh citation metadata for v0.7.0 by @reacher-z in #266
- Update verified Featured in lists by @reacher-z in #263
- Add Awesome Works using ClawBench by @reacher-z in #262
- feat(judge): answer/rubric judge — primitive for importing rubric benchmarks (re #170/#188/#189/#190) by @reacher-z in #259
- feat(harness): null + random-click baselines — the benchmark floor (closes #245) by @reacher-z in #257
- docs: changelog by @Perry2004 in #270
- feat(edgebench): package ClawBench V2 as an EdgeBench/SForge benchmark by @reacher-z in #253
- build: v0.8.0 release by @Perry2004 in #271
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.7.0
What's Changed
- Refactor/replace extension by cdp by @Perry2004 in #227
- docs: add Hugging Face Space metadata checklist by @beanscg in #226
- chore: devcontainer-lock.json by @Perry2004 in #228
- revert: remove harbor as a harness by @Perry2004 in #229
- Feat/223 harbor compatibility by @Perry2004 in #232
- Clarify current leaderboard CSV and provenance fields by @koriyoshi2041 in #231
- build: 0.7.0 release by @Perry2004 in #233
New Contributors
- @beanscg made their first contribution in #226
- @koriyoshi2041 made their first contribution in #231
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.6.0
What's Changed
- feat(harness): add Harbor (Terminus 2) agent harness by @reacher-z in #220
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.5.0
What's Changed
- ci: add 3 sec delay before download from testpypi for it to refresh by @Perry2004 in #207
- feat: display token usage in USD by @Perry2004 in #208
- refactor: structural multi-harness by @Perry2004 in #210
- fix: make token cost calculation available to all harnesses #211 by @Perry2004 in #212
- fix: tui highlight not moving with selection #213 by @Perry2004 in #214
- build: v0.5.0 minor release by @Perry2004 in #215
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.4.1
What's Changed
- Fix/201 task mismatch by @Perry2004 in #205
- build: v0.4.1 release by @Perry2004 in #206
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
v0.4.0
What's Changed
- docs: sync BibTeX key + primaryClass with arXiv listing (READMEs) by @reacher-z in #195
- chore: corpus-aware issue + PR templates (default v2) by @reacher-z in #198
- docs: V1 → V2 comparison table (7 axes, sourced from 6 sub-agent audits) by @reacher-z in #197
- feat: inline LLM judge + V2 leaderboard 6-tab toggle + reproduce-leaderboard recipe by @reacher-z in #196
- Cleanup repo root #202 by @Perry2004 in #203
- build: v0.4.0 release by @Perry2004 in #204
Full Changelog: https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md
ClawBench v0.1.0
ClawBench v1.0.0
Initial public release of ClawBench -- a benchmark for evaluating AI agents on 153 everyday online tasks across 144 live production websites.
Highlights
- 153 tasks spanning 15 life categories (daily life, travel, education, job search, etc.)
- 144 live websites -- real production sites, not sandboxed clones
- Isolated Docker containers with Chromium for each run
- Request interceptor that blocks the final irreversible action to prevent real-world side effects
- Five-layer recording: MP4 replay, screenshots, HTTP traffic, DOM actions, agent messages
- Interactive TUI for model selection, test case picking, and run management
- 6 frontier models evaluated: Claude Sonnet 4.6 (33.3%), GLM-5 (24.2%), Gemini 3 Flash (19.0%), Claude Haiku 4.5 (18.3%), GPT-5.4 (6.5%), Gemini 3.1 Flash Lite (3.3%)
Links
- Paper: https://arxiv.org/abs/2604.08523
- Project Page: https://claw-bench.com