Skip to content

Commit a3cd368

Browse files
vivekchandclaude
andcommitted
Add ClawMetry to Agent Evaluation & Observability
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rt5fq2BWxieLpR69MAbwjr
1 parent b458ef6 commit a3cd368

1 file changed

Lines changed: 1 addition & 0 deletions

File tree

‎README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1147,6 +1147,7 @@ Entries may carry one or more status tags so readers can judge maturity at a gla
11471147
- [Langfuse](https://github.com/langfuse/langfuse) - Open-source LLM engineering platform — traces, evals, prompt management. **Acquired by ClickHouse Jan 2026**; March 2026 shift to an observations-centric data model, April 2026 added Langfuse Cloud Japan + Experiments + Langfuse Academy + LLM-as-a-Judge API; v4 self-host release queued. ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2Flangfuse%2Flangfuse&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)
11481148
- [OpenLLMetry](https://github.com/traceloop/openllmetry) - Open-source observability for LLM applications based on OpenTelemetry. ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2Ftraceloop%2Fopenllmetry&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)
11491149
- [Weights & Biases Weave](https://github.com/wandb/weave) - Toolkit for developing, evaluating, and monitoring AI applications. ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2Fwandb%2Fweave&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)
1150+
- [ClawMetry](https://github.com/vivekchand/clawmetry) - Self-hosted observability and opt-in kill switch for coding agents (Claude Code, Codex, Cursor, Gemini CLI, OpenClaw and others) that reads the session logs runtimes already write on disk, so there is no SDK and nothing in the request path. [Website](https://clawmetry.com) ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2Fvivekchand%2Fclawmetry&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)
11501151
- [SWE-bench](https://github.com/SWE-bench/SWE-bench) - Benchmark for evaluating LLMs on real-world software engineering problems. ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2FSWE-bench%2FSWE-bench&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)
11511152
- [Terminal-Bench](https://www.tbench.ai/) - 🆕 Benchmark for terminal-based coding agent evaluation. Maintained by Harbor Framework. ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2Fharbor-framework%2Fterminal-bench&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)
11521153
- [Harbor](https://github.com/harbor-framework/harbor) - 🆕 Framework for evaluating and optimizing agents and LLMs at scale — run Terminal-Bench 2.x and custom benchmarks across thousands of cloud sandboxes, generate RL rollouts; a Stanford × Laude Institute collaboration. Apache-2.0. ![GitHub stars](https://img.shields.io/badge/dynamic/json?label=Stars&query=%24.stargazers_count&url=https%3A%2F%2Fapi.github.com%2Frepos%2Fharbor-framework%2Fharbor&color=yellow&logo=github&logoColor=white&style=flat&cacheSeconds=300)

0 commit comments

Comments
 (0)