Skip to content

Commit 0d9aa3a

Browse files
Your Nameclaude
andcommitted
feat(bench): B15 -- cross-language competitor A/B (calm vs CodeGraph vs Ctxo vs Context+)
Extends B13 (real CALM-vs-CodeGraph A/B, only 2-3 corpora) to all 6 of CALM's Tier-0 languages, reusing B12's corpus registry verbatim (fd/Rust, flask/Python, gin/Go, express/JS, zod/TS, spring-petclinic/Java), and adds 2 new real competitors -- both schema-verified against a live spawned MCP server in this repo's own sandbox before being wired in, not read off their READMEs: - Ctxo (github.com/alperhankendi/Ctxo, MIT, npm) -- has its own PreToolUse safety gate, directly contesting CALM's "only tool with a pre-edit gate" framing in docs/comparison.md. - Context+ (github.com/ForLoopCodes/contextplus, MIT, npx, 1971 stars) -- has its own memory/RAG tools, directly contesting CALM's "only tool with cross-session memory" framing, and opens its own README with an unqualified "99% accuracy" claim with zero methodology. Task: file-recall on "who calls this symbol", against B12's independent git-grep oracle (also fixed in this session: symbol-collision exclusion in B2's SCIP oracle, string-literal false positives, and the Java package-private-method gap this same benchmark's own dry run surfaced). Aggregate result (N=8 symbols/corpus, single pass, calm@822e238): calm 68/72 (94.4%), CodeGraph 69/72 (95.8%), Context+ 72/72 (100%), Ctxo 23/24 (96%, java only -- see README for a fully disclosed, 3-round investigation into why go/js/ts are excluded: setup verified genuinely complete but query results structurally empty, root-caused to a real Ctxo-side persistence issue via direct SQLite inspection, not a harness bug). Also fixes a real, live-verified Context+ finding: its get_blast_radius does plain substring text matching, not symbol-aware resolution (caught it matching a Java `print` method against unrelated CSS rules like `print-color-adjust`) -- doesn't cost it recall in this benchmark's scoring, but is a real precision problem worth weighing against its own "99% accuracy" claim. Registers B13/B14/B15 in the top-level benchmarks/README.md table (B13/B14 were previously implemented but never added) and adds a claims.registry.jsonl entry for B15. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
1 parent 297af80 commit 0d9aa3a

4 files changed

Lines changed: 771 additions & 1 deletion

File tree

benchmarks/README.md

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -25,8 +25,11 @@ nhanh, không thay thế nó.
2525
| B10 | Real Competitor A/B | `calm` vs CodeGraph vs Semble — tool call thật trên cả 3 MCP server thật (không phải số tự báo cáo) | **Superseded by B11**[`b10_real_competitor_ab/`](b10_real_competitor_ab/) (giữ lại, xem B11 cho methodology đã fix) |
2626
| B11 | Extended Real Competitor A/B | `calm` vs CodeGraph vs Semble vs grepai vs Serena — sửa các lỗ hổng methodology của B10 (oracle đúng-sai cho mọi task, N=5 thay vì N=1, thêm task risk_gate_refusal + memory_recall test thật tính năng khác biệt của `calm`) | **Implemented**[`b11_extended_competitor_ab/`](b11_extended_competitor_ab/) |
2727
| B12 | Tier-1/Tier-2 Tool-Surface Correctness | Lái thật 9 MCP tool (repo_overview/search/source/file_overview/callers/edit_context/edit_lines/edit_symbol/diff_impact/hotspots) qua JSON-RPC trên 6 repo OSS ngoài (Tier-0: Python/Rust/Go/JS/TS/Java), full-power build, ground truth độc lập (regex + git grep, không phải call-graph precision) — khác trục đo với B2/resolution | **Implemented**[`b12_tier1_tier2_tool_correctness/`](b12_tier1_tier2_tool_correctness/) |
28+
| B13 | CALM vs CodeGraph Multi-Repo A/B | Mở rộng B12's corpus registry thêm 1 competitor thật (CodeGraph) + task freshness-under-live-edit; canonical 100% (31/31) vs 87.1% (27/31) file-recall trên callers | **Implemented**[`b13_codegraph_multirepo_ab/`](b13_codegraph_multirepo_ab/) |
29+
| B14 | Risk Calibration | `calm guard`'s `aggregate_risk` có track đúng commit gây bug thật không (SZZ-lite trên chính git history của CALM) — trục đo khác hẳn correctness, đo calibration | **Implemented**[`b14_risk_calibration/`](b14_risk_calibration/) |
30+
| B15 | Cross-Language Competitor A/B | `calm` vs CodeGraph vs Ctxo vs Context+ trên cả 6 ngôn ngữ Tier-0 (không chỉ self-repo/Rust như B11) — 2 competitor mới, cả hai đều thách thức trực tiếp claim "duy nhất" của `calm` (Ctxo có pre-edit safety gate, Context+ có memory/RAG) | **Implemented**[`b15_cross_lang_competitor_ab/`](b15_cross_lang_competitor_ab/) |
2831

29-
Ngoài chuỗi B1-B11 (đo lợi thế `calm` so với naive/competitor), còn một track riêng đo **chất lượng
32+
Ngoài chuỗi B1-B15 (đo lợi thế `calm` so với naive/competitor), còn một track riêng đo **chất lượng
3033
resolution đa ngôn ngữ** cho kế hoạch 8-ngôn-ngữ Formal-tier
3134
(`docs/superskills/plans/2026-07-07-eight-lang-formal-tier.md`) — không thuộc số B, vì trục đo khác
3235
hẳn (độ rộng/độ chính xác hỗ trợ ngôn ngữ, không phải calm-vs-naive): **Resolution** — tier
Lines changed: 204 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,204 @@
1+
# B15 — Cross-Language Competitor A/B (`calm` vs CodeGraph vs Ctxo vs Context+)
2+
3+
Extends [B13](../b13_codegraph_multirepo_ab/README.md) (real CALM-vs-CodeGraph A/B, but only
4+
2-3 corpora) to **all 6 of CALM's Tier-0 languages** — reusing
5+
[B12](../b12_tier1_tier2_tool_correctness/README.md)'s corpus registry verbatim (fd/Rust,
6+
flask/Python, gin/Go, express/JS, zod/TypeScript, spring-petclinic/Java) — and adds **2 new real
7+
competitors**, both live-verified against a real spawned MCP server in this repo's own sandbox
8+
before being wired in (schemas, tool names, and setup quirks below were discovered by calling the
9+
tools, not by reading their READMEs):
10+
11+
- **[Ctxo](https://github.com/alperhankendi/Ctxo)** (MIT, npm) — has its own PreToolUse safety
12+
gate (`ctxo gate`), directly contesting CALM's "only tool with a pre-edit gate" framing.
13+
- **[Context+](https://github.com/ForLoopCodes/contextplus)** (MIT, npx, 1971★) — has its own
14+
memory/RAG tools, directly contesting CALM's "only tool with cross-session memory" framing, and
15+
opens its own README with an unqualified **"99% accuracy"** claim with zero methodology — exactly
16+
the marketing anti-pattern the "Nghiên cứu competitor" section of the top-level
17+
[`benchmarks/README.md`](../README.md) already warns against repeating.
18+
19+
CodeGraph's adapter (`codegraph_callers`, version pin, `CODEGRAPH_MCP_TOOLS` env) is reused
20+
verbatim from B13 — unchanged, still on npm latest `1.5.0` as of this run.
21+
22+
## Task measured
23+
24+
File-recall on "who calls this symbol" — the same shape B13 already used, now driven through 4
25+
different real tool schemas:
26+
27+
| Tool | Call | Notes |
28+
|---|---|---|
29+
| `calm` | `callers(symbol, path)` | structured JSON; file lives inside a qualified `symbol` string, not a separate field (B13 already found this the hard way) |
30+
| CodeGraph | `codegraph_callers(symbol, file)` | free-text response, paths extracted via regex |
31+
| Ctxo | `search_symbols(pattern)``find_importers(symbolId, edgeKinds=["calls"])` | **two** real MCP calls — `find_importers` needs a `symbolId` (`file::name::kind`), not a bare name |
32+
| Context+ | `get_blast_radius(symbol_name, file_context)` | free-text response (`" path.ext:\n L<n>: ..."`), paths extracted via regex |
33+
34+
Oracle: B12's `ground_truth.py`, reused verbatim — word-bounded `git grep` + independent
35+
definition-regex extraction, **including 3 real oracle bugs found and fixed while building this
36+
exact benchmark** (see "Bugs found building this" below) on top of the 2 found auditing CALM on
37+
2026-08-18 (symbol-collision in B2's SCIP oracle, string-literal false positives in this same
38+
`git_grep_call_sites`).
39+
40+
## Scope limits — read before citing any number from this benchmark
41+
42+
- **Ctxo has no plugin for python or rust.** Verified live, not assumed from its docs: `ctxo
43+
install python` and `ctxo install rust` both **404 on the real npm registry**
44+
(`@ctxo/lang-python`, `@ctxo/lang-rust` don't exist). Silently skipped on those 2 languages —
45+
an absent row is not the same claim as a losing row, and is recorded as
46+
`arms_skipped_unsupported` in `results.json`, not hidden.
47+
- **Every tool answers a slightly different question under the shared "file-recall" label.**
48+
CodeGraph free-texts an impact summary; Ctxo's `find_importers` is edge-typed (asked for
49+
`edgeKinds: ["calls"]` specifically, but its underlying resolver may still have its own notion
50+
of what counts); Context+'s `get_blast_radius` explicitly documents itself as "usages," not
51+
"calls" — and (see below) was caught doing plain substring text search, not symbol-aware
52+
matching, on at least one real query. Read the raw per-symbol rows in `results.json` before
53+
treating any aggregate percentage as a clean ranking — same caveat B11's own README gives for
54+
its raw token-ratio numbers.
55+
- **Single pass** (`--n-repeats 1` by default) per symbol per corpus in this run. `--n-repeats 3`
56+
is available and recommended before treating any one row as final, matching B13's discipline.
57+
58+
## Bugs found building this (verified live, not asserted)
59+
60+
Three real, previously-undocumented ground-truth bugs were found and fixed *while building this
61+
benchmark*, before any published number — same "audit the oracle before trusting a miss" discipline
62+
as the 2026-08-02 and 2026-08-18 fixes to this same file:
63+
64+
1. **Java method-definition pattern required an explicit access modifier.** Package-private
65+
methods — the standard JUnit 5 convention for test methods (`void testFoo() { ... }`, no
66+
`public`/`private`/`protected`) — were never recognized as *definitions*, so a test method named
67+
after the production method it exercises (e.g. `void initUpdateOwnerForm() throws Exception`
68+
testing `OwnerController.initUpdateOwnerForm()`) got miscounted as a real *call site* of the
69+
production method. Verified live on spring-petclinic: CALM, CodeGraph, **and** Ctxo all scored
70+
0/1 "missing" a call that was never real — the oracle's sole "hit" was the test method's own
71+
declaration line. Fixed: the modifier group is now optional, same as the pre-existing
72+
class/interface patterns in the same file already treat it.
73+
2. **Ctxo's plugin installer hard-requires a `package.json`, even for non-npm languages.**
74+
Verified live on spring-petclinic (pure Maven, zero npm anywhere in the repo): `ctxo install
75+
java` refused outright with "No package.json in the current project" — its plugin-install
76+
mechanism always shells out through npm, unconditionally, regardless of which language plugin
77+
is being installed. Worked around with a throwaway `package.json` written before setup and
78+
removed immediately after (exactly what a real user hitting this on a Java-only repo would do).
79+
3. **`sample_symbols` returned 0 candidates on one otherwise-healthy run**, never reliably
80+
reproduced (the same corpus produced 8 real candidates seconds later when queried by hand) —
81+
not disk pressure (21GB free at the time), smells like a transient subprocess/IO hiccup. Added
82+
a one-shot retry rather than silently reporting a corpus as sample-less.
83+
84+
## Real finding, not a harness bug: Context+'s `get_blast_radius` false-positived on plain substring text
85+
86+
Live-reproduced, not inferred: querying `get_blast_radius(symbol_name="print", file_context=
87+
".../PetTypeFormatter.java")` on spring-petclinic returned **16 usages in 3 files**, one of which
88+
was `src/main/resources/static/resources/css/petclinic.css` — matching literal CSS text like
89+
`print-color-adjust: exact;` and `.d-print-inline-block`. This is plain substring text matching
90+
conflating a Java method name with unrelated CSS class names that happen to contain the same
91+
letters, not symbol-aware call-graph analysis. It doesn't cost Context+ any *recall* in this
92+
benchmark's scoring (the real oracle file was still present in its returned set, so it still scores
93+
a hit) — but it's a real, disclosable precision problem, worth weighing against the "99% accuracy"
94+
claim on Context+'s own README, and worth reading the raw `contextplus_files` column for, not just
95+
the recall fraction.
96+
97+
## Results (2026-08-18, `calm` @ `822e238`, N=8 symbols/corpus, single pass)
98+
99+
File-recall on "who calls this symbol", per language (hit/total oracle files):
100+
101+
| lang | calm | CodeGraph | Ctxo | Context+ |
102+
|---|---:|---:|---:|---:|
103+
| python | 9/9 (100%) | 9/9 (100%) | *(no plugin — see below)* | 9/9 (100%) |
104+
| rust | 9/9 (100%) | 9/9 (100%) | *(no plugin — see below)* | 9/9 (100%) |
105+
| go | 11/11 (100%) | 10/11 (91%) | **1/11 (9%) — see disclosure below, not a capability claim** | 11/11 (100%) |
106+
| javascript | 7/8 (88%) | 8/8 (100%) | **0/8 (0%) — see disclosure below, not a capability claim** | 8/8 (100%) |
107+
| typescript | 11/11 (100%) | 9/11 (82%) | **1/11 (9%) — see disclosure below, not a capability claim** | 11/11 (100%) |
108+
| java | 21/24 (88%) | 24/24 (100%) | 23/24 (96%) | 24/24 (100%) |
109+
| **aggregate (java only for Ctxo — see below)** | **68/72 (94.4%)** | **69/72 (95.8%)** | **23/24 (96%)** | **72/72 (100%)** |
110+
111+
**Read the "Ctxo go/js/ts: a real, unresolved integration-reliability finding" section below before
112+
citing the go/js/ts Ctxo numbers for anything** — they are published for transparency (raw data in
113+
`results.json`), but this benchmark's own investigation could not confirm they measure Ctxo's real
114+
call-graph quality, so the aggregate row above excludes them and counts only Ctxo's java result
115+
(the one arm verified end-to-end with real, non-empty query results).
116+
117+
**Reading the rest of the table**: CodeGraph and Context+ both land at or near 100% on every
118+
language they run on — CodeGraph misses 3 files total (go/1, typescript/2, both same-file or
119+
private-symbol edge cases, not investigated further here), Context+ misses none in this sample
120+
(see the CSS false-positive section below for why "0 misses" doesn't mean "flawless"). `calm`'s 4
121+
misses were spot-checked, not just counted: the javascript one (`User`, `examples/view-locals/
122+
user.js`) is a real, disclosable gap — the symbol is invoked exclusively via `new User(...)`, and
123+
CALM's JS/TS call extractor doesn't currently treat a `new`-expression as a call site for the
124+
constructed class name. The java ones are inheritance/cross-file dispatch cases (`getName`/`isNew`
125+
defined in a base class, invoked through a subclass instance) — a much harder class of problem for
126+
any syntactic (non-type-checking) resolver, consistent with this suite's own prior findings on
127+
Rust `Self::`/method-name collisions.
128+
129+
## Ctxo go/js/ts: a real, unresolved integration-reliability finding
130+
131+
Not a harness bug (ruled out through 3 independent rounds of increasingly rigorous verification,
132+
each one changing the setup code and re-running) and not (as far as this investigation could tell)
133+
an oracle bug either. Documented in full because burying an inconvenient result is exactly what
134+
this benchmark suite exists to not do:
135+
136+
1. **Round 1**: raw numbers were go 1/11, javascript 0/8, typescript 1/11, java 23/24. Hypothesis:
137+
`npm install -D @ctxo/lang-<x>` silently failing to materialize `node_modules` on larger
138+
real-dependency-tree corpora. Fixed (verify + retry) — numbers **did not change**.
139+
2. **Round 2**: direct SQLite inspection of `.ctxo/.cache/symbols.db` after a run that printed
140+
"[ctxo] Building codebase index... Found 141 source files" / "Index complete: 141 files indexed"
141+
found **all 3 tables (`files`, `symbols`, `edges`) empty (0 rows)** — the CLI's own stated
142+
progress does not match what it actually persisted. Also found the CLI prints its real progress
143+
to **stderr, not stdout** (a genuine bug in the first fix's own verification logic, which only
144+
checked stdout). Fixed (read both streams, require a `package.json` marker file specifically —
145+
not just the containing directory — before treating the plugin install as real, retry up to 3x).
146+
Re-ran the full 6-language sweep: numbers **still did not change** — go 1/11, javascript 0/8,
147+
typescript 1/11, java 23/24, byte-for-byte identical to before the fix, despite setup metadata
148+
now unambiguously showing a real, complete, verified index build (`plugin_materialized: true`,
149+
`indexed_ok_marker: true`, real per-language file counts in the captured output).
150+
3. **What this rules in/out**: the failure is reproducible, stable across 2 independent full runs
151+
with materially different (and progressively more careful) setup code, and specific to the
152+
`@ctxo/lang-typescript` (covers both javascript and typescript) and `@ctxo/lang-go` plugins —
153+
`@ctxo/lang-java` consistently works (23/24, and CALM/CodeGraph/Context+ all score normally on
154+
the exact same go/js/ts corpora in the exact same run, which rules out a corpus-level or
155+
MCP-transport-level problem). The pattern (small throwaway-package.json corpus works, larger
156+
real-dependency corpus doesn't) is consistent with an async persistence race in Ctxo's own
157+
indexer specific to those 2 plugins, but this investigation could not pin the exact mechanism
158+
within reasonable scope, and does not have access to Ctxo's own source to confirm.
159+
4. **Why this isn't scored as "Ctxo fails on go/js/ts"**: a 0-9% recall number here would be
160+
measuring "did this specific CLI+MCP-server pipeline reliably persist an index in this sandbox,
161+
for these 2 plugins, on this run" — not "how good is Ctxo's call-graph resolution once it has a
162+
working index" (java's 23/24 answers that question much better). Publishing the low number as a
163+
capability claim would be exactly the kind of misleading, oracle-unaudited result this whole
164+
benchmark suite's own house rules (see `benchmarks/README.md`'s "Nghiên cứu competitor" section
165+
on the Semgrep 250%-vs-50-71% lesson) argue against repeating in the other direction.
166+
167+
## Run
168+
169+
```bash
170+
cargo build --release -p calm-cli # default features already full power, nothing extra needed
171+
benchmarks/.venv/bin/python benchmarks/b15_cross_lang_competitor_ab/run_benchmark.py \
172+
--langs python,rust,go,javascript,typescript,java \
173+
--arms codegraph,ctxo,contextplus \
174+
--n-repeats 1
175+
```
176+
177+
`--langs`/`--arms` accept comma-separated subsets for a faster partial run (e.g. `--langs java
178+
--arms contextplus` for a single-corpus dry run). `.work/<lang>` corpora are thrown away after each
179+
language's pass unless `--keep-corpus` is set; `results.json` is written incrementally, one
180+
language at a time, so a crash partway through still leaves every completed language's data intact.
181+
182+
## Version pins
183+
184+
- CALM: whatever `--calm-bin` points at (default `target/release/calm`) — `results.json`'s
185+
`meta.calm_git_sha` records the exact commit independently of `calm --version` (which only prints
186+
the Cargo.toml package version, ambiguous across unreleased commits — B13 already learned this
187+
the hard way).
188+
- CodeGraph: `@colbymchenry/codegraph@1.5.0`, pinned explicitly in every spawn (not a bare package
189+
name — B13 found the bare form can resolve to a stale npx cache).
190+
- Ctxo: `@ctxo/cli@0.11.4` (latest on npm as of 2026-08-18).
191+
- Context+: `contextplus@1.0.8` (latest on npm as of 2026-08-18).
192+
193+
Exact pins for the run in "Results" above (`calm_worktree_dirty_at_run: true` — this benchmark's
194+
own files were themselves uncommitted working-tree changes at run time, not core engine code):
195+
196+
| | |
197+
|---|---|
198+
| calm | `822e238efc54df32da36505cf25890c2302ee06f` |
199+
| python (flask) | `36e4a824f340fdee7ed50937ba8e7f6bc7d17f81` |
200+
| rust (fd) | `41532d114e2ba565fb5367d606c111b29b96450c` |
201+
| go (gin) | `34dac209ffb6ef85cc78c5d217bbb7ad001d68fd` |
202+
| javascript (express) | `a3714473feb3d2908add734d340e7755fd85e0a3` |
203+
| typescript (zod) | `912f0f51b0ced654d0069741e7160834dca742ee` |
204+
| java (spring-petclinic) | `51045d1648dad955df586150c1a1a6e22ef400c2` |

0 commit comments

Comments
 (0)