Repository navigation
fix(bench): reject duplicate raw quads in UPDATE comparisons - #6487
Conversation
Keep complete-quad snapshot and probe comparisons strict when canonicalization would hide equal-count duplicate redistribution. Pin both-engine graph scope and preserve canonicalization, lexical adjudication, and count diagnostics. Co-Authored-By: GPT-6 Astra <noreply@openai.com>
There was a problem hiding this comment.
🟢 Approval recommended
The change is narrowly scoped to the bench comparator, is well-defended by targeted tests, and does not alter engine semantics or canonicalization behavior.
Pull request overview
Tightens the sparq-bench UPDATE differential comparator to prevent false Verdict::Same outcomes when raw snapshot/probe output redistributes identical duplicate N-Quads lines that get collapsed by canonicalization.
Changes:
- Added strict detection/rejection of repeated raw N-Quads lines on the equality fast-path (after canonical forms match and raw totals match).
- Extended module documentation to clarify the comparator’s “unique emission” contract is specific to complete-quad snapshots / fully-projected probes (not general SPARQL result bags).
- Added focused regression and contract tests covering duplicate redistribution, self-comparison invalidity, probe scope/projection, and preservation of canonicalization error behavior.
File summaries
| File | Description |
|---|---|
| crates/sparq-bench/src/update_fuzz.rs | Adds raw-duplicate rejection to the comparator’s Same fast-path and introduces regression/contract tests to prevent future false-equality. |
Review details
- Files reviewed: 1/1 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
sparq engine
Details
| Benchmark suite | Current: 1e800e1 | Previous: 674f50b | Ratio |
|---|---|---|---|
store_bytes_per_triple |
88 bytes |
88 bytes |
1 |
dict_bytes_per_term |
53 bytes |
53 bytes |
1 |
comp_store_bytes_per_triple |
44 bytes |
44 bytes |
1 |
store_bytes_per_triple_small |
88 bytes |
88 bytes |
1 |
wasm_bundle_bytes |
1562916 bytes |
1562916 bytes |
1 |
This comment was automatically generated by workflow using github-action-benchmark.
|
The Generated by Claude Code |
…or-raw-duplicates
🔎 Codex reviewer —
|
🔎 Codex reviewer —
|
The UPDATE differential comparator can report
Samewhen two outputs have equal raw row counts but redistribute repeated identical quads. Blank-node canonicalization collapses those repetitions before comparison. The strict equality path now rejects a repeated raw complete-quad line, reporting its side, original line and occurrence count while preserving the existing unequal-count diagnostic.This addresses #6483 within the private complete-quad snapshot and fully projected probe contract. It does not change engine results, canonicalization or the semantics of general SPARQL result bags. Duplicate-free blank-node relabelling and the same triple in distinct named graphs remain accepted. The borrowed set adds O(N log N) comparisons and O(N) references on the strict equality path; no measured performance improvement is claimed.
Validation on the exact committed source: all 26 actual module tests passed, including six new boundary/regression tests and both-engine default/named-graph controls. Removing only the new guard compiled successfully and made the redistribution regression fail by observing
Same. Scoped actual-module Clippy with warnings denied, touched formatting and diff checks passed. These native tests used pinned recorded dependencies; they do not replace the normal Linux workspace and feature gates. The first run was interrupted by a resource-monitor race during normal temporary-directory cleanup; the unchanged compiled candidate passed after a controller-only correction. Both runs' evidence is preserved. Local author preflight encountered the established Bash 3mapfilelimitation, so Linux privacy checks remain required.Implementation: GPT-6 Astra with extra-high reasoning. Actual independent Claude Opus 5 with extra-high reasoning approved commit
0b4554b924a80432cc1b572bd19f8e58cdbfb4e6for normal CI, with no blocking findings. The issue stays open until the fix lands and relevant post-merge evidence is verified.Benchmark
Local
sparq-cli bench, operators suite (2,000 entities, 16k triples), release-fast binaries of main4105e5489dand this PR. Best of 5 iterations per round, minimum over 5 interleaved rounds; row counts match on every query. Every query is within noise of main; the geomean ratio is 1.002.Generated by Claude Code