Repository navigation
fix: canon directional literals, terse K: boundary, mcp nl_query config, gpu OOB GROUP BY keys - #6659
Conversation
…ig, gpu OOB keys - sparq-canon (#5359): standard (rdf-canon-backed) paths fail closed with a typed CanonError::DirectionalLiteral on RDF 1.2 directional-language literals instead of a generic Bridge error from the oxttl-0.1 re-parse; documented in crate docs + README, pinned by tests/directional_literal.rs. - sparq-terse (#4662): the K:/V()/PREFIX scanners treat non-ASCII bytes and a preceding PN_LOCAL_ESC as name characters, so `?aéK:x` and `ex:a\-K:x` pass through unchanged. - sparq-mcp (#4833): nl_query and ask share one base NlqConfig; new run_nl_query_with takes an explicit config and applies the same check_dictionary repair branch Nlq::ask does. - sparq-gpu (#4603): out-of-range GROUP BY keys set a device flag and return GroupKeyOutOfRange (the CPU oracle returns the same error instead of panicking); upload_table rejects non-power-of-two / full tables in release; the probe walk is bounded at mask+1 steps (CPU mirror too); run() polls with a 60 s timeout instead of waiting indefinitely. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
🔎 Codex reviewer —
|
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
🔎 Codex reviewer —
|
…ed queue The #4603 bounded poll panicked on timeout. On an unwind the caller's Gpu is dropped, and wgpu-core 30's Drop for Queue calls the HAL's wait_for_idle() with no timeout, so it could block forever on the same stalled submission. Kernel calls now return Result<_, GpuError>. On a poll failure, run() marks the Gpu stalled and returns GpuError::Stalled. Later calls fail fast without touching the device, and Drop for Gpu forgets an extra clone of the Arc-backed wgpu Device and Queue, so their destructors (and the unbounded idle wait) never run. This stays safe code (forbid(unsafe_code)). group_aggregate's out-of-range key error becomes GpuError::GroupKeyOutOfRange, with From<GroupKeyOutOfRange>. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
🔎 Codex reviewer —
|
|
Local ci-fast gate (GitHub Actions outage; Jesse approved local-gate merges). PR head
Squash-merging under the local-gate rule. Generated by Claude Code |
Requested by Jesse · project thread
Summary
Four small, independent fixes, each with a regression test that failed before the change.
"hello"@en--ltr) reaching a standard rdf-canon-backed entry point used to fail as a genericCanonError::Bridge, because the oxttl-0.1 re-parse is RDF 1.1. Every standard path now checks for these literals up front and returns a typedCanonError::DirectionalLiteral. The enum is#[non_exhaustive], so adding the variant is non-breaking. The crate docs and README now say so, and the error message points to therdf12-triple-termsprofile, which canonicalizes these literals natively; a test asserts that. This PR doesn't conflict with fix(canon): label-independent tie-break for tied RDFC-1.0 N-degree hashes #6645, which only vendorsrdf-canonand adds a test file; it doesn't touch the bridge insrc/lib.rs.K:sigil scanner has the same ASCII-only name-boundary bug askeyword_hintsdid #4662). TheK:,V(andPREFIXtoken-boundary checks used an ASCII-only predicate, so?aéK:xwas misread as theK:sigil and a valid query was rejected withUnknownKeyword. The newpreceded_by_nametreats any non-ASCII byte (outside strings, IRIs and comments it can only be aPN_CHARSname character) and a precedingPN_LOCAL_ESC(\-, etc.) as part of a name. Tests cover the Unicode variable, theex:a\-K:xlocal-name case (also broken before) and theV(scanner.askandnl_querynow start from onebase_config(). A newrun_nl_query_with(graph, question, config, llm)applies the samecheck_dictionaryrepair branch thatNlq::askhas, so the two tools can't drift if that flag is ever turned on. The default behaviour doesn't change, andboth_tools_accept_an_ungrounded_predicatestill passes. The new test fails if the branch is disabled.skills/agent-tools/SKILL.mdis updated.>= groupsnow sets a flag word on the device, with no host scan, and bothGpu::group_aggregateand the CPU oraclecpu::group_aggregatereturnErr(GroupKeyOutOfRange). Before, the GPU silently dropped the row and the CPU panicked. Checked on lavapipe: with the old kernel behaviour the new differential test fails withOk([...9999 rows...])where the CPU returnsErr.upload_tablenow rejects tables that aren't a power of two or have noEMPTY_KEYslot, in release builds as well. Before, this was only adebug_assert!.mask + 1steps.run()now waits at mostPOLL_TIMEOUT(60 s) for its own submission and then panics, instead of usingwait_indefinitely().write_tabledeliberately skips the O(n) table check. It's the timed refill in the e2e benchmark legs, and the bounded walk already stops a full table from hanging.group_aggregatenow returns aResult. The callers in the example and tests, andskills/gpu-kernels/SKILL.md, are updated; the skill's sample also no longer uploads a full table.Closes #5359
Closes #4662
Closes #4833
Closes #4603
Base gate (always required)
cargo build --workspacesucceeds. (Left to CI; only the touched crates were built locally.)cargo clippy --workspace ...is clean. (Left to CI. Ran crate-scopedcargo clippy -p {sparq-canon,sparq-terse,sparq-gpu,sparq-mcp} --all-targets --all-features -- -D warnings: clean.)cargo fmt --all).cargo testpasses for every crate this PR touches:sparq-canon(default and--all-features),sparq-terse --all-features,sparq-gpu(run on a real adapter: Mesa lavapipe/llvmpipe Vulkan, so the GPU differential tests actually ran and weren't skipped),sparq-mcp --features nlq.Targeted re-evaluation
CanonError::DirectionalLiteral,sparq_mcp::nlq::run_nl_query_with,sparq_gpu::{GroupKeyOutOfRange, POLL_TIMEOUT}, and thegroup_aggregatesignatures (GPU and CPU) now returningResult. Updatedskills/agent-tools/SKILL.mdandskills/gpu-kernels/SKILL.md, plus the sparq-canon README and crate docs.Ratchets and conventions
Security
research/gpu-threat-model.md.Performance check (local, before vs after main)
Non-canonical, shared-box measurements: 4 cores, other jobs running, release profile. Each binary was built from
origin/mainand from this branch in separate target dirs. The runs alternate main/PR to cancel drift. Noise band is the p10–p90 spread of the per-run times, relative to the median.canonicalize_quads, 100k quads (20k blank nodes, 4 graphs, lang/typed literals); 15 runs, best of 3canonicalize_triples, default-graph subset (25k triples); 15 runs, best of 3terse_to_sparql, 6,000 queries (K: sigils, Unicode names); 25 runs × median of 7Verdict: no regression beyond noise. The up-front directional-literal scan is one pass over object terms and doesn't show up next to the RDFC-1.0 work. The terse boundary check costs O(1) per candidate sigil byte; best-of times are identical. sparq-gpu was not timed: lavapipe is too noisy.
🤖 Generated with Claude Code
https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
Generated by Claude Code