Repository navigation
perf(serialize): id-level graph_to_turtle writer (#4898) - #6658
Conversation
graph_to_turtle / graph_to_turtle_with decoded every triple into owned oxrdf terms, cloned subjects several times per row, rendered each predicate into a fresh String key, and rendered every IRI twice (header dry run + body). The new writer works on dictionary ids: IRI / blank-node renderings are cached per id, literals render from borrowed term_parts records, inline integers skip the Term rebuild, and the used-prefix set is collected while rendering. Output is byte-identical to write_turtle(&graph_triples(g), ..), pinned by a new test. Also: - escape_string / escape_iri copy unescaped runs in bulk. - write_prefix_header uses a compiled PrefixTable (no probe render / clone). - write_iri tie-break compared the best LOCAL length to the candidate NAMESPACE length; it now implements the documented longest-namespace rule. - Turtle (plain and pretty) dropped the RDF 1.2 base direction of directional language-tagged literals (@ar--rtl re-parsed as @ar). - New example turtle_vs_oxttl: same-box comparison vs oxttl's TurtleSerializer on a document-shaped graph. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
🔎 Codex reviewer —
|
Turtle's PN_LOCAL must start with PN_CHARS_U | ':' | [0-9] | PLX, so an unescaped leading '-' or '.' is invalid. is_simple_pn_local accepted both, and with longest-namespace selection overlapping prefixes (a -> http://ex/ns, z -> http://ex/) turned <http://ex/ns-foo> into the unparsable a:-foo. The helper now rejects those first characters, so candidates fall back to a shorter valid prefix (z:ns-foo) or the full IRI. Adds a parse round-trip regression test with overlapping namespaces over the id-level writer, the generic and pretty Turtle writers, and the compact and pretty TriG writers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
🔎 Codex reviewer —
|
1 similar comment
🔎 Codex reviewer —
|
…JSON-LD The TriG `@prefix` header (buffered, streaming and pretty) and the compacted JSON-LD `@context` were collected from the graphs' triples only, but the writers also prefix-compact each named graph's NAME. With a: -> http://ex/ and b: -> http://ex/ns, GRAPH <http://ex/nsLongEnoughLocalPart> rendered as `GRAPH b:LongEnoughLocalPart` under a header declaring only a:, which is invalid TriG (and a JSON-LD graph @id whose prefix is missing from the context re-expands to a different IRI). Every used-prefix set now comes from one walk, `note_dataset_iris`, over every IRI position a dataset writer may compact: subject / predicate / object IRIs (through triple terms, including literal datatypes) and the name of every emitted named graph. The longest-namespace rule is kept. The TriG headers no longer clone the union of all graphs' triples, and the pretty header uses the compiled PrefixTable instead of a probe render per IRI. Tests: trig_declares_prefixes_used_only_by_graph_names round-trips the reviewer's dataset through all six TriG entry points (parse back, dataset isomorphism); jsonld_context_declares_prefixes_used_only_by_graph_names. The turtle_vs_oxttl example now also times the TriG writers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
🔎 Codex reviewer —
|
|
Local ci-fast gate (GitHub Actions outage; Jesse approved local-gate merges). PR head
Squash-merging under the local-gate rule. Generated by Claude Code |
Requested by Jesse · project thread
Summary
Closes #4898:
graph_to_turtlewas slower than Oxigraph 0.5's Turtle writer on a document-shaped graph.Root cause.
graph_to_turtle/graph_to_turtle_withcalledwrite_turtle(&graph_triples(g), ..). That first decoded every triple into ownedoxrdfterms (threeStringallocations per row, plus the literal datatype). It then cloned each subject several times per row, rendered every predicate into a freshStringkey and hashed it, and rendered every IRI twice (a header dry run into a probe string, then the body). On a document graph (many subjects, few predicates, long text literals), this per-row bookkeeping was most of the cost.Fix (all in the feature-gated
sparq-engine-serializecrate, behindserialize-rdf):graph_to_turtleandgraph_to_turtle_with:Dict::term_partsrecord, with the datatype suffix cached.Termrebuild.u32ids. When the input is already grouped, asiter_ids(SPO) order is, no reordering happens.escape_string/escape_iri: copy unescaped runs in bulk instead of pushing perchar.write_prefix_header, which the generic and streaming paths use, now uses a compiledPrefixTable: no probe render, no subjectTermclone.turtle_vs_oxttl: a same-box comparison against oxttl 0.2'sTurtleSerializer(the writer behind Oxigraph 0.5'sRdfSerializer) on a deterministic document-shaped graph. Elements are headings, paragraphs and list items with varied-length text, order integers, links, and some@enlabels.Output.
cmpof before/after dumps, both default and custom prefixes).id_writer_matches_generic_writerpinsgraph_to_turtle_with == write_turtle(&graph_triples(g), ..)and a parse round-trip. It covers inline and big integers, plain, typed, custom-datatype, language and directional literals, escape-heavy strings, blank nodes, triple terms, non-simple local names, and nested and equal-length namespaces, across four prefix maps.turtle_row_order_groups_like_write_turtle_bodycovers the non-grouped ordering path.Two behaviour fixes the new test found. Both are in the shared generic path, so the old and new writers still agree:
write_iritie-break bug. The guard compared the best match's local-part length against the candidate's namespace length. So it neither kept the longest namespace nor broke ties by label order, both of which the doc comment promises. It now follows the documented rule. Output changes only when two registered namespaces both match the same IRI with a simple local name. The coverage test that pinned the old arithmetic was rewritten aswrite_iri_longest_namespace_wins."x"@ar--rtlas"x"@ar, so a re-parse lost the direction. It now emits--rtl/--ltr.Minor: a prefix label containing
:(invalid Turtle) used to be mis-detected by the header'ssplit_once(':')re-parse. The compiled table now marks the actual prefix used.Measurements (non-canonical work-box measurements)
Shared 4-core dev box running other builds concurrently. Each number is the minimum of 200 to 300 iterations. These are signals, not canonical numbers; the canonical/EC2 bar from epic #2600 has not been run.
Corpus:
turtle_vs_oxttl 1500= 1,500 document elements, 8,396 triples, customdoc:/ont:prefixes.graph_to_turtle_withTurtleSerializer, triples pre-decodedTurtleSerializer, including id → term decodeReproduce:
cargo run --release -p sparq-engine-serialize --features serialize-rdf --example turtle_vs_oxttl -- 1500 200.On this synthetic shape the old writer was about 2–3× slower than oxttl, a bigger gap than the ~32% the reporter measured on their real graph; the new writer is about 2.4× faster than oxttl. I did not have the reporter's ruddydoc fixture, so their graph is not measured here.
The streaming writer (
graph_to_turtle_streaming) still uses the generic path. Only its prefix header got faster. It stays byte-identical tograph_to_turtle, which the existingstreamed_equals_bufferedtest checks.Base gate (always required)
cargo build --workspacesucceeds. (Not run: crate-scoped only on this shared box; left to CI.)cargo clippy --workspace --exclude sparq-py --all-targets -- -D warningsis clean. (Rancargo clippy -p sparq-engine-serialize --all-targets --all-features -- -D warnings: clean. Full workspace left to CI.)cargo testpasses for every crate this PR touches. (cargo test -p sparq-engine-serialize --features serialize-rdf,streaming-serialization: 123 lib tests and all integration tests pass, after merging currentorigin/main.)Targeted re-evaluation (check the rows that apply to your change)
pubitem. I still updatedskills/data-formats/SKILL.mdto document the longest-namespace rule and the new example.serialize-rdfis opt-in), so the byte-identity gate is unaffected. The scripts were not run.Cargo.lockgains only edges to crates already in the lock:rustc-hash2.1.3 (optional, enabled only byserialize-rdf) andoxttl0.2.3 (dev-dependency). No new crate versions.cargo deny/cargo auditare not installed on this box, so they were not run.Ratchets and conventions
TODO/FIXMEmarkers.write_iricomment, SKILL.md).Security
unsafe(the crate keepsforbid(unsafe_code)). The datatype cache is keyed by the dictionary-owned&straddress and length while the graph is borrowed immutably.🤖 Generated with Claude Code
https://claude.ai/code/session_01ScyGGohDhirnbLbrUSnA5n
Generated by Claude Code