diff --git a/docs/ingestion.md b/docs/ingestion.md index b1575f6..afaae5d 100644 --- a/docs/ingestion.md +++ b/docs/ingestion.md @@ -79,9 +79,8 @@ units, multivalued, required, pattern}}}`. Converts the dict into real, hash-identified objects: - **Properties are built first.** Each one's `hash_id` is computed from - `name`/`description`/`range`/`units` only — `slot_uri` and - `registry_version` are stored on the object but excluded from the hash - (they're provenance/operational metadata, not identity). + `name`/`description`/`range`/`units` only — `slot_uri` is stored on the + object but excluded from the hash (it's origin metadata, not identity). - **Classes reference their properties by hash_id**, sorted. This is why a class's own hash depends on its full induced property set — two classes with the same properties (regardless of declaration order) hash the same. @@ -91,8 +90,8 @@ Converts the dict into real, hash-identified objects: have every `is_a` target resolvable within its own import closure (that's `SchemaView`'s job), so this recursion always terminates cleanly. - **Every entity gets one `ProvenanceEntry`** for this ingestion — `source`, - `attributed_to` (agent), `generated_at`, `activity` — entirely separate - from the hash computation. + `attributed_to` (agent), `generated_at`, `activity`, `registry_version` — + entirely separate from the hash computation. ### 3. `write_registry_entities()` + `write_structural_edges()` (`neuro_ghost/db.py`) @@ -116,9 +115,8 @@ and `seed.py` — schema.org is ingested through the exact same path, just with | `range`, `units` | `RegistryProperty` | Identity-defining. | | `properties`, `is_a`, `mixins` | `RegistryClass` | Identity-defining — all stored as hash_id references. | | `class_uri` / `slot_uri` | `RegistryClass` / `RegistryProperty` | Ontology IRI preserved from the source. **Not** part of the hash — two schemas using different IRIs (or none) for the same content still collapse to one node. | -| `registry_version` | both | Operational stamp. Not part of the hash. | | `provenance` | both | List of `ProvenanceEntry`. Accumulates, never affects `hash_id`. | -| `source`, `attributed_to`, `generated_at`, `activity`, `derived_from` | `ProvenanceEntry` | PROV-O–grounded (`slot_uri: prov:wasAttributedTo` etc.) | +| `source`, `attributed_to`, `generated_at`, `activity`, `derived_from`, `registry_version` | `ProvenanceEntry` | PROV-O–grounded (`slot_uri: prov:wasAttributedTo` etc.). `registry_version` lives here, not on the entity — the same entity can be attested by different sources at different times, each under a different registry version, so a single scalar on the entity doesn't fit (same reasoning as dropping `source_label`). | Field names were deliberately aligned with LinkML's own metamodel (`description`, `range`, `class_uri`/`slot_uri`, `is_a`, `abstract`) rather @@ -167,6 +165,33 @@ Saying "parse_linkml extracts X" and "the registry stores X" are different claims — test each layer for what it actually is, not for what you assume the other layer does with it. +`tests/test_ingest_registry.py::test_required_does_not_affect_property_identity` +takes this one step further, end to end: two schemas declare the exact same +`age` slot except one marks it `required: true`. Ingesting both must produce +exactly one `RegistryProperty` node, not two — proving `required` doesn't +leak into identity, in the real graph, not just in an isolated object. + +## Open question: when should ProvenanceEntry.registry_version be set? + +The whole-registry version (the semver in `data/registry.json`, e.g. `1.7.0`) +only gets computed once, at the very end of a submission, when +`export_json.py` reads the last entry in `data/provenance.json` and bumps it. +`seed.py`/`ingest_linkml.py`/`align.py` all run *before* that — so the version +an entity is actually going to belong to doesn't exist yet at the moment it's +ingested. + +That rules out the obvious fix ("pass the current, pre-bump version to +ingest_linkml.py/align.py") — it would record the version this submission is +replacing, not the one it's part of. The consistent fix would be computing +the bump once, up front (alongside where the Action already determines +`bump` type in "Parse issue metadata"), and threading that single value +through every step including `export_json.py` — instead of `export_json.py` +computing it independently after the fact. + +**Not implemented.** `schema_submission.yml` does not pass `--registry-version` +to `ingest_linkml.py` or `align.py` — those entities' `ProvenanceEntry. +registry_version` stays `None` for now, pending a decision on the above. + ## Known gaps (as of this writing) - **`index.html`** (frontend) still expects the pre-rework JSON shape diff --git a/neuro_ghost/export_json.py b/neuro_ghost/export_json.py index 48261cb..a539136 100644 --- a/neuro_ghost/export_json.py +++ b/neuro_ghost/export_json.py @@ -96,8 +96,8 @@ def export_snapshot(conn, registry_version: str) -> dict: "hash_id": hash_id, "iri": class_uri or "", "name": name or "", - "definition": desc or "", - "is_abstract": bool(is_abstract), + "definition": desc or "", + "is_abstract": bool(is_abstract), "sources": _attesting_sources(conn, "RegistryClass", "HAS_PROVENANCE", hash_id), "properties": [ {