Skip to content

Commit 53bb7b9

Browse files
committed
Fix corpus recovery, indexing guards and GraphRAG retries
Keep config and mutation outcomes attached to the winning corpus registry load. Use the canonical embedding identity during reindex and disable both hidden KG retry layers. Clarify native gateway spend ranges and missing evidence. @codex review
1 parent 2a3555a commit 53bb7b9

14 files changed

Lines changed: 1876 additions & 124 deletions

docs/exec-plans/active/graphrag-continuation-2026-09-04.md

Lines changed: 140 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -831,3 +831,143 @@ shared-code scope: 70 files, 1,006 symbols and 95 affected flows. The CLI still
831831
caps its displayed names with a larger limit, while reporting that its counts and
832832
risk cover all changes; exact staged paths were checked separately. A fresh PR
833833
review and full CI run follow.
834+
835+
The amended PR90 head `9c3bb56b` passed full CI with 2,922 tests and 68
836+
environment-gated skips in 860.85 seconds. GitHub Codex review reported no major
837+
issues on that exact head. PR90 merged as `2a3555a7` on September 5 at 08:42 UTC;
838+
post-merge CI passed, including the Docker build and container test. Production
839+
activation completed at 09:05 UTC on that exact commit. All nine preexisting
840+
GitNexus tooling paths were preserved byte-for-byte. The native ledger migration
841+
verified 140 completed migrations; native gateway readiness and the deployed marker
842+
were verified after startup.
843+
844+
### Separate operator follow-ups
845+
846+
The Cost & Capacity dashboard follow-up uses UTC midnight for Today and retains
847+
explicit selected-range and seven-day meanings. Native model/lane counters use
848+
reset-aware queries and normalize absent lane labels to unattributed. The pinned
849+
collector experiment confirmed that an asynchronous OTLP export failure does not
850+
necessarily emit a callback-failure counter sample. Scoped terminal-error logs
851+
therefore remain separate evidence. Their zero is guarded by the presence of logs
852+
from the same gateway service; absent or foreign-only logs remain unavailable.
853+
Five real Loki cases, the real Prometheus reset/midnight matrix and nine dashboard
854+
tests pass. Sol approved the final presence correction. No production dashboard
855+
change has been made yet.
856+
857+
Shared config loads now pin corpus/global scope and reject older selection epochs,
858+
including A to B to A. The private Graph browser suite passed 14 cases after a
859+
test-observer correction: Vite's actual timestamped module URL must be reused to
860+
inspect the rendered app's store. Frontend lint, 21 unit tests and build passed in
861+
that Graph overlay. Sol then found a further upstream recurrence: an older corpus
862+
registry response can resurrect a deleted corpus after a newer forced refresh.
863+
The new held-response regression reproduced that recurrence against the prior
864+
store. Registry publication now fences success, failure, loading, URL/storage
865+
canonicalization and events by request generation. All 16 Graph browser cases
866+
pass, including old and new concurrent registry loads. All nine accounting UI
867+
cases also pass with these shared-store changes. Final frontend lint, 23 unit
868+
tests, build, nine dashboard tests with the real Prometheus query engine, Ruff,
869+
banned-pattern and generated-type checks pass. Sol approved the final registry
870+
correction; E5 and this config correction are frozen for their own publication.
871+
872+
The cloud-embedding transport continuation has begun as a separate source slice.
873+
It will route existing OpenAI embedding calls through the native gateway with
874+
explicit identities while preserving local embedding paths and dense contracts.
875+
It is excluded from the pending PR90 deployment and NASA rebuild.
876+
877+
878+
### September 5: live accounting and corpus recovery follow-through
879+
880+
The approved NASA schema attempt `e3e16387536941029e634ea4bddbc9da` completed
881+
at 09:09:14 UTC after 33 seconds. The browser recovered the cached schema after a
882+
usage interruption without generating another paid proposal. Its 26 node types,
883+
34 relationship types and 89 directed patterns cover alarms, programs, trajectories,
884+
anomalies, failure causes and corrective actions; schema hash is
885+
`faa5b42a4876bd8e2c8ba43751fe856d2e69c288e45301efd1d600378e16375c`.
886+
Native reconciliation matched one request and $0.0649835 provider-reported spend,
887+
with 10,602 input and 3,848 output tokens. The delivered Langfuse generation and
888+
historical Mimir schema-proposal counters agree with that usage and cost. Native
889+
ledger content is absent; Langfuse input/output contain redaction placeholders.
890+
The browser correctly retains the unverified gateway-attempt-policy qualifier.
891+
892+
The ordinary NASA rebuild attempt `f899295ba26d4b179d64ef714378bf69` refused
893+
at 15:10:47 UTC before any extraction dispatch: the embedding guard reported stored
894+
deterministic versus configured provider. The active generation remains unchanged.
895+
The authoritative Postgres column is provider; the guard incorrectly read the absent
896+
nested metadata key as deterministic. NASA, Epstein and code share that promotion
897+
metadata shape. The guard correction reads only the canonical column and refuses
898+
unknown identities. Its 10-case matrix reproduced nine failures on the old guard;
899+
all 13 relevant promoted-lane cases pass with the correction, and Sol approved it.
900+
No production metadata was manually repaired.
901+
902+
GitHub PR91 identified two further corpus recovery cases: an initial registry
903+
failure followed by a pending retry could switch config to global, and stale forced
904+
refresh callers could return before the winning registry published. Both behavioral
905+
regressions failed on the prior stores. Successful registry resolution now has
906+
explicit state, and forced/shared callers follow successive winning request promises.
907+
All 18 real private Graph browser cases pass, including both new recovery cases.
908+
The first chain harness attempt held three identical requests and reached its
909+
fixture deadline; the corrected test holds at most two while exercising A-to-B-to-C
910+
supersession. This intermediate 18-case result was superseded by the final 26-case matrix
911+
and nine accounting cases described below.
912+
913+
The first focused D3 embedding suite passed 133 tests with one environment-gated
914+
skip. Its core Sol review found three valid issues: application processes could
915+
retain the upstream key, catalog upserts filled an invalid embedding base URL, and
916+
HTTP 408/409 did not receive configured application retries. Corrections and broader
917+
family tests are in progress. D3 remains excluded from deployment.
918+
919+
### September 5: failed full NASA replacement and remaining blockers
920+
921+
The explicitly forced staged replacement `0152e29560bf4d1fa9216375891b9d4f`
922+
started at 15:21:38 UTC. Docling completed at 15:42:36; semantic KG requests began
923+
at 15:42:39. Extraction failed with a request timeout at 15:46:17, before resolution
924+
or community summaries. The active generation `3054ecc26d3649a086758e04ece30488`
925+
and its 1,002 chunks/dense points remain intact. Failed staging rows, graph and
926+
Qdrant collection were reclaimed, and the run fence was released.
927+
928+
The native census contains 131 actual HTTP attempts with nine uncertain outcomes.
929+
The ledger subsequently recorded 130 successful provider rows totaling $3.0555729;
930+
the browser's manual refresh at 15:57:55 matched that amount and retained incomplete
931+
coverage with one missing request. Six native requests exceeded the configured
932+
30-second timeout (maximum 41.093 seconds), and five provider completions arrived
933+
after application failure. Four application OpenAI SDK retry delays were logged
934+
even though native gateway retry fields were zero. The official wrapper also has
935+
an independent rate-limit retry handler. Both hidden retry layers are now explicitly disabled in the pending PR91
936+
source. The real HTTP matrix failed 33 cases on the old source; all 92 relevant
937+
KG/census tests pass after the correction, and independent Sol review approved it.
938+
939+
The graph telemetry's 1,002 attempted/failed count is a whole-file exception count,
940+
not actual dispatch evidence. All chunks ran inside one official pipeline execution;
941+
successful per-chunk results were in memory and cannot be recovered from the
942+
content-free ledger or redacted traces. The estimate of $3.2875 also omitted most
943+
serialized schema/prompt overhead and understated output: observed requests averaged
944+
7,646 input and 493 output tokens. A retry must follow corrected estimation and
945+
timeout/retry handling; durable extraction reuse and truthful progress are being
946+
assessed. No additional paid full run has been started.
947+
948+
D3's reviewed 38-source suite passed 249 tests with one provider-capability skip.
949+
Sol's further catalog-family finding reproduced 13 failures in the broader upsert
950+
matrix; all 32 endpoint cases now pass. Existing model families are preserved, and
951+
new families derive from model identifiers rather than request capability labels.
952+
Final Sol review approved this correction. Ruff, mypy (179 source files), banned
953+
patterns, generated types, contract bundle and the 454-key config reality check pass.
954+
D3 is still undeployed.
955+
956+
PR91's final failure correction returns an explicit non-rejecting registry-load
957+
outcome and preserves the newest settled result for older mutation callers. An
958+
actual failed winning refresh reproduced the old mutation falsely reporting success.
959+
Sol approved the source correction. All 26 real Graph browser cases and all nine
960+
native-accounting browser cases pass on the final stores. The browser test driver
961+
now keeps async operations rooted in the actual rendered page: raw CDP evidence
962+
identified a collected evaluation promise, rather than application navigation. A
963+
separate update assertion now uses the supported corpus name field. Independent
964+
Sol review approved the driver changes. One earlier private rerun collided with
965+
pytest corpus cleanup sharing Neo4j; the final suites ran exclusively. Production
966+
was unaffected. Frontend lint, 23 unit tests, build, Ruff, mypy (178 source files),
967+
banned patterns and generated-type checks pass. GitNexus reports 14 changed files,
968+
66 symbols and 36 expected shared-store flows at critical risk.
969+
970+
The extraction recovery plan is recorded separately in
971+
`graphrag-extraction-recovery-2026-09-05.md`. Checkpoint persistence and a forecast
972+
that includes approved schema/prompt overhead are under implementation; neither
973+
is deployed or claimed accepted. No additional paid NASA rebuild has started.
Lines changed: 86 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,86 @@
1+
# GraphRAG Extraction Recovery Implementation Plan
2+
3+
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4+
5+
**Goal:** Preserve validated semantic extraction work across failed replacements and show accurate chunk progress without changing graph generation or promotion semantics.
6+
7+
**Architecture:** Store official, validated chunk graphs in the existing Postgres database before the official extractor adds lexical and run-specific identity. New runs reuse only exact matching extraction inputs, then pass deep copies through the existing official postprocessor, pruner, scoped writer, resolver, and promotion checks. Native gateway records remain the sole billing evidence.
8+
9+
**Tech Stack:** Python 3.12, Pydantic, asyncpg/Postgres, neo4j-graphrag 1.19.0, Neo4j, existing run records, React/Zustand and generated TypeScript.
10+
11+
**Spec:** `docs/exec-plans/active/graphrag-continuation-2026-09-04.md`, September 5 failed NASA replacement evidence, and the requirements below.
12+
13+
## Global Constraints
14+
15+
- Mac checkout is source only; all execution, databases, tests, builds and browser acceptance run on LXC100.
16+
- Preserve the official extraction/postprocessing/pruning/writing pipeline and current whole-generation promotion invariants.
17+
- Preserve the separate D3 embedding and D13 estimate changes; stage exact owned files or patches.
18+
- No raw prompts, API keys, gateway responses or spend ledger in the checkpoint table. Domain graph data remains in the corpus's existing private Postgres storage.
19+
- No automatic fresh paid dispatch after malformed, mismatched or unreadable checkpoint data.
20+
- Reusable identity excludes run ID, credentials, timeout and concurrency. It includes corpus, source path/file digest, exact chunk content/order/provenance, approved schema, rendered prompt/examples, model alias/upstream/sanitized endpoint, model parameters and implementation recipe versions.
21+
- An unchanged provider model slug cannot prove an unchanged provider implementation. Record this limitation; explicit deindex clears checkpoints.
22+
- A cache hit is an extraction reuse, not a new native request or new spend.
23+
- Run GitNexus impact before edits and detect-changes before committing. Public telemetry impact is CRITICAL; Postgres client impact is HIGH. Review against exact changes, then independent Sol xhigh and the normal PR loop.
24+
25+
## Task 1: Typed and fenced Postgres checkpoints
26+
27+
**Files:** Create `server/models/graph_extraction_checkpoint.py`; modify `server/db/postgres.py`; create `tests/integration/test_graph_extraction_checkpoints.py` in the existing pytest stack.
28+
29+
**Interfaces:** The new Pydantic envelope owns identity and the official `Neo4jGraph`, not a second graph DTO. Define `GraphExtractionCheckpointIdentity`, `GraphExtractionCheckpoint`, `graph_extraction_cache_key(identity)`, and the three Postgres operations below. A separate preparation operation may handle corpus recipe rotation if it keeps the same transaction/fence rule.
30+
31+
```python
32+
async def get_graph_extraction_checkpoint(
33+
self, repo_id: str, cache_key: str,
34+
) -> GraphExtractionCheckpoint | None: ...
35+
36+
async def put_graph_extraction_checkpoint(
37+
self, repo_id: str, owner_run_id: str,
38+
checkpoint: GraphExtractionCheckpoint,
39+
) -> None: ...
40+
41+
async def prepare_graph_extraction_checkpoint_file(
42+
self, repo_id: str, owner_run_id: str, recipe_hash: str,
43+
file_path: str, current_keys: list[str],
44+
) -> None: ...
45+
```
46+
47+
- [ ] Add real Postgres tests first: idempotent duplicate write, conflicting duplicate content, wrong corpus/key, malformed payload, absent/replaced run fence, corpus deletion, explicit deindex, and failed staging cleanup retention.
48+
- [ ] Run that suite on the private LXC Postgres and retain the behavioral failure evidence.
49+
- [ ] Add `graph_extraction_checkpoints` with `(repo_id, cache_key)` primary key, a corpus FK with cascade deletion, recipe/file lookup fields, creation time, and JSONB envelope. Register it in corpus-owned deletion inventory.
50+
- [ ] Canonicalize identity to stable sorted JSON and SHA-256. Validate the envelope and official graph on reads. Check current base-corpus fence inside the same transaction for writes and pruning. Reject conflicting rows; preserve identical rows.
51+
- [ ] Keep only the current recipe partition and prune obsolete keys only after a file's full current chunk set is known. Successful completion removes deleted-file entries; failed staging reclamation does not remove matching base-corpus entries.
52+
- [ ] Run the real persistence/lifecycle matrix and Ruff/mypy; independently review transaction ownership and deletion races before integration.
53+
54+
The identity model must carry explicit fields for the global constraints above. Avoid opaque caller-provided hashes without retaining their typed provenance. The envelope additionally records its originating run and pruning counts; its graph contains no staging/run/corpus scope properties. SQL identifiers are fixed, values parameterized.
55+
56+
## Task 2: Preserve completed official extraction results
57+
58+
**Files:** Modify `server/indexing/graphrag_pipeline.py`; create a focused `server/indexing/extraction_checkpoint.py` adapter if necessary; extend `tests/unit/test_graphrag_census.py`, `tests/unit/test_graphrag_pipeline.py`, and `tests/integration/test_graphrag_pipeline_live.py`.
59+
60+
**Interfaces:** Build an execution-scoped checkpoint context from the base corpus, owner run, Postgres client, exact approved extraction recipe, file and chunk provenance. Pass it explicitly to the existing extractor. `_DrainingExtractor.extract_for_chunk` consumes that context; the public official return type stays `Neo4jGraph`.
61+
62+
- [ ] Add a real HTTP fixture with several successful domain graphs followed by a timeout/refusal. Restart a separate worker against the same Postgres database and assert that only missing chunk identities are dispatched.
63+
- [ ] Prove each changed identity input prevents reuse, while changing only run ID, timeout or concurrency preserves reuse. Include two corpora with otherwise identical input.
64+
- [ ] Inside `extract_for_chunk` (already admitted by the official semaphore), validate a matching stored envelope. Return a deep copy on a hit. A miss calls the existing official method exactly once; SDK and wrapper retries stay disabled.
65+
- [ ] Validate reserved scope properties, entity-ID consistency and the official graph shape. Apply official `GraphPruning` before persistence, preserving pruning counts. Valid empty graphs can be reused but cannot bypass generation-level nonempty invariants.
66+
- [ ] Persist before counting reusable success. Retain and drain an initiated commit task on cancellation using the existing producer/drain ownership. Never release the fence while a checkpoint writer still owns work.
67+
- [ ] Pass the returned copy through unchanged official lexical postprocessing and final file pipeline. Assert that reused graphs acquire only the new run/staging scope and retain document, `NEXT_CHUNK`, and source links.
68+
- [ ] Exercise cancellation during commit, abrupt worker exit, fence takeover, corrupt rows and storage outage with real components. Failure must not silently issue a replacement paid call.
69+
70+
## Task 3: Accurate persisted progress and operator acceptance
71+
72+
**Files:** Modify `server/models/index.py`, `server/api/index.py`, `server/indexing/run_records.py`, the existing Indexing UI graph summary, its browser fixture/spec, and generated contracts. Keep D13 estimate helpers under their existing owner's control.
73+
74+
**Interfaces:** Extend `GraphExtractionTelemetry` with explicit reuse/cancellation/unfinished counts, defaulting new fields for historical records. A typed progress callback publishes admitted/completed outcomes from the extractor into the existing locked run-summary update path. Existing native census and costs remain independent.
75+
76+
- [ ] Add state-transition tests covering queued, admitted, validated, checkpointed, reused, failed and cancelled chunks. Initial queued tasks must not all count as attempted.
77+
- [ ] Remove semantic-path whole-file `attempted += len(chunks)` and `failed += len(chunks)` on exception. Preserve separate writer/promotion failure evidence and completed extraction counts.
78+
- [ ] Count success after durable acceptance, reuse separately, and worker duration on actual misses. Provider-latency estimates must exclude reused successes from their denominator.
79+
- [ ] Persist progress through existing run-record locks without clobbering later accounting or newer progress. Verify deliberately delayed progress/accounting writes in both orders.
80+
- [ ] Update the Indexing UI to show selected, attempted, reusable successes, reused, failed and unfinished work truthfully after reload. A failed run remains unpromoted even with successful checkpoint rows.
81+
- [ ] Regenerate TypeScript and contract bundles on LXC100; run changed-surface tests, Ruff, mypy, banned patterns, type/contract/config checks, full CI and the real browser matrix.
82+
- [ ] Obtain independent Sol xhigh review, complete the PR loop and deployed-marker verification. Only then change NASA's existing per-chunk timeout to a reviewed value, show the corrected estimate, run the approved replacement, and verify domain retrieval/source provenance in the real browser.
83+
84+
## Self-review and acceptance evidence
85+
86+
The three tasks cover persistence, official-pipeline reuse and operator truth separately. Identity mismatch, stale-writer races, corruption, deletion and cancellation are explicit requirements, not fallback cases. Store interfaces above are shared contracts; rename them coherently if implementation evidence requires a change. Existing native charges for the failed NASA run remain $3.0555729 with incomplete request coverage; this plan cannot recover outputs already lost by that run.

0 commit comments

Comments
 (0)