Overall result: PASS. P3-005 executed UAT-011 through UAT-018 exactly as
defined in uat-test-scripts.md, with UAT-015 executed
as its two required subcases (15A and 15B), for nine recorded results. An
independent replay in fresh synthetic state later re-executed all nine entries
and passed. This document is the consolidated durable evidence record; it
supersedes the original 1,019-line report draft, preserving its material
observations at proportional length.
This is a synthetic educational execution with no PHI. It is not clinical or regulatory validation and does not claim production fitness, certified HL7/FHIR conformance, or Epic/Beaker build experience.
| Fact | Recorded value |
|---|---|
| Task | P3-005 - v1.1 Recovery UAT Evidence Pilot |
| Dispatched baseline | 6f20a41dedf0087e22388b65230b70a0d58c94f8 (main), explicitly dispatched once by Austin |
| Recorded execution | 2026-07-25 21:45 UTC, by Codex |
| Independent replay | 2026-07-25 23:08 UTC, fresh synthetic state, new REPLAY-* request IDs |
| Python / platform | 3.12.13 on Linux x86_64; isolated virtual environment from requirements-dev.txt |
| Data | Repository synthetic recovery corpus only; no PHI |
| State | Fresh in-memory database per procedure (UAT-011/012 share one by design); separate file-backed scratch database for UAT-018, deleted after use |
- Baseline, before the recorded UAT sequence:
python -m pytest -qpassed 164 tests across eight suites;python -m src.demo_runexited 0 with all five scenarios; the UAT definitions, recovery corpus manifest, and the three frozen v1.1 records were byte-identical to the dispatched baseline. - Post-run, after the recorded execution and again after the independent replay: 164 tests passed and all five demonstration scenarios passed.
- A pre-recording dry run in disposable state caught one scratch-driver query
error (
ORDER BY audit_idcorrected toevent_id), was restarted clean, and passed all nine entries before the recorded execution began. Dry-run state was destroyed and is not recorded evidence.
| UAT | Requirements | Recorded execution | Independent replay |
|---|---|---|---|
| UAT-011 | R-020, R-021 | PASS | PASS |
| UAT-012 | R-022, R-023, R-024, R-027, R-028 | PASS | PASS |
| UAT-013 | R-025, R-026 | PASS | PASS |
| UAT-014 | R-029, R-030 | PASS | PASS |
| UAT-015A | R-032, R-037, R-038 | PASS | PASS |
| UAT-015B | R-032, R-037, R-038 | PASS | PASS |
| UAT-016 | R-034, R-035, R-036, R-041 | PASS | PASS |
| UAT-017 | R-033, R-040 | PASS | PASS |
| UAT-018 | R-030, R-031, R-039 | PASS | PASS |
One concise entry per UAT. Every action went through the public service
(recovery.retry_queue_item, recovery.redrive_queue_item,
recovery.get_recovery_history) or public ingestion; raw SQL was read-only.
Ingesting original fixture 09 against an open order produced an OPEN queue
item classified SPECIMEN_INCOMPATIBLE / SPECIMEN / REDRIVE_ONLY with both
timestamps null; ingesting fixture 06 against a finalized order produced a
TERMINAL item classified ORDER_FINALIZED / ORDER_STATE / TERMINAL with
terminal_at set and resolved_at null. Retained reasons named the
incompatible specimen and the finalized order. PRAGMA foreign_key_check
returned []. PASS.
A corrected re-drive of the OPEN fixture-09 item (request REC-UAT12-A)
produced one REDRIVE_CORRECTED / SUCCEEDED attempt linking queue 1 to new
message 3, which was FILED with two results and one INBOUND_RESULT_FILED
event; the queue became RESOLVED. The field-complete immutability comparison
matched before and after on every named field: original message_id,
control_id (RCO-09), status (ERRORED), created_at, payload bytes
(481) and SHA-256
4ad9b9b0991841f261003d266c35f493b0a52fb128787da4a4487f0cb7d63786, and the
identical queue raw_payload bytes and SHA-256. The new message had a
distinct ID and the corrected-payload fingerprint
554bb6c37dfeec5794e68d9cabb0535bb43becf02946aa8d60364c1e591d2e9d. PASS.
Fixture 05 without a matching order queued OPEN as
ORDER_NOT_FOUND / ORDER_MATCHING / RETRY_OR_REDRIVE. After the matching
order was created, retry_queue_item produced a RETRY_ORIGINAL / SUCCEEDED
attempt: the original and new payloads were both 485 bytes with identical
SHA-256 312eae192cbb723ef06eb9e7d1eed8db148e9d488a6d9633b4c3cf1954db6437,
proving the retry used a byte-for-byte copy of the immutable linked original
message payload. Message IDs were distinct, the original stayed ERRORED, the
new message was FILED, and the queue became RESOLVED. PASS.
Three independent cases all returned REJECTED attempts with no resulting
message: a request against a TERMINAL fixture-06 item (queue stayed TERMINAL,
order stayed FINALIZED); an unchanged retry against an OPEN REDRIVE_ONLY
fixture-08 item (queue stayed OPEN, nothing filed); and a retry whose target
order had since been finalized, which moved the queue OPEN to TERMINAL with
terminal_at set while order state, result counts, and filing-event counts
were unchanged. PRAGMA foreign_key_check returned []. PASS.
Re-driving the still-invalid original fixture-11 payload produced a FAILED
attempt with its resulting message preserved as ERRORED and zero
fish_result, RESULT_ENTERED, and INBOUND_RESULT_FILED rows; the queue
remained OPEN with both timestamps null. A later corrected re-drive under a
new request ID SUCCEEDED, FILED its message, and RESOLVED the queue. The item
held exactly one FAILED and one SUCCEEDED attempt. PASS.
The public workflow dependency workflow.enter_fish_result was temporarily
replaced so call 2 raised a handled InboundError after call 1 had really
written; the replacement was restored unconditionally in finally and
workflow.enter_fish_result is real_enter was confirmed afterward. The
injection-point evidence recorded state["calls"] == 2.
Observed in-transaction pre-fault milestone, captured read-only after call 1
returned and before call 2 raised: fish_results = 1, RESULT_ENTERED = 1,
conn.in_transaction is True.
Observed after the handled FAILED outcome: fish_result = 0,
RESULT_ENTERED = 0, INBOUND_RESULT_FILED = 0; the order remained
IN_PROCESS, the queue remained OPEN with both timestamps null, the attempt
was preserved as FAILED with its message ERRORED, and
conn.in_transaction is False. The first write was therefore proven to occur
inside the open transaction and to be rolled back, while the approved
handled-failure records committed. A later new request ID SUCCEEDED, FILED
its message, and RESOLVED the queue with exactly one SUCCEEDED attempt.
PASS.
After a successful corrected re-drive, an identical replay of the same
request ID returned the existing SUCCEEDED attempt with message, FISH-result,
attempt, and filing-event counts identical before and after (2/2/1/1). A new
request ID against the resolved item produced a REJECTED attempt with no
resulting message. Reusing the original request ID with a different actor
raised RequestIdConflictError, added exactly one REQUEST_ID_CONFLICT
audit event, changed no counts, and left exactly one SUCCEEDED attempt.
PASS.
A fresh ORDER_NOT_FOUND item was driven through FAILED (no matching order),
SUCCEEDED (after the order was created), and REJECTED (after resolution)
retries, then a mismatched-actor reuse of the first request ID.
get_recovery_history returned the three attempts in attempt_id order,
each carrying the full payload SHA-256
312eae192cbb723ef06eb9e7d1eed8db148e9d488a6d9633b4c3cf1954db6437 and a
human-readable outcome_detail; the conflict did not appear as a history
row. The complete observed outcome_detail values (independent replay,
request IDs REPLAY-UAT17-H1/H2/H3) were:
- FAILED:
Recovery processing failed (ORDER_NOT_FOUND); attempted message preserved as ERRORED, queue item 1 left OPEN. No order matches accession number ACC-REC-0500-NOMATCH. - SUCCEEDED:
Recovery succeeded via RETRY_ORIGINAL; new message 3 FILED to order 1 (2 probe result(s)); queue item 1 moved OPEN -> RESOLVED. - REJECTED:
Queue item 1 is RESOLVED; recovery is not permitted against a closed queue item.
The filing audit detail was message_id=3 probes=2 (RUNX1T1_RUNX1, CBFB),
and the conflict audit detail was request_id=REPLAY-UAT17-H1 reused with mismatched actor('analyst01'!='different-actor'); original attempt_id=1 (outcome=FAILED) left unchanged. The recorded 21:45 UTC execution observed
the same filing detail and the equivalent conflict detail for its
REC-UAT17-H1 request. PASS.
A successful corrected re-drive ran against a file-backed scratch database.
Immediately after the service call, conn.in_transaction was False and
PRAGMA foreign_key_check returned []. After closing and reopening the
file, the queue item was durably RESOLVED with resolved_at set and
terminal_at null, exactly one SUCCEEDED attempt persisted, the resulting
message persisted as FILED with the corrected-payload SHA-256, and the
post-reopen PRAGMA foreign_key_check returned []. The scratch database
was deleted in unconditional cleanup and did not exist afterward. PASS.
The first independent review of the original evidence handoff blocked
acceptance because required exact snippets and details were missing
(AF-2026-011). Rather than reconstructing the deleted scratch runner as
historical fact, a clearly labeled independent replay was executed on
2026-07-25 at 23:08 UTC in fresh synthetic state with new REPLAY-* request
IDs and a new file-backed UAT-018 database. All nine entries passed,
directly re-observing the UAT-015B in-transaction pre-fault milestone
(fish_results=1, RESULT_ENTERED=1, in_transaction=True), the
post-rollback zero counts, and the complete UAT-017 detail strings quoted
above. The replay supplied the previously missing evidence; it did not alter
or relabel the original 21:45 UTC observations, and it did not by itself
accept the pilot. Because the same agent produced both the recorded run and
this replay, the replay is a second execution in fresh state, not a review
by an independent party; acceptance of this report requires genuinely
independent review.
Every result above is reproducible from durable repository artifacts alone:
- Procedures:
uat-test-scripts.md, UAT-011 through UAT-018, executed exactly as written through the public service. - Automated equivalents:
tests/test_recovery_service.py(54 tests, including invariants I-01 and I-02 and the mid-operation rollback),tests/test_recovery_schema.py(29), andtests/test_failure_classification.py(20). - Standard validation commands:
pip install -r requirements-dev.txt,python -m pytest -q(expect 164 passed), andpython -m src.demo_run(expect five scenarios, exit 0).
No scratch runner, database, console log, or virtual environment was committed; the drivers were disposable by contract and are not part of the evidence chain.
Two process failures occurred around this execution and are recorded as
findings in AUDIT_FINDINGS.md:
- Publication preflight failure (AF-2026-010). The execution gate verified
the baseline, branch state, and test environment but not an authenticated
publication path. Only after the recorded UAT and local evidence commit were
complete did the builder discover the workspace could not push or use
gh. Publication required a separately authorized equivalent-tree fallback through the connected GitHub app; matching tree and blob SHAs preserved byte-for-byte evidence equivalence. - Incomplete first evidence handoff (AF-2026-011). The first published
report omitted contract-required exact snippets, the actual UAT-015B
pre-fault capture code, and a full UAT-017
outcome_detailexample, and the PR description omitted the result table. Independent review blocked acceptance; the independent replay above supplied the missing evidence.
Both disclosures are part of this durable record so the evidence chain is honest about how it was produced.