feat(research): expose explicit yield dimensions - #7092
Conversation
|
Important Review available on request
Reviews should be triggered manually for repositories with fewer than 10 stars. Select Trigger review above or comment ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
ll7
left a comment
There was a problem hiding this comment.
Exact-head self-review for d97684be52ca4942e3a59b1fccbba20e8b25ab6c.
No blocking findings for this bounded reporting-only slice. The implementation:
- exposes four closed research-yield dimensions with explicit query text, denominators, and
complete buckets; - rejects unknown dimension names, fields, buckets, missing dimensions, malformed counts, and
denominator mismatches; - preserves source snapshot path and SHA-256 identity and renders definitions in Markdown;
- does not infer counts from live issue/PR state and does not wire launch-gate behavior.
The 28-test focused suite, CLI smoke, Ruff check/format, docs proof-consistency, and exact diff
checks passed. The fixture/report output is implementation proof only, not benchmark or scientific
evidence. Domain-Aware Approval remains pending for this research-reporting surface.
🔍 PR Contract Check SummaryOverall Status: 🟢 PASSED
ℹ️ Info
This check is mechanized. Please resolve any blockers to pass CI. |
|
Exact-head implementation gate for PR #7092.
Gate conclusion: implementation proof accepted for this exact head; merge remains blocked pending ordinary review and Domain-Aware Approval. gate-verdict: accepted @ d97684b |
ll7
left a comment
There was a problem hiding this comment.
Exact-head self-review for PR #7092.
- Reviewed commit:
fbb30293952fbbe80bfd36a8287d5e6566eddc81. - Base:
4a790accb05575707fe4d66f546ebb9e4bbbfd6. - Scope remains report-only: explicit frozen-snapshot research-yield dimensions; production
admission and launch-gate wiring remain deferred. - Validation at this head: 28 focused tests passed, report CLI smoke passed with source SHA-256,
Ruff check passed, Ruff format check passed, andgit diff --checkpassed. - No implementation blocker found in this bounded slice.
Domain-Aware Approval remains pending for the research-yield classifications. This is not a
merge-ready, campaign-admission, benchmark-evidence, execution-authorization, or publication
approval.
ll7
left a comment
There was a problem hiding this comment.
Exact-head self-review for PR #7092 / issue #7090.
- Reviewed head:
ee908dbc25ec94d3d19e8dc65c750c79e200bff7. - Base:
4aac85935985434cf25c25442a1bc5bc216b78e3. - The report-only slice computes four explicit research-yield dimensions with closed query,
denominator, bucket, and unknown-field contracts. It does not admit production results, gate
launches, submit compute, or establish research claims. - Validation at this exact head: 28 focused tests passed; CLI smoke emitted
research_yield_report.v1with SHA2566bebd421...70a7af; Ruff check, format, and diff checks
passed. - No implementation blocker found in this slice. Domain-aware review remains pending; this is
diagnostic reporting, not benchmark evidence or publication approval.
ll7
left a comment
There was a problem hiding this comment.
Exact-head self-review for PR #7092 / issue #7090.
- Reviewed head:
0fc1d2b6bf450f9b609b1e7df16c4d0f9551f12e. - Base:
ade42921d6b00406f69ae3a39352bff6fa66a53d. - The report-only slice computes four explicit research-yield dimensions with closed query,
denominator, bucket, and unknown-field contracts. It does not admit production results, gate
launches, submit compute, or establish research claims. - Fresh-base validation: 28 focused report tests passed after merging current
origin/main; the
prior CLI smoke, Ruff/format, docs, and diff proofs remain applicable to the unchanged slice. - No implementation blocker found in this slice. Domain-aware review remains pending; this is
diagnostic reporting, not benchmark evidence or publication approval.
ll7
left a comment
There was a problem hiding this comment.
Exact-head self-review for PR #7092 / issue #7090.
- Reviewed head:
cd2b2ec2a9d554a41113fb02b6146c17ef39e49c. - Base:
a956b839e877ce18d03a5aa7903ff923952547da. - The report-only slice computes four explicit research-yield dimensions with closed query,
denominator, bucket, and unknown-field contracts. It does not admit production results, gate
launches, submit compute, or establish research claims. - Fresh-base validation: 28 focused report tests passed after merging current
origin/main; prior
CLI smoke, Ruff/format, docs, and diff proofs remain applicable to the unchanged slice. - No implementation blocker found in this slice. Domain-aware review remains pending; this is
diagnostic reporting, not benchmark evidence or publication approval.
ll7
left a comment
There was a problem hiding this comment.
Exact-head refresh review
- Reviewed PR #7092 at head
65c169649917728c941736ba05d11a36052e28e8against base2cce5e2f0916524916a4592062e47b1b7b1dc433. - Scope remains four files: campaign-manifest guidance, the research-yield report, its answerability-focused tests, and the versioned fixture.
- The implementation preserves frozen source identity, requires explicit query/denominator/bucket accounting, and rejects unknown or inconsistent dimensions.
- Parent proof: 28 focused tests passed; report CLI smoke emitted all four dimensions and source SHA-256; Ruff check/format, docs proof, and
git diff --checkpassed. - Final readiness: follow-up
ok, checklist errors none; domain status is explicitlypending_domain_approvalbecause this is research-yield reporting. - No benchmark, compute, evidence admission, publication, or merge was performed or inferred.
No blocking implementation findings at this exact head. Domain review remains required before merge readiness.
Exact-head refresh review
No blocking implementation findings remain at this exact head. Domain review remains required before |
ll7
left a comment
There was a problem hiding this comment.
gate-verdict: accepted @ 6f6c7f7
merge-ready: no
pr-metadata: reconciled @ 6afa98db8e608f6afe7d7765b80b41694a25a084b135757224cc556bb8d9ef8c
Implementation review accepted for the reporting-only slice. The four explicit dimensions are query-defined source metadata and remain unable to authorize campaign admission or support a scientific result. The exact-head repair correctly rejects duplicate snapshot IDs, non-finite lag values, and malformed in-memory snapshots through the same validation path used for file-backed snapshots. Focused tests (31), CLI smoke, scoped Ruff/format, documentation/evidence integrity, and exact diff checks passed. Domain-aware approval remains pending, so no merge-ready label is applied.
Exact-head refresh review
No blocking implementation findings remain at this exact head. |
ll7
left a comment
There was a problem hiding this comment.
Exact-head research-yield review
- Reviewed PR #7092 at exact remote head
96be21e92734dfb3270f4a97f81def9216800147against currentorigin/main5c4f96468a2bacd3a04133df519f6d8756040073. - The report now keeps empirical answers, infrastructure/preflight throughput, explicit lag summaries, and four query-defined research-yield dimensions separate. Snapshot validation rejects duplicate records, non-finite lags, unknown dimensions/buckets/fields, and denominator mismatches rather than inferring values from live issue or PR state.
- Focused research-yield/manifest tests passed 35. Full
BASE_REF=origin/main scripts/dev/pr_ready_check.shpassed through core and optional lanes with a clean tree. - Current exact-head hosted checks are terminal green: 29 passing checks and 2 intentional skips.
gate-verdict: accepted @ 96be21e92734dfb3270f4a97f81def9216800147;merge-ready: no. Domain-aware approval remains pending for the reporting/evidence-boundary change.- This is diagnostic process reporting only: no campaign, planner, benchmark, scientific, productivity, or paper-facing result was generated or admitted.
96be21e to
3740a82
Compare
|
gate-verdict: accepted @ 3740a82 Exact-head research implementation review for PR #7092 / issue #7090:
The reporting slice preserves a strict, frozen-snapshot boundary: it does not infer live GitHub state or scientific value, and it rejects malformed or incomplete dimensions. No campaign, compute, evidence admission, benchmark result, publication, or paper/dissertation-facing claim was made. Domain-aware approval remains pending, so this review intentionally does not add |
|
gate-verdict: accepted @ 3740a82 Current exact-head refresh for PR #7092 / issue #7090:
Domain-aware approval remains pending. This PR is not merge-ready and authorizes no merge, compute, evidence promotion, publication, or scientific claim. |
ll7
left a comment
There was a problem hiding this comment.
Review gate
Accepted implementation proof for the exact current-base revision:
- Head:
3740a828b5bc8df42fc982f621d243ee34a0eca9 - Base:
65ebc834e68e5a19a5a8711c521913dbfa023262 - Research-yield/answerability focused tests: 31 passed.
BASE_REF=origin/main scripts/dev/pr_ready_check.sh: passed with a clean stamp; optional lane and all ratchets passed.- Hosted checks for this head: 34 total, 32 success, 2 skipped, and no failures.
The implementation gate is accepted at this exact head. This remains merge-ready: no: the change is diagnostic process reporting, domain-aware approval remains required, and the report cannot authorize campaigns or establish scientific value. No benchmark execution, compute, evidence admission, publication claim, or merge authorization is inferred from this review.
ll7
left a comment
There was a problem hiding this comment.
Review evidence correction
The exact-head implementation review remains valid at 3740a828b5bc8df42fc982f621d243ee34a0eca9 against 65ebc834e68e5a19a5a8711c521913dbfa023262.
The canonical latest-run REST monitor now reports 31 checks, 29 success, 2 skipped, 8 superseded reruns, and overall success. The raw check-run total includes superseded contract reruns; the PR body records the canonical latest-run count. merge-ready: no remains unchanged: domain-aware approval is pending, and no campaign, compute, evidence admission, or scientific claim is authorized.
Exact-head audit handoff
Proof
Boundary and dispositionThis is a frozen-snapshot research-process/reporting contract. It does not infer live GitHub or
|
ll7
left a comment
There was a problem hiding this comment.
Exact-head review: #7092
- Reviewed head:
52c8693f1b1729e2f82ee3c7ae0ee01d21640196 - Current base used for local proof:
origin/mainat50a511d2336c37ae3957660a0141494d516be668 - Scope: reporting/process support for explicit research-yield dimensions; this does not establish a scientific, benchmark, or paper-facing result.
Evidence at this exact head:
- Focused tests: 62 passed (
tests/benchmark/test_research_answerability.pyandtests/docs/test_next_issue_shortlist.py). - Ruff, format, and
git diff --check: passed. - Final
scripts/dev/pr_ready_check.shagainst the currentorigin/main: passed with a clean tree. The readiness stamp records base50a511d2, head52c8693f, and statuspassed. - Hosted checks for the exact head: terminal green; the known skipped checks are non-blocking, and superseded duplicate contract runs were cancelled.
- No unresolved, non-outdated review threads were found.
Decision boundary:
- Domain approval remains pending.
- No campaign, compute, benchmark-evidence, or paper-claim admission is implied by this review.
merge-ready: nountil the domain owner accepts the reporting contract and the PR is explicitly promoted.
ll7
left a comment
There was a problem hiding this comment.
Exact-head review: #7092
- Reviewed PR head:
1ecd16c3f0cfa0378eef48f6962eeb68b5572726 - Current base:
origin/main=59e7f9e9ca4898dcdb977b8809e6ccedfc434601 - Local final readiness: passed with a clean tree, current-base stamp, core and optional/predictive
lanes, coverage diagnostics, TODO-docstring and broad-exception ratchets. - Focused proof: 62 answerability/yield and current-main shortlist tests; report CLI JSON/Markdown
generation from the frozen fixture; docs/evidence integrity; docs proof consistency; Ruff,
format, and diff checks passed. - Hosted exact-head CI: terminal success with 29 successful checks and 2 expected skips; expected
head matched and no superseded run was used as evidence.
The report preserves explicit query-defined dimensions and source snapshot identity. It does not
run a campaign, establish a scientific or benchmark result, rank planners, promote evidence, or
admit paper, visual, or dissertation claims. Fallback/degraded/blocked/unavailable states remain
explicit boundaries.
Domain-aware approval is still pending. merge-ready: no; no merge or scientific-admission
decision is made by this review.
ll7
left a comment
There was a problem hiding this comment.
Reviewed PR #7092 / issue #7090 at exact live head 1ecd16c3f0cfa0378eef48f6962eeb68b5572726 against current origin/main 59e7f9e9ca4898dcdb977b8809e6ccedfc434601.
Exact-head implementation and reporting-boundary review:
- Focused answerability/yield suite: 31 passed.
- Current-main regression slice
tests/docs/test_next_issue_shortlist.py: 31 passed. - Frozen-snapshot dimension vocabulary, query/denominator/bucket checks, finite-lag validation,
duplicate-record rejection, CLI help, docs/evidence integrity, evidence-registry ratchet,
Ruff, format, andgit diff --checkpassed. - Full
PR_READY_MODE=finalcore and optional readiness passed with a clean exact-head stamp. - Hosted exact-head CI is terminal success: 31 checks, 29 success, 2 skipped, no failures,
expected-head match true.
Research/evidence boundary accepted for implementation integrity only: this reports explicit
process/evidence-readiness dimensions from a frozen snapshot. It does not reconstruct live issue,
PR, campaign, planner, or metric state; run campaigns; submit compute; admit evidence; close
issues; publish; or establish scientific, benchmark, planner, paper, or dissertation results.
Fallback, degraded, unavailable, blocked, and diagnostic-only states remain explicit.
Domain-aware approval remains pending for evidence classification, experimental comparison,
benchmark interpretation, and paper-facing claim boundaries. Counts must not be promoted into
scientific progress or empirical answers without the required domain decision.
gate-verdict: accepted @ 1ecd16c
pr-metadata: reconciled @ c0795d9f360b7377c1ea4cc12edb003b759db90bab11d3069b38e084de4230ec
merge-ready: no
ll7
left a comment
There was a problem hiding this comment.
Reviewed PR #7092 / issue #7090 at exact live head 5f18e9a4c1a0afb4a1aa3b8282a2b8d82f412cdb against current origin/main 60124e812c693a8b14f0130ed7565f11588cc9ee.
Exact-head implementation and reporting-boundary review:
- Focused answerability/yield suite and current-main regression slice: 62 passed.
- Frozen-snapshot dimension vocabulary, query/denominator/bucket checks, finite-lag validation,
duplicate-record rejection, CLI help, docs/evidence integrity, evidence-registry ratchet,
Ruff, format, andgit diff --checkpassed. - Full
PR_READY_MODE=finalcore and optional readiness passed with a clean exact-head stamp
output/validation/pr_ready/review-pr-7092-20260815c.json. - Hosted CI for the refreshed head must settle before final gate evaluation.
Research/evidence boundary accepted for implementation integrity only: this reports explicit
process/evidence-readiness dimensions from a frozen snapshot. It does not reconstruct live issue,
PR, campaign, planner, or metric state; run campaigns; submit compute; admit evidence; close
issues; publish; or establish scientific, benchmark, planner, paper, or dissertation results.
Fallback, degraded, unavailable, blocked, and diagnostic-only states remain explicit.
Domain-aware approval remains pending for evidence classification, experimental comparison,
benchmark interpretation, and paper-facing claim boundaries. Counts must not be promoted into
scientific progress or empirical answers without the required domain decision.
gate-verdict: accepted @ 5f18e9a
pr-metadata: reconciled @ 7584bf6af4f3a3b82d6c514bfffaf000ff75750e2ac3edf1810f265f2852ebad
merge-ready: no
Exact-head audit handoff
Proof
Boundary and dispositionThis is a frozen-snapshot research-process/reporting contract. It does not infer live GitHub or
|
ll7
left a comment
There was a problem hiding this comment.
Terminal exact-head review: PR #7092 / issue #7090
Reviewed the current implementation at exact head e0f7354e0d1499cc1f4ebe5dc35b8a4c2708f5e8 against current origin/main=752e05f5574543bae4fbbde61be7ff79f501d9d9.
Implementation evidence:
- The linked worktree is clean and the diff remains limited to the four intended frozen-snapshot reporting, fixture, test, and documentation paths.
- Current changed report/answerability validation passed
31tests. - Source-backed CLI generation emitted all four explicit query-defined dimensions, preserved the frozen source SHA-256, and generated both JSON and Markdown outputs.
- The prior shared contract and benchmark transition/lazy-import bundles remain
63and31passed respectively; Ruff, formatting, docs/evidence integrity, proof/catalog consistency, evidence-registry ratchet419/419, and diff checks passed. - Full
BASE_REF=origin/main PR_READY_MODE=final PR_READY_ADVISORY=1 scripts/dev/pr_ready_check.shpassed at2026-08-15T16:59:33.941743+00:00on this exact head and current base with a clean tree, core/optional lanes, and ratchets. Stamp:output/validation/pr_ready/review-pr-7092-current-20260815-v2.json. - Hosted checks are terminal and healthy: 31 total, 29 successful, 2 intentional skips, 0 pending, 0 failed; 2 superseded historical runs are excluded by the queue classifier.
- Base freshness is fresh and the review-thread audit is complete with 0 unresolved threads.
- PR metadata is reconciled at
d01c4935b0ed173b7b685766b8d77fcde29b01393b57d446e65c17bd44bb0f19.
The explicit query/denominator/bucket reporting contract and fail-closed snapshot validation are accepted for the recorded diagnostic scope. Domain-Aware Approval remains pending for answerability/reporting semantics, evidence classification, and research-progress interpretation. The report does not infer live workflow state or scientific progress, and this review establishes no empirical result, benchmark ranking, evidence admission, publication, compute, or paper-facing claim.
gate-verdict: accepted @ e0f7354
pr-metadata: reconciled @ d01c4935b0ed173b7b685766b8d77fcde29b01393b57d446e65c17bd44bb0f19
merge-ready: no
ll7
left a comment
There was a problem hiding this comment.
Exact-head research-result/reporting review
Reviewed PR #7092 / issue #7090 at exact head ba2f029042d55833ddff192749dfe0ef43574ab2
against current origin/main 9b96aac9cf7a1a190662d6a499fcbbffbd0b8a77.
- The isolated worktree was refreshed by merging the live base; the current-base delta remains
exactly the four report/manifest/fixture/test paths. - Focused yield/answerability proof passed: 31 tests. Related manifest, campaign-runner, and
benchmark lazy-import regression proof passed: 28 tests. - The CLI emitted all four explicit query-defined dimensions and preserved frozen source SHA-256
6bebd421e7eaaa74883b2a9a6c477357cfe43ac90704050a342c4fd41070a7af; JSON/Markdown outputs
remain scratch diagnostics. - Final readiness passed at
2026-08-16T13:57:59.133097+00:00UTC with a clean tree, exact
base/head, core and optional/predictive lanes, ratchets, and diff checks. Metadata was
reconciled through REST with digest1b071f64ddf0ebaaee649e16c4ab7695127a4d1bf14e9046136fbb5aed8bc296.
Implementation integrity is accepted at this exact head. This remains diagnostic workflow
reporting only: it does not infer scientific progress from issue/PR state and establishes no
benchmark ranking, empirical result, evidence admission, publication, or paper-facing claim.
Domain-aware approval remains pending. This review does not add merge-ready, authorize a
campaign, or promote generated report output to durable evidence.
gate-verdict: accepted @ ba2f029042d55833ddff192749dfe0ef43574ab2
pr-metadata: reconciled @ 1b071f64ddf0ebaaee649e16c4ab7695127a4d1bf14e9046136fbb5aed8bc296
merge-ready: no
compute: none
ll7
left a comment
There was a problem hiding this comment.
Exact-head research-result/reporting review — PR #7092
- Reviewed head:
379e848bbce84742e887f4c7bf9e9c394c25f448. - Reviewed base:
462032df2abc3e086655935288c806b9df8bda2b(origin/main). - Rebased/merged cleanly onto current
main. The 4-path diff remains diagnostic research-yield reporting infrastructure. - Local validation: 31 focused yield/answerability tests passed. Ruff check and format passed.
git diff --checkpassed. - Domain-aware approval: pending maintainer decision.
- Reconciled PR metadata:
pr-metadata: reconciled @ 4142fc0eb94c1b27881bde286b60ba437318a3ee07bc6ac38f269220caa88c54.
gate-verdict: accepted @ 379e848bbce84742e887f4c7bf9e9c394c25f448
merge-ready: no (domain-aware approval pending)
Maintainer Decision RequiredWhat is complete:
What is missing for merge:
|
Exact-head review: PR #7092 / issue #7090
Why this is not author-reservedThe change is diagnostic workflow tooling. It adds strict schema validation and rendering for four Validation at this exact head
Remaining gateOnly terminal green CI at Claim boundary unchanged: diagnostic reporting contract only. No campaign, compute, SLURM
|
ll7
left a comment
There was a problem hiding this comment.
Exact-head review: PR #7092 / issue #7090
- Reviewed head:
780377325d544464bdd5f2ec9566049f640523f4 - Base:
a1892cf453973cd19e7bbba158a9f4132009bcee(origin/main), merged in cleanly (no conflicts) - Scope verified against the PR body: exactly four files —
scripts/analysis/report_research_yield.py,
tests/benchmark/test_research_answerability.py,tests/fixtures/research_yield_snapshot.v1.json,
docs/benchmark_campaign_manifest.md.
Why this is not author-reserved
The change is diagnostic workflow tooling. It adds strict schema validation and rendering for four
closed, query-defined process dimensions (duplicate/competing PRs, post-merge repairs, admitted
result packets, blocked-age buckets) read from an already-frozen research_yield_snapshot.v1. It
edits no claim ledger, no manuscript or dissertation prose, no locked evidence record, and no
preregistration; it admits no evidence and changes no benchmark metric, split, planner, or scoring
rule. It changes no orchestrator authority or auto-merge behavior. report_research_yield.py is not
invoked by any workflow, gate, or check — its only reference in the repository is the example command
in docs/benchmark_campaign_manifest.md — so it cannot decide what is admitted; it only reports
counts the snapshot already declares. Every added code path is fail-closed (unknown dimensions,
unknown fields, unknown/missing buckets, duplicate record IDs, non-integer or negative counts,
non-finite lag values, and denominator/bucket-sum mismatches all raise). The admitted_result_packets
dimension is a count copied from the frozen snapshot, not an admission decision. The docs change
records the reporting contract without asserting any result. The decision-required label was
conservative; the residual uncertainty here is reducible by review, so it is being removed.
Validation at this exact head
- Focused answerability/yield suite: 31 passed.
- Downstream check: no workflow, script, or gate consumes
report_research_yield.py. ruff check,ruff format --check,git diff --check: passed.check_base_sensitive_gates.py --pr 7092:gate_required: false(selectorordinary).check_docs_evidence_integrity.py --base-ref origin/main: 4 changed files passed.- Unresolved review threads: 0. Requested reviewers: 0.
- PR body reconciled to this head before this review.
- Hosted CI at this head is queued, not green: at review time the repository had 100+ queued and
15 in-progress Actions runs, and none of this PR's 25 checks had started. That is runner
starvation, not a defect of this PR.
Remaining gate
Only terminal green CI at 780377325d544464bdd5f2ec9566049f640523f4. No further push is planned, so
the gate-verdict trailer below stays current at that exact head. decision-required is removed
because the domain question it encoded is answered above; merge-ready is intentionally not
applied yet and is the single remaining action once checks conclude green.
Claim boundary unchanged: diagnostic reporting contract only. No campaign, compute, SLURM
submission, evidence admission, publication, ranking, or scientific-progress claim is made or
authorized by this review. Merge authority remains with the guarded merger.
gate-verdict: accepted @ 780377325d544464bdd5f2ec9566049f640523f4
base-policy: ordinary-cas @ 780377325d544464bdd5f2ec9566049f640523f4
pr-metadata: reconciled @ 3158fb64cf5d60ce781c2ae4d126f989c212de96932f2f1cd0b8df27edbdb18c
merge-ready: no (awaiting terminal CI only; not a domain block)
PR reconciliation — superseded by #7469PR #7469 contains this PR's complete research-yield reporting surface ( Keeping both branches open creates overlapping ownership of the same schema, fixture, tests, and documentation. Close this PR as superseded; preserve its commits as review history. Any review findings specific to the yield dimensions must be carried to #7469. Canonical owner: #7469. |
## Summary Harden the `goal-pr-review` skill for the situation observed on 2026-08-18: many reviewer lanes plus the autonomous factory acting on the same PR set, a red shared `main`, `decision-required` labels that were mostly reducible, and PR bodies whose metadata was reconciled byte-wise but not truth-wise. Docs-only; no runtime behavior changes. ## Linked Issues - Relates to `#7448` (fabricated exact-head SHAs in gate-verdict trailers) - Relates to `#7491` (merge-ready + stale "not merge-ready" narrative reaching main) - Relates to `#7482` (shared-main namespace baseline that made unrelated PRs look red) ## Stack / Dependency - Base dependency: none - Required prior PRs: none - Stack follow-up issues: none - Safe to review independently: yes - Review dependency reason, if any: none ## What Changed - `.agents/skills/goal-pr-review/SKILL.md` (Codex mirror is a directory symlink; no separate copy): - **Concurrent Writers** section: read-live-before-mutate with an active-writer window; advisory `review-claim: <lane> @ <head> until <UTC>` marker; content-identical head moves (main refresh only) transfer findings but require re-publishing exact-head carriers after green CI; factory/owner label sweeps are authoritative ("one label away" reporting, no re-apply in-run); successor/superset detection across open PRs on the same issue. - **Shared-Main Baseline Before Per-PR Diagnosis**: run the base-sensitive marker suite on `origin/main` first; classify matching PR failures as `shared_main_blocked`; route one bounded repair instead of per-PR "fixes". - **Decision-Required Triage**: author-reserved taxonomy (claim/evidence admission, preregistration authorization, release/settings/secrets, authority *expansion* vs fail-closed narrowing) vs reducible; new `author_decision` parking state mapped to `AUTHOR_DECISION_REQUIRED`; ≤25-line decision-packet format; keep the branch mergeable while parked. - Step 4 body check: every 40-hex SHA in the body must resolve and equal live carriers (never prefix-complete); not-ready sentences must be re-narrated before `merge-ready`. - Step 9 + **Shared Resource Budget**: stop waiting under runner starvation and publish what exists; API-quota guard; no background pollers; worktrees on repo disk / free-space check; detached-HEAD inventory-test gotcha. - Output Requirements now include a terminal state, "one label away" detail, packet location, parked/racing writers. - `CHANGELOG.md`: Unreleased/Changed entry. ## Why It Matters - Added value: turns the ad-hoc rules ten parallel review agents had to be told out-of-band into the skill contract itself; each item is tied to an incident (evidence voided by mid-run rebases on 6 PRs; #7092/#7102 closed and a full `merge-ready` sweep while merges were queued; #7482 re-diagnosed per PR; 12 of 15 `decision-required` PRs proved reducible; #7374/#7435 bodies said "not merge-ready" while labeled/merged; API quota exhausted twice; `/dev/shm` full twice). - Expected impact: fewer racing writes, fewer wasted CI reruns, fewer author packets for reducible decisions, no unverifiable SHAs or stale narratives in squashed history. - Why this is worth merging now: the same review-drain pattern is running daily. ## Research Result Guidance - Target claim / hypothesis / blocker this should affect: NA — workflow docs only - Comparator or baseline, if applicable: NA - Evidence tier: docs-only - Result classification: NA - Decision or stop rule, if applicable: NA - Parent issue, claim map, registry, context note, or synthesis surface to update: none - New research/benchmark/metric/paper-facing analysis tool, if any: NA - support helper (skill contract text) ## Domain-Aware Approval - Required for this PR: no - docs-only skill contract; no evidence classification or claim surface changed - Domains reviewed: NA - Status: not required - Approver/review source or waiver: NA - Validity checklist. Keep these machine-detected labels unchanged: - Target claim/hypothesis: NA - Comparator or split/evidence validity: NA - Fallback/degraded exclusions: NA - Claim boundary: NA - Implementation integrity vs experimental validity: NA ## Falsification / Non-Transfer Check NA — no empirical claim. ## Next Empirical Action NA. ## Validation / Proof - `uv run python scripts/dev/check_skills.py` → Validated 55 skills, typed registry, generated README, and routing tests. - `uv run pytest tests/dev/test_check_skills.py tests/dev/test_factory_v2_skill_contract.py tests/dev/test_token_efficient_thread_profile_snapshot_command.py -q` → 64 passed, 1 skipped. - `uv run python scripts/tools/sync_ai_config.py --check` → 7 symlinks validated. - `uv run python scripts/dev/check_docs_evidence_integrity.py --files …` → 2 changed files passed. - `uv run pre-commit run --files .agents/skills/goal-pr-review/SKILL.md CHANGELOG.md` → passed. ## Performance Evidence NA — docs only. ## Risks / Rollout - The `review-claim` marker is advisory until tooling (factory / `pr_loop_policy.py`) recognizes it; the text says so. - `author_decision` is a documented parking state, not a new `pr_loop_policy.py` classification; a follow-up could add machine detection of `decision-required` + packet presence. ## Docs / Provenance - Skill contract updated in place; CHANGELOG entry added. ## Downstream Propagation NA — not an evidence-producing PR. ## Follow-Up Issues - Optional: teach `pr_loop_policy.py` / merge gate to recognize the `review-claim` marker and the not-ready-sentence check (#7491 already tracks the gate side). ## Reviewer Notes Diff is additive apart from one table row (`failed_ci` mapping) and the Output Requirements list. Long lines are inside Markdown table rows, matching the existing file style. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Refs #7090.
Summary
Extend the frozen
research_yield_snapshot.v1reporting contract with four explicit, query-defineddimensions: duplicate/competing pull requests, post-merge repairs, admitted result packets, and
blocked-age categories. The report preserves each dimension's query, denominator, and complete
bucket counts instead of inferring research yield from live issue or pull-request state.
The v1 contract remains strict: older snapshots without the required
dimensionsblock arerejected. Generated reports remain scratch diagnostics and are not durable benchmark evidence.
Research and Evidence Boundary
without turning implementation activity into scientific progress?
reporting dimensions and does not compare experimental arms or rerun a campaign.
evidence admission, or paper/dissertation claim is made.
denominator mismatches, and non-finite lag values before campaign, compute, or evidence admission.
produced.
Changes
query/denominator/bucket contracts.
packets, and blocked-age categories.
snapshots.
non-finite values, denominator mismatch, and legacy snapshots.
docs/benchmark_campaign_manifest.md; production campaign-admission composition remainsowned by friction: wire answerability into production campaign admission and complete research-yield dimensions #7090.
Domain-Aware Approval
research-progress interpretation.
implementation loop; maintainer/domain review remains pending.
dimensions without inferring scientific progress from issue or PR state.
digest, explicit denominators, and closed bucket definitions; no experimental split or metric
is changed.
malformed counts, non-finite lags, and denominator mismatches fail closed; no fallback or
degraded evidence is admitted.
or scientific-progress claim is established.
domain review and any downstream research interpretation remain separate.
Validation / Proof
780377325d544464bdd5f2ec9566049f640523f4.a1892cf453973cd19e7bbba158a9f4132009bcee(origin/main), merged in cleanly with no conflicts.ordinary(check_base_sensitive_gates.py --pr 7092reports no base-sensitive files).git diff --check: passed.Downstream / Delivery Boundary
does not close or launch it.
leaderboard, or paper-facing result was produced.
scientific-progress claim without separate evidence review.
gate-verdict: accepted @ 780377325d544464bdd5f2ec9566049f640523f4base-policy: ordinary-cas @ 780377325d544464bdd5f2ec9566049f640523f4merge-ready: no (awaiting terminal CI only)