calibration: week of 2026-07-13 - #78
Merged
Merged
Conversation
…hot caveat Also notes the 07-17-snapshot nature of the counts (#942 was since reopened for re-review; late auto-applies roll into the next window).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Weekly review-prompt calibration for the week of 2026-07-13, synthesized from the pre-fetched raw verdict-labeled issues in HarperFast/ai-review-log. First sweep run from the GitHub Actions runner (successor to the disabled claude.ai routine). The curated calibration/false-negative log supplements were empty this week; the signal arrived as a single bulk triage batch (all verdicts applied 2026-07-16/17) rather than daily sweeps, and 13
verdict:pendingissues remain open, so this is a partial read of the week.Verdict mix
78 issues triaged: useful 41 · noise 10 · partial 27 (53% useful) — more partial-heavy than recent weeks (83–85%), consistent with a back-log catch-up batch weighted toward the harder, second-reviewer-corroborated cases.
Per model:
Per prompt ref: single-ref week — every Claude entry ran at
9cf49d2(baseline phase); no before/after read possible. (Four entries carry noPrompt reffield — three gemini noise runs + one gemini partial.)Both Claude cells clear the ~15-entry floor, but the sonnet-5 vs sonnet-4-6 partial-rate gap (22% vs 54%) is confounded — the models reviewed different PRs (sonnet-5 → recent
harpercore; sonnet-4-6 → mostlyharper-pro+oauth), not a controlled A/B — so it must not be read as a model-quality signal. gemini (10, below floor) is provider-harness output, out of scope for Claude-prompt edits.Recurring NOISE patterns → none to add
9/10 noise calls are the already-covered vacuous "no blockers" on a no-reviewable-surface diff: CI
retention-days(#1000), dep/lockfile repin (#922), Dockerfile bump (#932), docs-only (#935), + five geminioauthruns (#974, #972, #970, #934, #933). All named by the existing "Mechanical, no-logic diffs" bullet.The one distinct call: a confidently-wrong false positive — a sonnet-4-6 run blocked
sync-core.sh:50claimingnode -e'sprocess.argv.slice(1)is off-by-one and re-raised it despite @kriszyp's empirical counter-evidence (#947, human-confirmed noise). Single, narrow point → watch-listed.Recurring PARTIAL / FALSE-NEGATIVE patterns → none to add (all covered or pending)
Three corroborated clusters, each already named by current or in-flight prompt text:
Number(' ')→0 (#942),Number()on ISO/boolean@expiresAt(#975), unvalidatedNumber()parse (#954),if(-1000)truthy →setTimeout(#982),'0'-is-falsyPARK_WARN_MS(#995), truthiness dropping explicitnull→!= null(#967). Exactly theuniversal.md"Falsy vs nullish guards" + "Numeric coercion and range" bullets, both shipped within the last two cycles (ref9cf49d2) — same-week recurrences of freshly-live text. Reinforce, no new text.Promise.allpartial-startup teardown + unhandledfetch(#992), bare.pipe()wherestream.pipeline()was needed (#953). Covered by "Setup/teardown that leaks on partial failure" + "Error and abort paths that crash or leak". The.pipe()→pipeline()idiom (1 pt) → watch-listed.oauth-heavy, human-HIGH): limiter before auth → unauthenticated quota exhaustion (#930), cleartexthttp:issuer + non-atomic JTI replay (#924),try/catchswallowing a missingnm→ ABI gate bypassed fail-open (#937). Overlaps the still-open, unmerged #73 ("Severity deflation on security/hardening", "Fail-closed / validate-at-registration") — those bullets are not yet in ref9cf49d2, which is why the reviews missed them. Per the "don't duplicate an open calibration PR" rule: noted, not duplicated.Corroborated-but-covered / one-offs (no edit): partial optional-chain
TypeError(#957) & non-array crash (#945) → "Unvalidated shape"; test-validity under-calls (#952, #955) → "Test validity" meta-check; multi-trailing-slash normalization (#939, 2nd point for the 06-22 param-routing watch-list); singletons (#946, #944, #948, #981); #976 is correctly a real-finding-left-unresolved, not an FN. Three gemini partials (#964, #940, #923) are provider output, out of scope.Changelist
CALIBRATION.md: prepended the## Week of 2026-07-13entry.universal.mdtext or by the still-open #73 security-deflation / fail-closed additions; the misses are same-week recurrences of freshly-shipped (or not-yet-merged) guidance, not new blind spots. Conservative call per the rubric — watch whether the mix moves once that text has run a full cycle and calibration: week of 2026-06-29 #73/calibration: week of 2026-07-06 #77 merge.Open calibration PRs noted (neither duplicated nor contradicted): #73 (2026-06-29,
universal.md, unmerged), #77 (2026-07-06, log-only).🤖 Generated with Claude Code