Recording the one number that says whether the reminder path works, so the three tiers have
something to beat.
The measurement
Agent 5 of notes/2026-09-02-audit-and-replan.md swept 1456 transcripts across every
project on this machine:
| figure |
value |
| sessions where the checkpoint or prompt hook fired |
866 |
sessions that invoked skill-compounder |
96 |
| sessions that did both |
91 |
| conversion |
10.5% |
This project alone: 249 nudges, 3 invocations.
So roughly nine in ten nudged sessions read the reminder and carry on. That is the
denominator #19 is about: recognition during work is absent, and a reminder injected into
a thread that is deep in another task is mostly not acted on.
What to do with it
- Store the sweep as a script under
scripts/, so the figure is re-derivable rather than
quoted. A number nobody can re-run is not a baseline.
- Re-run it after tier 0 and tier 1 land. A note takes seconds and a reminder arrives
keyed to the moment, so both should convert better than 10.5% or the tiers have not
earned their place.
- Report per project as well as overall. 3 of 249 here against 91 of 866 everywhere
suggests the rate depends on what the session is doing, and one aggregate hides that.
What this does not settle
.claude/CLAUDE.md records the 12-edit and 20-minute hook constants as unvalidated, and
names bin/skillreport as the instrument that would settle them. That instrument cannot
compute its conversion figure at all right now: the hook writes a unary tally the reader
rejects as non-numeric. See the measurement-fixes issue. Until that is fixed, this
transcript sweep is the only conversion number available, and it says nothing about which
edit count is right.
Do not tune the constants against this figure.
Wave 1 of the three-tier epic. Depends on the measurement-fixes issue for anything beyond
the raw rate.
Recording the one number that says whether the reminder path works, so the three tiers have
something to beat.
The measurement
Agent 5 of
notes/2026-09-02-audit-and-replan.mdswept 1456 transcripts across everyproject on this machine:
skill-compounderThis project alone: 249 nudges, 3 invocations.
So roughly nine in ten nudged sessions read the reminder and carry on. That is the
denominator #19 is about: recognition during work is absent, and a reminder injected into
a thread that is deep in another task is mostly not acted on.
What to do with it
scripts/, so the figure is re-derivable rather thanquoted. A number nobody can re-run is not a baseline.
keyed to the moment, so both should convert better than 10.5% or the tiers have not
earned their place.
suggests the rate depends on what the session is doing, and one aggregate hides that.
What this does not settle
.claude/CLAUDE.mdrecords the 12-edit and 20-minute hook constants as unvalidated, andnames
bin/skillreportas the instrument that would settle them. That instrument cannotcompute its conversion figure at all right now: the hook writes a unary tally the reader
rejects as non-numeric. See the measurement-fixes issue. Until that is fixed, this
transcript sweep is the only conversion number available, and it says nothing about which
edit count is right.
Do not tune the constants against this figure.
Wave 1 of the three-tier epic. Depends on the measurement-fixes issue for anything beyond
the raw rate.