- Add private, source-bound QA evidence pins for runtime errors, assertion violations, and approved goal witnesses. Humans and agents can recheck one exact indexed witness after an edit without another search run; pins intentionally retain no story prose, choice labels, observed values, or report content. A pin is evidence memory, not a replacement for a new broad bounded check after meaningful edits.
- Make hosted progress preserve the last measured work window while the checker moves through later phases or reconnects. The page now updates that evidence in place when newer work arrives instead of replacing it with a generic interstitial; cancellation and uploaded-file deletion language remain explicit.
- Add a conservative forced-choice-cycle specialist. It reports an exact replay path when the only offered choice returns to the same author-visible control state, prunes that branch across sequential DFS, beam, and random passes, and stops a whole portfolio only for the proven direct root-cycle case. Turn, visit-count, random, and EXTERNAL-sensitive stories remain outside this specialist's claim; changing counters and optional exits remain unflagged.
This release makes safe author-defined story rules available in the hosted checker without changing Inkcheck's bounded-search promise. The new web rule builder adds one explicit numeric invariant to an ordinary hosted run; it does not use AI, infer author intent, or spend hidden directed-search budget.
- Add a hosted typed-rule request contract. A browser can submit up to twelve safe rules using the same non-executable assertion grammar as
inkcheck.yml, CI, and MCP; malformed, empty, or oversized rule sets fail before exploration. - Add the first hosted author flow for one optional numeric rule: detected
VARnames assist input; the writer chooses a comparison, a numeric literal or second variable, andalwaysorterminalscope; the exact typed rule is previewed before the run. - Generate a private temporary
inkcheck.ymlonly inside the hosted job directory, reuse the CLI's existing validation/exploration/report pipeline, and delete that configuration with the uploaded files after every result, error, or cancellation. - Surface rule violations as report findings with observed values and exact indexed replay witnesses. A quiet bounded run remains
not_observed, never verification; only an exhaustive pass earnsexhaustively_verified. - Keep assertion-directed extra search experimental and opt-in. It cannot become a default, receive scorecard credit, or make a coverage-improvement claim until the separate preregistered specialist-promotion evaluation passes.
This stable release ships Inkcheck's anytime QA foundation without claiming that bounded search can cover every large story or that an experimental allocator has become universally better. The established deterministic portfolio remains the search default. Workload-aware concurrency, durable result windows, campaigns, exact replay, and human/agent controls make long checks more useful and inspectable; dynamic long-tail expansion, rotation, and stopping remain shadow-only after the three-family promotion gate did not establish broad author-facing value.
- Run the local CLI and one-shot MCP portfolio with workload-aware automatic concurrency. A reusable 1,024-state live pilot activates workers only for sustained eligible workloads; small, exhaustive, depth-bound, saturated, constrained, shared-search, and additive-goal work stays sequential. Fixed concurrency and a hard single-worker opt-out remain available.
- Add durable, source-bound local reports, compact checkpoints, saved-finding lookup, exact witness replay, regression pins, and resumable MCP search sessions. Partial evidence survives cancellation and declared resource boundaries, while stale or incompatible capabilities fail closed.
- Add deterministic campaign budgets and immutable result windows for agents and humans. Named Quick, Balanced, Deep, Overnight, Campaign, and Fixed modes enforce aggregate ceilings, protected regression/long-tail reserves, source invalidation, bounded forecasts, and attributable allocation/stop reasons.
- Add explicit assertion and approved-goal campaign children without mutating or reducing the exact base frontier. Campaign-new evidence is identity-deduplicated while complete child provenance remains available in separate reports.
- Add independent protected long-tail portfolio children plus observed-versus-campaign-new yield, rediscovery, and discovery-spacing evidence. The checked three-family gate found resource and terminal-diversity gains on The Intercept but no new runtime error, assertion violation, approved goal, authored knot, or visible outcome; Dog Ink Adventure and Heresy II reached a time or memory bound before long-tail allocation. Live allocation therefore remains unpromoted.
- Add privacy-minimal MCP defaults, a packaged Inkcheck agent skill, ten golden QA exercises, and an executable no-hidden-hints readiness scorer. Full story content and witnesses remain explicit drill-down boundaries; Inkcheck itself remains deterministic and non-AI.
- Add trustworthy human, NDJSON, MCP, and hosted progress that distinguishes work-budget use, discoveries, binding limits, cancellation, and final outcome. Hosted jobs retain privacy-safe progress across restart without retaining uploaded story source.
- Preserve the stable product boundary in documentation and the product/engineering scorecard: anytime value is currently 7/10 and demonstrated generalization 4/10. Adaptive first-window sizing, compact large checkpoints, provider-attributed cost, broader agent evaluation, and bounded specialist explorers remain future roadmap work.
This prerelease turns the qualified workload classifier into the local CLI and one-shot MCP portfolio default, while retaining explicit sequential/fixed controls and the hosted deployment ceiling. It also packages the post-beta.2 agent-readiness, compact-output, checkpoint, campaign-specialist, and progress-trust work. Automatic concurrency changes execution timing, not search allocation or bounded-coverage meaning. It is published under npm's next tag; stable latest remains 0.5.1.
-
Promote workload-aware concurrency to the local CLI and one-shot MCP portfolio default (#169).
autouses the production-eligible live-pilot handoff with a conservative four-lane ceiling; small, exhaustive, depth-bound, authored-frontier-saturated, one-core, memory-constrained, shared-search, and additive-goal work stays sequential. Explicit1is a hard opt-out, explicit 2-16 ceilings preserve fixed concurrency, hosted jobs retain their separately configured explicit ceiling, and compact machine output exposes the full versioned activation decision without authored content. -
Reuse workload-aware activation work instead of replaying it (#174). A deterministic inside-out DFS pilot now forms the prefix of the ordinary first portfolio round, stays live in the parent process, and either continues sequentially or overlaps untouched persistent-worker passes under one global state and memory envelope. The repeated 80-cell 100K gate retained exact evidence, schedule, and proof with zero duplicate evaluations; matched 5M The Intercept runs retained exact evidence at both depths, correctly stayed sequential at depth 30, and reduced depth-100 wall clock from 657.7s to 489.0s while moving first meaningful evidence from 88.9s to 0.4s. This cleared the executor's production-eligibility gate for the #169 integration above.
-
Add a research-only 1,024-state workload-aware concurrency activation evaluator. It rejects exhaustive, depth-bound, and authored-frontier-saturated pilots, reports duplicate work and high uncertainty, and remains explicitly ineligible for production until pilot state can migrate without exceeding the user ceiling.
-
Add explicit terminal progress status, binding stop reason, and result outcome across NDJSON, human terminal, and hosted jobs. Exhaustive completion, state/depth/time/memory/frontier/worker limits, compile failures, cancellation, service restart, and unexpected errors remain distinct from whether the report found runtime/assertion issues. Unexpected CLI failures now emit a best-effort terminal error event; responsive signal cancellation remains #37.
-
Persist hosted async-job status and privacy-safe progress events in an optional private TTL-backed file store. Production Compose enables it in the existing volume; restarts preserve counters and return an honest retry state without retaining uploaded source, filenames, findings, or reports, and expired or malformed records are purged.
-
Stream privacy-safe live discovery events for CLI, human-terminal, and hosted progress consumers. Events expose only monotonic numeric deltas and cumulative counts for endings, runtime errors, visited knots, assertion failures, goals, stages, and visible outcomes; final reports remain authoritative for identities and story content.
-
Extend the promotion harness with matched sequential-versus-concurrent portfolio runs, an explicit 2-16 candidate worker ceiling, and result-window time-to-1/5/10 meaningful-evidence milestones. Runtime errors, assertion violations, authored knots, and visible endings earn timing credit; raw terminal multiplicity and worker heartbeats do not.
-
Enforce concurrent exploration's declared heap envelope across the parent and all worker isolates. Workers publish heap use through shared memory, the parent triggers one cooperative memory stop when aggregate use binds, partial evidence survives, and machine reports expose planned heap shares plus the observed aggregate high-water mark.
-
Give hosted deployments an independent
INKCHECK_WEB_PORTFOLIO_CONCURRENCYceiling (default 1, maximum 4), separate from simultaneous-job concurrency and the local CLI ceiling. -
Replace #94's failed fixed concurrent allocator with persistent worker-owned pass engines across the production adaptive rounds. A matched 5M The Intercept verification preserves identical terminal, runtime, knot, and complete-schedule SHA-256 digests while reducing wall clock from 601.1s to 369.5s (38.5%) at 801 MiB versus 546 MiB peak RSS (46.8% higher). The fixed executor is removed; concurrency remains opt-in at one by default pending the broad #56, cancellation, hosted-cap, and aggregate-resource gates.
-
Add the first #94 worker-backed portfolio executor behind explicit CLI/config/MCP concurrency. It uses rolling bounded worker slots, deterministic weighted grants and canonical merge order, aggregate budget heartbeats, state/memory/deadline controls, safe constrained-machine fallback, compact execution evidence, and a distinct worker-failure partial-report contract. Concurrency remains one by default: matched fixed-allocation 5M The Intercept candidates lost 42.4% of terminal identities, and the shipped rolling form was slower than the adaptive sequential baseline. Adaptive concurrent allocation remains required before default promotion.
-
Add the versioned agent-readiness benchmark foundation: a deterministic unfamiliar-project runtime/assertion repair fixture, no-hidden-hints protocol, checked expected evidence, machine scorer, bootstrap/call/reference/safety/proof targets, and separate tool/skill/model/environment failure attribution. The default five-tool MCP profile and bundled skill measure 2,738 bootstrap tokens under the preregistered byte estimate, while
INKCHECK_MCP_PROFILE=fullpreserves named-tool compatibility. Real results from two distinct agents remain a release gate rather than synthetic evidence. -
Make MCP inspection, compilation, statistics, and one-shot exploration compact and privacy-minimal by default with explicit drill-down/full detail, source-bound inventory pages, and incremental session-event cursors. Raise the bounded campaign-window ceiling to 1,024 while retaining byte/resource guards.
-
Ship a versioned Inkcheck agent skill in the npm artifact. Its compact inspect-compile-search-replay-fix-verify loop uses progressive Ink and finding references, preserves author intent and bounded-evidence language, and includes ten golden QA exercises spanning compile errors, runtime exhaustion, assertions, state, turns, randomness, externals, unreachable content, and stale witness paths.
-
Stabilize runtime finding IDs across search strategies and witness shortening. Approximate source mappings remain visible metadata but no longer participate in identity; exact locations may still distinguish otherwise identical generic failures. Promotion evidence now separates semantic runtime retention from approximate location drift so a line-mapping difference cannot become a false critical regression.
-
Store new exact shared-search checkpoints as compact streamed gzip artifacts while retaining backward-compatible reads of schema-v1
.jsonfiles. Stable checkpoint IDs and logical resume state are unchanged; the writer enforces ceilings against compressed durable bytes without constructing one duplicate artifact-sized JSON string, and corrupt compression fails closed before envelope validation. -
Add explicit assertion and approved-goal campaign children. Matching MCP campaigns can spend separately accounted root-started specialist windows without mutating or reducing the exact resumable base frontier; deterministic ledger purposes, source-bound child reports, hard campaign resource ceilings, privacy-minimal metadata, and cross-report evidence deduplication keep the work attributable. Exhaustive base campaigns can still run an explicitly requested specialist while retaining the base proof boundary.
-
Add the first authored-story campaign-child evaluation harness and checked-in 5M-grant evidence. The study records blocked resource cells instead of treating them as misses, separates specialist intent/critical yield from broad evidence deltas, and documents one credible Intercept assertion lead, one rejected rule hypothesis, one staged-goal miss, and current checkpoint/specialist economics.
-
Let explicit specialist children run after a protected base closes at its state ceiling, while time, memory, disk, deadline, cancellation, and invalidation stops remain binding. Oversized or unstringifiable shared checkpoints now retain a partial report instead of throwing
Invalid string length; campaign workers reserve commit headroom and may persist a measured over-ceiling peak only when the allocation explicitly reports a memory stop.
This prerelease adds policy-bound campaigns for agents and humans without changing the fixed portfolio default. Campaign forecasts and knee observations remain bounded planning evidence, never coverage claims. It is published under npm's next tag; stable latest remains 0.5.1.
- Add the deterministic v0.6 campaign-policy and aggregate-ledger foundation: scarce/balanced/abundant postures, protected regression and long-tail reserves, source/config binding, hard resource ceilings, partition descriptors, and explicit bounded-evidence semantics. This does not activate dynamic production allocation or add a campaign execution surface yet.
- Add durable MCP campaign result windows with
start_campaignandcontinue_campaign. Campaigns reuse the exact shared frontier, survive fresh processes, enforce aggregate state/time/deadline/memory/disk limits between windows, retain immutable report/checkpoint provenance, invalidate safely on source edits, and preserve the latest partial report on cancellation or a hard boundary. Independent child strategies, concurrency, and cross-run finding merge remain unavailable. - Add versioned MCP campaign controls and compact decision explanations. Agents can choose
quick,balanced,deep,overnight,campaign, or backward-compatiblefixedmodes; override bounded resource, value, stop, and ceiling fields; and inspect stable policy IDs, allocation reasons, preferred-yield rates, throughput/resources, uncertainty-labelled next-window ranges, knee evidence, binding constraints, and report IDs for drill-down.kneestopping requires three dry preferred-yield windows and still honors protected long-tail work. Forecasts remain empirical observations over one exact shared trajectory, never coverage or asymptote claims. - Add
inkcheck campaignfor Quick, Balanced, Deep, Overnight, Campaign, and Fixed human intents. It emits immutable source-bound result windows with stable finding IDs, work/yield/forecast evidence, trigger reasons, and continuation state, and preserves the latest partial report when a deadline or between-window cancellation stops the campaign. - Add hosted Quick and Balanced intents, uncertainty-labelled progress, completed result-window metadata, and source-bound cancellation windows. Hosted cancellation retains final progress counts and explicit omitted-finding counts without retaining the story report or inventing finding identities.
This prerelease packages the 0.6 measurement, bounded-resource, durable-evidence, and interactive-agent foundations for public evaluation. It is published under npm's next tag; stable latest remains 0.5.1. The fixed portfolio remains the production default, and the anytime allocation policy remains shadow-only because current promotion evidence shows baseline parity rather than a broad advantage. Campaign intents, deadlines, concurrency, and specialist dispatch remain 0.6 work rather than implied beta features.
-
Add transactional MCP
add_goalprobes (#144/#66). An agent can attach one safe typed or staged goal to a current result-window session and spend up to 5M explicit additional states from the story root. The exact base checkpoint, report, grant, and evidence remain untouched; inspection exposes separate base/directed/total accounting plus bounded opaque probe summaries. Full goal conditions and witnesses live only in the private source-bound goal report and explicit content-revealing response. Base plus cumulative directed grants retain the 100M campaign ceiling. Search-session schema v4 adds goal audit/accounting fields while reading v1-v3 sessions; persisted directed frontiers and in-frontier reprioritization remain deferred. -
Add private MCP runtime regression pins (#142/#66).
pin_regressionturns one current replayable runtime finding into an idempotent session-bound artifact containing indexed choices, story seed, and hashes of the expected runtime errors;check_regressionreplays it after edits without spending search states and returnsfixed,still_failing, orpath_changed. Pins and audit events omit runtime text, labels, transcript, variables, and prose; storage is private, atomic, capped at 1 MiB per pin and 100 pins per project. Search-session schema v3 adds pin/check audit events while reading v1/v2 sessions. Ending and assertion pins remain explicitly unsupported until they have truthful domain-specific replay contracts. -
Add revision-bound MCP
replay_witness(#140/#66). An agent can explicitly turn one stable finding from a session's latest report into a current-source deterministic playtest without loading the full report. Replay fails on stale source/revision, foreign IDs, or unsupported witnesses; successful execution advances the session revision and records only finding/report IDs plus replay status in bounded metadata. Transcript, choice text, and variables are returned only across this explicit content-revealing boundary. Search-session schema v2 adds the audit event while reading and upgrading v1 foundation sessions. -
Add durable MCP result-window sessions (#138/#66).
start_search,inspect_search,continue_search, andcancel_searchlet agents spend base-shared work in synchronous windows, reopen bounded privacy-minimal evidence in a fresh MCP process, and continue the exact source/config-bound frontier without replay. The first window defaults to 1M states, each call may add at most 5M, and the cumulative ceiling is 100M. High-entropy bearer capabilities, revision checks, private atomic metadata, bounded events/files, explicit retained-versus-discarded cancellation, and fail-closed source/capability checks keep this first slice honest; mid-window preemption and assertion/goal/variable-aware sessions remain later #66 work. -
Complete the local report-storage lifecycle (#136/#63): report directories/files are private, durable atomic writes sync before and after rename, and new saves fail cleanly above 256 MiB per report or 1 GiB per project without deleting stable evidence.
artifacts deleteand per-entrypointartifacts prune --keep Nare preview-first, require--apply, and cap each deterministic cleanup batch at 100 reports. Capabilities expose every ceiling. -
Add bounded saved-finding lookup and exact CLI witness replay (#134).
artifacts findingspages through privacy-minimal summaries without story prose, variables, or witness paths;artifacts findingfetches one complete stable finding; andartifacts replayrecompiles current source and follows the saved indexed choices with the recorded story seed. Cursors are report-bound, duplicate IDs fail closed, and stale/path-changed reports cannot execute historical witnesses. -
Persist and resume exact base-shared CLI checkpoints (#132).
--save-checkpointwrites a private, atomic, source/config-bound artifact when live work remains;resume <id> --max-states Ncontinues to a larger total grant in a fresh process and automatically saves the next generation.checkpoints list/showexposes bounded freshness metadata, corrupt/stale/incompatible artifacts fail closed, and deterministic retention keeps three generations per entrypoint under 512 MiB single-file and 1 GiB project ceilings. Capabilities advertise schema v1 and the CLI-only resume surface; unsupported search modes and tracker state are rejected rather than approximated. Shared search now also uses the documented global default depth of 100 instead of its stale internal value of 30. -
Add the first exact-resume foundation for base shared search (#130).
exploreSharedResumable(...)can serialize its complete live frontier to source-bound schema-v1 JSON and continue to a larger total state grant without replaying earlier work. Split runs are regression-tested against uninterrupted execution, including a pause partway through a choice list; changed source/configuration, corrupt references, incompatible schema versions, completed runs, and resource-stopped runs fail closed. Durable CLI files and assertion/goal-aware checkpoints remain separate follow-up work. -
Add opt-in source-bound local report artifacts (#128).
--save-reportwrites a versioned envelope atomically under.inkcheck/reports/and returns a stable content-derived ID;artifacts list/showreopens it in a fresh session and labels the evidencecurrent,stale, orpath_changedagainst the present entrypoint. Corrupt, tampered, and incompatible artifacts fail closed. Capabilities and agent docs distinguish available report persistence from still-unavailable resumable search. -
Bound shared-search retained memory without imposing a low universal cap (#98). Expanded checkpoint JSON and dead witness ancestry are released when no pending descendant needs them, stale policy views compact deterministically, and pass telemetry separates pending/active payload, ancestry, indexes, references, and findings from process heap/RSS. Optional CLI/config/MCP checkpoint count and byte envelopes preserve partial evidence and report the distinct
truncatedBy.frontiercause. Adversarial low-dedup/deep-branching ladders and matched 64/128 MiB Intercept envelope cells document scaling, collection, and clean stops; disk spill is gated on broader evidence rather than assumed. -
Make Ink runtime randomness reproducible under an explicit initial seed (#117).
--story-seed, project config, MCP exploration/playtest, capabilities, reports, replay instructions, and the promotion harness now distinguish Ink's runtime RNG from the existing--seedsearch-sampling control. Both default to 1; authoredSEED_RANDOM(...)still works, and reports remain honest that one run does not enumerate every story seed. -
Preserve production allocation whenever policy-v2 replay is warming up or gated off (#118). The cumulative integer floor now begins only when a previously approved policy overlay controls a window, so replay cannot change evidence while reporting
allocationApplied: false; policy-controlled windows retain exact floor accounting. The 100-state/depth-300 deep-chain pair now keeps all seven baseline knots for both seeds. -
Add the #56 search-promotion harness and a 20-family consent/license manifest. One command runs isolated fixed-portfolio versus policy-replay matrices across budgets, depths, and seeds; passes real typed assertion rules into both strategies; records stable critical/authored/terminal identities, proof/truncation, pass/frontier telemetry, repeat determinism, elapsed time, and peak RSS; emits JSON, Markdown, or timing-free deterministic output; and highlights worst-family regressions without declaring a winner. The first 240-pair matrix is checked in as a concise evaluation: 222 pairs were neutral, assertion evidence matched, policy replay had one low-budget deep-suffix regression and no broad largest-budget gain, so default allocation remains frozen.
-
Replace shadow policy v1's absolute 1,000-state recency grace with deterministic scale-normalized renewal (#113). Policy v2 measures marginal yield against recent grant/consumption windows, gives signals a one-to-two-window horizon, requires a three-window warm-up, bounds the experiment slice, and applies replay overlays only for renewed critical evidence or explicit goal progress. Early-choice 100/500/2,000 pairs now match the fixed portfolio; 100K/1M/5M high-water pairs remain neutral at 45 endings and 22 knots; small combination-lock exhaustion proof is preserved. Production allocation remains unchanged pending #56.
-
Give research policy replay auditable cumulative integer probe floors (#106). Fractional promises pool across windows, whole states rotate to the largest service debt, completed passes release future service, and every replay round records planned grants plus cumulative promise/grant/debt accounts. Tiny-window and exact 5M accounting tests pass; production allocation remains unchanged because the early-choice regression (#113) still blocks promotion.
-
Add portfolio-marginal value curves for #105 while retaining pass-local diagnostic curves. Runtime/assertion credit is identity-based, approximate runtime fallback lines are normalized for allocation credit, and exact endings/outcomes/knots/goals/stages are paid once in actual scheduler order. Policy replay now reads marginal evidence, removing its matched 50-state sparse-error and late-ending regressions; early-choice authored-coverage losses remain and continue to block promotion.
-
Add a living product/engineering truth scorecard with explicit 10/10 targets, candid current ratings, evidence/gaps, and a release reassessment protocol. The README now states that planned value comes from a reproducible broad hybrid plus bounded Ink-aware specialists, and corrects the current 8% pass floor from a “guarantee” to a fractional intent whose small-window service debt remains #106.
-
Add the first opt-in #103 policy-applied replay harness. It uses the existing portfolio engines and deterministic windows, applies only recorded
reallocateshares, logs every decision/allocation, and leaves the default scheduler unchanged. A paired sparse-error fixture documents a critical v1 regression at 50 states: the fixed portfolio finds the runtime error and proves exhaustion while the candidate misses it and remains partial, before recovering at 75 states. This evidence blocks promotion and identifies duplicate pass-local credit and small-window floor rounding as follow-up defects. -
Add a manifest-driven shadow-policy budget-ladder evaluator for #56. It compares exact runtime, assertion, knot, visible-outcome, and terminal-state identities against a declared high-water run; classifies stop risk without an aggregate score; emits stable JSON or Markdown; requires source license/consent metadata; and explicitly treats independent larger runs as bounded comparisons rather than continuation prefixes or coverage oracles. It does not activate policy decisions or change search behavior.
-
Add the v0.6 anytime decision engine in strict shadow mode (#92). JSON and MCP reports now include a deterministic, versioned continue/reallocate/probe/stop recommendation with explicit evidence, uncertainty, binding constraint, and protected per-pass probe floors. Findings use lexicographic value tiers instead of an opaque score, with runtime/assertion evidence kept highest and separately visible. The recommendation is never applied (
mode: shadow,applied: false), so this release gathers auditable policy evidence without changing search allocation, stopping behavior, findings, or coverage claims. -
Complete the factual #91 curve contract with explicit marginal deltas and internal unique-state novelty. Dedicated fixtures cover no discoveries before a bound, increasing discovery gaps, and a long dry interval followed by late recovery; wall time remains observational progress telemetry rather than contaminating deterministic curves.
-
Add deterministic discovery summaries for 0.6 shadow-mode consumers (#91): event count, first/latest discovery positions, current dry distance, latest gap, and longest observed gap survive bounded curve compaction and stream through NDJSON. These are measured facts only; no plateau, knee, or stopping inference is introduced.
-
Extend 0.6 discovery curves with meaningful QA value classes and actual portfolio order (#91): exact terminal states remain separate from normalized visible outcomes, while assertion violations, reached goals/stages, runtime errors, and authored knots have independent counters. Portfolio reports retain a bounded merged curve, and NDJSON progress streams the new privacy-safe counts. Scheduling and stopping remain unchanged.
-
Begin the Inkcheck 0.6 anytime-decision measurement foundation (#91): every exploration pass records a deterministic discovery curve with separate ending, runtime-error, and authored-knot counts. Curves compact to at most 64 samples regardless of state budget, preserve early/latest evidence and dry-gap measurements, and do not yet alter scheduling or stopping behavior.
-
Stop hosted
humanFindingsfrom advising CLI flags a web user cannot set (#49).buildHumanFindingstakes anaudience: "cli" | "hosted"option; the hosted server passes"hosted", so the limit-bound unvisited-knot next step now says to run inkcheck locally for a deeper check instead of naming--max-depth/--max-states. CLI output is unchanged. -
Add safe typed search goals with an explicit additive budget. General exploration keeps the full
maxStatesallocation;goalMaxStates/--goal-statesoptionally adds deterministic goal-proximity work and defaults to zero. Reports, progress events, config validation, discovery, CLI, and MCP expose baseline, goal, and combined budgets, with a shared 100,000,000-state ceiling. -
Add ordered staged goals for late variable dependencies. Each cumulative milestone has a deterministic witness and status; a bounded prerequisite miss blocks downstream stages instead of claiming they are unreachable. Stages share the explicit additive goal budget and never displace baseline exploration.
- Add safe author-defined story assertions (#64): versioned config and structured MCP input accept typed variable/literal comparisons composed with
all,any, andnot, scoped always, at terminal states, or on entering a named knot. Every search engine evaluates the same prevalidated rules on visited states; violations fail CI and carry stable IDs, observed values, exact indexed witnesses, and replay operations. Reports distinguish violations, bounded runs where no violation was observed, and exhaustive verification. No JavaScript, shell, or arbitrary Ink expressions are accepted. - Add strict project configuration schema v1:
inkcheck.ymlcan commit a project-relative entrypoint and bounded CI defaults,inkcheck validate-config [path] [--json]reports actionable path-specific errors, explicit CLI flags win, and unknown keys fail rather than pretending future assertions/goals are active. The JSON Schema is packaged andcapabilitiesnow advertises config schema v1. - Add idempotent project bootstrap commands:
inkcheck initcreates a validated minimal config after unambiguous entrypoint discovery, andinkcheck agent-kit --format codexscaffolds pinned CI, generated-artifact ignore rules, and compact version-matched agent instructions. All writes are preflighted; a conflicting target aborts the operation before any file is created or overwritten.
- Report portfolio-wide cumulative counts in live progress (#55). The portfolio scheduler emitted each progress event from the just-run pass's snapshot, so the endings/errors/unvisited-knots counts a consumer showed as a running total bounced up and down as passes interleaved. Progress now comes from the scheduler's portfolio-wide dedup sets, which are monotonic by construction: endings and runtime errors only rise, unvisited knots only fall, within a run. Fixes the hosted web progress indicator and the CLI
--progress=human/ndjsonoutput together. - Raise the default choice-trail depth from 30 to 100 (
DEFAULT_MAX_DEPTH), shared across the exploration engines, the shape profiler, the CLI, and the MCPexplore_storytool. A depth of 30 cut off ordinary stories out of the box — The Intercept's paths pass ~19 choice-bearing knots and real playthroughs go deeper — so the plaininkcheck story.inkrun under-explored unless the author knew to reach for--auto. 100 is generous headroom over real story depths; a story deeper than that still truncates cleanly withtruncatedBy.maxDepthand adeepenrecommendation, and--autoraises the limit per-story from the shape profile. Explicit--max-depthstill wins and--autostill never lowers a limit. (examples/deep-chain.inkis now 130 knots deep so it keeps demonstrating that plain defaults truncate while--autoreaches the ending and proves exhaustiveness.) - Stop the hosted web checker from timing out on normal-sized stories and returning a misleading "your story is so detailed and long" limit error (#71). Three changes: the hosted
--max-depthdefault drops from 1,000 (the system's escalation ceiling, which let one deep-loop trail consume the whole state budget) to 100, ample headroom over real story depths — The Intercept's deepest pass reaches depth 65 and its result is byte-for-byte identical at depth 100 vs 1,000; the graceful--max-timenow reserves a real margin below the hard SIGKILL deadline (15%, at least 30 s, vs a fixed 10 s that was too tight to flush a multi-MB partial report), so the partial report the engine already computed is returned instead of discarded; and the hosted timeout ceiling drops from 450 s to 300 s — a 5-minute cap that returns a strong partial report beats a 7m30s wait that returned nothing. A genuinely wedged run that still has to be killed now gets an honest time-limit message that does not blame the story's size. - Add versioned agent report schema v1 (#59): JSON reports now identify Inkcheck/schema versions, fingerprint the compiled story, record effective configuration and binding limit, and enrich compile/runtime/ending findings with stable IDs, normalized kinds, suggested actions, and documentation identifiers. Every explored ending/runtime witness carries zero-based
choiceIndicesalongside human choice text and exactplaytest_storyreplay instructions; duplicate labels are unambiguous, and playtest reportspath_changedfor stale indexed witnesses. - Add versioned agent discovery foundations (#58):
inkcheck capabilities [--json]and MCPinkcheck_capabilitiesexpose schemas, limits, modes, and explicit unavailable feature flags;inkcheck inspect <story.ink> [--json]and MCPinspect_storyreturn a deterministic, bounded, source-only project map without compiling or exploring. Inspection follows project-local includes, reports shape/semantics/externals/knots/variables, and rejects missing or outside-root includes. - Add opt-in
--search=shared-variable(and matching MCP mode), which dedicates 12.5% of shared-frontier selections to uncommon observed variable snapshots and transitions while retaining novelty, deep, and seeded views. The heuristic is deterministic, bounded, and mechanical; it does not interpret variable meaning. A checked-in comparison table records both improvements and regressions, so portfolio remains the default. - Document the search-strategy promotion policy: default portfolio allocation is frozen until an alternative passes a broad predeclared matrix with structural-family regression gates, multiple budgets/depths/seeds, resource measurements, and no lost runtime-error or assertion evidence at the largest comparable budget.
- Add opt-in
--search=shared(also available as MCPexplore_story.search) for an experimental shared-state multi-frontier engine. Deep, novelty-first, and seeded views share one deduplicated graph and expand each state at most once while compact parent links preserve repro paths. New pass telemetry records unique states, peak pending states/bytes, variable states/transitions observed, and rare variable transitions. The existing adaptive portfolio remains the default while benchmark evidence accumulates; variable rarity is observed but does not steer search yet. - Add a wall-clock time budget so a slow run degrades gracefully instead of being killed.
--max-time <s>stops exploration cleanly at the deadline and returns the partial report it has (newtruncatedBy.timecause), mirroring the memory guard for time. The hosted web checker now passes--max-timejust under its hard SIGKILL timeout, so a story too slow to finish returns a partial report with its findings-so-far instead of failing with a timeout and losing them; the hard kill remains only as a backstop for a wedged process. On a time stop thenextRunverdict isinvestigate(raise--max-timeor use the local CLI), neverbroaden. - Drop internal "beam width" / "frontier cap" jargon from human-readable truncation advice. A pure beam prune still means the story is bigger than the run covered — which the depth/state hints and the states-explored count already convey — so
truncationAdviceno longer surfaces the beam's internal cap in--human/text/Markdown reports (or in thehumanFindingsthe hosted web checker renders).
- Raise the state-budget ceiling to 100,000,000 and the CLI/MCP default to 10,000,000 (from 1,000,000 / 100,000). The large default is safe because exhaustible stories early-exit, the memory guard stops cleanly before an out-of-memory crash, and progress reporting lets you watch and interrupt a long run — so the practical limiter on a big run is memory (or the wall clock you choose), not the ceiling. The hosted web checker keeps a 1,000,000-state default and cap so one story cannot monopolize the shared server; larger jobs belong on the local CLI. Pin
--max-statesin CI when a bounded runtime matters more than depth of coverage. - Add a Performance and memory section to the README: per-million-states timing, the memory-term breakdown (dedup hash floor ~200 B/distinct state and dominant; DFS/beam/random flat; BFS repro frontier the one super-linear risk), a "~2 GB heap per 10M distinct states" worst-case rule of thumb, and the levers (
--max-old-space-size,--no-min-repro,--max-states) for a story too big to finish.
- Add a memory guard so large runs degrade gracefully instead of crashing. A V8 heap out-of-memory abort cannot be caught after the fact, so exploration now watches heap use and stops cleanly before the wall, keeping every finding so far and reporting
truncatedBy.memoryin a partial report. The cap defaults to 85% of the V8 heap limit (honoring--max-old-space-size) and is overridable with--max-memory <mb>; the guard is active in the CLI and the MCPexplore_storytool. On a memory stop thenextRunverdict isinvestigate(raise the heap, lower--max-states, or split the story), neverbroaden, since more budget would hit the wall sooner. - Recommend the next run (#30): every report now carries a
nextRunverdict from a small closed vocabulary —stop,deepen,broaden,reseed,investigate— computed as a pure, deterministic function of the report (plus the static shape profile for the deepen target), with ready-to-use flags, a rationale citing the report fields that drove it, and an evidence-backed expected gain. Proposed flags never exceed the hard ceilings; when no increase has evidence behind it, the verdict degrades toinvestigateand points at the unvisited knots worth reviewing. Available in CLI--json, the MCPexplore_storytool, and as one-line advice in text and Markdown reports. - Add
--next(#30): after the check, the CLI applies the recommendation and reruns automatically — up to three escalations, stopping on astop/investigateverdict, at the flag ceilings, or when an escalated run finds nothing new (fixpoint). On the 40-deep chain fixture,--nexttakes a defaults run that found nothing to a proven-exhaustive result in one automatic escalation. The per-run trail is recorded in--jsonoutput asruns; hop narration goes to stderr so machine output stays clean. - Add truthful human progress for local terminal runs: interactive terminals show the current phase, configured work-budget use, discoveries, throughput where stable, and elapsed time;
--progress=humanprints log-friendly snapshots, while--progress=offremains silent. Progress explicitly describes work budget rather than story coverage. - Make hosted checks asynchronous with private short-lived jobs, real CLI-derived progress events, SSE reconnect plus status polling, and server-side cancellation. Uploaded source is still deleted after success, cancellation, timeout, or failure; retained job metadata contains only progress/report data and expires quickly.
- Emit lifetime per-pass telemetry in JSON reports (#28): a
passesarray with, per pass, states explored vs granted, own finding counts, portfolio-marginal first discoveries (consistent with the schedule's per-round sums), dedupe hits, max depth reached,lastDiscoveryAtState(the cheap discovery-curve signal), truncation causes, and exhaustiveness — plus peak frontier size and prune count for the beam. Standalone pass runs and the CLI's BFS repro slice attach their own entry, so--jsonconsumers see every pass that contributed to a report without parsing progress logs. Telemetry reports facts and leaves stop/continue judgments to the consumer: a long gap since a pass's last discovery does not prove the pass is done.
- Spend the exploration budget adaptively (#29): portfolio passes now run interleaved in ten deterministic rounds, with each round's grants reallocated toward passes whose findings are still growing (guaranteed per-pass floor so dry spells never defund a pass), and the whole portfolio stops the moment a systematic pass proves every reachable state visited — a small fully-explorable story at the default 100,000-state budget now finishes in the ~10 states it actually has instead of resampling for the full budget. The executed schedule (grants, consumption, and marginal discoveries per pass per round) is recorded in
--jsonoutput. Runs remain fully deterministic; a budget-bound random slice is now always reported truncated (sampling never proves completeness) unless a systematic pass proved exhaustion. - Profile story shape before exploring (#27): new
--profileprints a static scan — variables and where they are assigned, choice density, the longest divert path — plus the depth limit and pass weights inkcheck would choose;--autoapplies them, raising--max-depthwhen static divert paths outrun the default (explicit flags always win, limits are never lowered), dropping sampling passes for variable-free stories, and boosting beam/random weights when variable state is set early. On a 40-choice-deep chain, default settings find no endings while--autoreaches the ending and proves the story exhaustive in ~111 states. Addsexamples/deep-chain.ink. - Make bounded-coverage limits explicit in every report (#22): all output modes state the limits the run used (depth, state budget, seed); truncated runs name the limit that actually cut coverage (
truncatedBy) with targeted raise-this-flag advice, informed by local The Intercept evidence that depth and state budget are separate axes. - Triage unvisited knots with an inbound-divert source scan (#22): each unvisited knot reports
inboundDivertsandstaticOrphanCandidate, so reports distinguish "no authored divert points here — possible orphan" from "has inbound diverts — likely beyond this run's limits", with matching next-step advice in--humanoutput. - Report
exhaustive: truewhen a systematic pass visits every reachable state without hitting a limit, and stop counting sampling-slice budget exhaustion as truncation in that case. This also fixes small fully-explored stories failing--strict(and being reported as partial) because the random slice always spends its whole sub-budget. - Label runtime-error repro paths with the pass that found them in text and Markdown reports (previously JSON-only).
- Add a frontier-capped, diversity-first beam-search slice (~15%) to the exploration portfolio (#21). The beam advances level-by-level like BFS but keeps at most 64 states per level, selected round-robin across variable-signature groups (novelty-ranked within each group: new knots, new variable signatures, new offered-choice sets). It is deterministic without a seed, bounds the frontier memory that made naive BFS impractical on large stories, and marks the run truncated whenever it prunes a reachable state. On the early-choice grid fixture, the beam alone finds all seven endings in ~2.2K states.
- Add a seeded random-sampling slice (~20%) to the exploration portfolio so early-choice state combinations get sampled instead of repeated; the deterministic DFS portfolio alone missed 4 of 7 endings on an adversarial early-choice fixture even at a 1M state budget (#20, #21).
- Add
--seedto the CLI and aseedinput to the MCPexplore_storytool; a fixed default seed keeps CI runs reproducible, and the used seed is reported inexplore.limits.seed. - Label every reported ending and runtime error with the search pass that found it (
foundBy, e.g.dfs:last,bfs,random:seed=1). - Add
examples/early-choice-grid.ink, the community-motivated fixture behind issue #20, plus regression tests that the portfolio now reaches all seven endings within a bounded budget.
- Raise the default exploration state budget to 100,000 and the maximum accepted
--max-statesvalue to 1,000,000 across CLI, MCP, and hosted checker validation. - Add a technical coverage/performance note for bounded exploration. The CLI spends each state budget across a deterministic portfolio of last-choice-first DFS, first-choice-first DFS, inside-out DFS, and a small BFS repro-shortening slice. In local The Intercept runs at default depth 30, higher budgets found more terminal states but still reported truncation: 50,000 states in 9.4s found 7 terminal states and 9 unvisited knots; 100,000 states in 19.9s found 10 terminal states and 9 unvisited knots; 500,000 states in 100.2s found 17 terminal states and 8 unvisited knots; 1,000,000 states in 205.5s found 25 terminal states and 8 unvisited knots. These results document real bounded QA behavior rather than exhaustive verification.
- Map content-exhaustion runtime errors to the authored choice that triggered the dead end when inkjs does not provide a runtime address, so CLI and human reports can include an approximate file/line reference for “ran out of content” failures.
- Add a self-hosted web checker with direct
.inkupload or paste, optional unchangedINCLUDEfiles/folders, consent gates, safe path validation, pilot access codes, rate limits, one-job concurrency, child-process timeouts, and immediate temporary-file deletion. - Support an exact browser-origin allowlist so a static community page can call the checker without opening the API to arbitrary websites.
- Add a hardened Docker/Caddy deployment whose application container has no runtime internet route.
- Document a production budget under $50/month and keep residential Windows hosting for development rather than the public trust boundary.
- Improve bounded exploration coverage with a complementary DFS portfolio and a smaller BFS repro-shortening slice while keeping the same public state limits.
- Document that
--max-statesis a total portfolio budget and that truncated reports can still contain useful endings, runtime errors, and coverage clues. - Keep human reports focused on actionable findings by omitting the hosted truncation coverage note from the finding list.
- Preserve turn and random runtime state whenever the source uses those Ink features.
- Exclude crashing terminal states from successful outcome counts.
- Disclose truncation, random behavior, and every
EXTERNALfunction stubbed to zero. - Make
--strictfail when traversal is truncated or external behavior is unavailable. - Replace ambiguous “ending” counts with distinct terminal-state language.
- Add
--markdownreports for GitHub Actions Step Summaries. - Validate CLI limits and reject unknown options with usage exit code 2.
- Verify downloaded official inklecate 1.2.1 archives with pinned SHA-256 hashes.
- Test Ubuntu, macOS, and Windows.
- Include the InkJam guide and machine-readable manifests in the npm package.
- Lead with mechanical, non-generative QA rather than AI integration.
- Add an InkJam-oriented guide for interpreting errors and coverage limitations.
- Document Ink-Tester as a complementary random-coverage tool.
- Remove an unreproducible large-story result from the README.
- Add published package authorship metadata.
- Initial CLI and MCP release.