This file contains implementation-ready future work derived from the toolkit strategy and ordered by the roadmap. Active repository work belongs in PLANS.md; roadmap outcomes are intentionally not repeated here.
Statuses are ready, planned, parked, conditional, or owner action. Release classification is an expectation to be confirmed when work is activated. Completed work is reflected in the roadmap and stable execution-history archive rather than retained as future backlog.
| Previous item | Disposition | Result |
|---|---|---|
| M2 evidence gaps (BL-008) | Completed and removed | Demonstrated framework resolution, component ownership, scoped completeness, and first-class MCP delta projection ship through shared evidence contracts. |
| M2 framework selection (BL-009) | Completed and removed | Bounded static Jinja2, Go web, Symfony/Twig, and React/Next providers ship with synthetic cross-provider validation and no second MCP. |
| M2 milestone closure | Completed | Stable 0.5.0 delivery and external readback establish M2; BL-010 remains a conditional M7 extension. |
| M3 generic external-MCP qualification (BL-011) | Completed and removed | Checksum-verified OCR 1.9.5, the production ocr-ci review path, a local model peer, and a real stdio MCP peer establish the documented direct-composition safe-use envelope and its claim limits. |
| M3 provider examples (BL-013) | Removed | Direct provider examples over model-selected unrestricted arguments are not the target architecture; valid synthetic adapter qualification moves behind the BL-023 broker boundary. |
| M4 accepted decisions (BL-014) | Completed and removed | Structured target-only decisions preserve deterministic identity, safe applicability, staleness, and bounded projections without suppression authority. |
| M4 project guidance (BL-015) | Completed and removed | Immutable target guidance has bounded discovery, deterministic applicability/precedence, changed-guidance exclusion, and full text through the built-in evidence MCP. |
| M4 milestone closure | Completed | Stable v0.6.0 and protected-target identity readback establish M4 without making later context enrichment part of it. |
| OpenSSF Best Practices publication (BL-022) | Completed historically and not reused | The stable execution history records the passing badge publication and closure; the next identifier is BL-023. |
| Native fuzzing campaign | Retained and revised | BL-019 keeps its activation requirements and adds future M5 parsers, handles, schemas, and hostile adapter responses to its candidate inventory. |
| File-based user configuration | Retained and clarified | BL-020 remains parked; M5 owns only its narrow protected-target context/DLP policy, not a general configuration framework. |
| Additional provider adapters | Retained and clarified | BL-021 remains conditional; future forge parity includes discussion and snapshot capabilities without blocking GitLab-first M5. |
M3 is established. BL-011 is complete and recorded above rather than retained as future work. BL-012 is an authentication extension whose trigger is not met; authentication never substitutes for resource authorization and does not block M3 or BL-023.
- Status: conditional
- Priority: low
- Roadmap theme: M3 External MCP hardening
- Dependencies: Established native remote Streamable HTTP and stdio proxy fallback. The completed BL-011 safe-use envelope applies to any direct provider composition.
- Activation trigger: A supported provider requires authorization-code OAuth rather than static environment-backed headers, and a reviewed stdio proxy is insufficient for pilot operations.
- Goal: Add a provider-neutral authentication lifecycle without placing long-lived OAuth material in repository content or OCR config; object authorization remains server-owned and separate.
- Scoped deliverables: Define authorization-code plus PKCE, browser callback ownership, refresh/revocation, secure token persistence, tenant/resource binding, dynamic-client-registration policy, sanitized audit events, and provider conformance fixtures before selecting an implementation boundary.
- Acceptance criteria: Tokens never enter argv, repository files, generated context, or logs; refresh/revocation and tenant changes fail closed; synthetic static-header and browser-OAuth fixtures preserve native HTTP and stdio fallback.
- Exclusions: Resource authorization by token presence, automatic reference discovery, writes, generic web access, or treating a permanent token as OAuth lifecycle support.
- Validation: Threat-model review plus synthetic authorization, PKCE, callback, refresh, revocation, tenant mismatch, persistence-permission, redaction, and OCR integration cases.
- Release classification expectation:
release-requiredonce public authorization behavior is selected.
- Status: ready
- Priority: high
- Roadmap theme: M5 Bounded review-context enrichment
- Dependencies: Established M1 evidence/MCP architecture, established M3 direct-composition safe-use envelope, established M4 protected-target policy/guidance boundary, and exact qualification of the OCR capabilities consumed by implementation.
- Activation trigger: Owner authorizes M5 implementation. The synthetic pilot must cover forge discussions and one external adapter class; BL-011/M3 are complete. BL-012 does not block activation when reviewed static credentials or a stdio proxy suffice.
- Goal: Deliver one provider-neutral context lifecycle without a second review engine.
- Release classification expectation: Work package A architecture/threat checkpoint is
no-release; runtime implementation isrelease-requiredand closes through one stable product delivery.
- Refine contracts, exact OCR requirements, current/upstream gates, and permitted single-pass model semantics before runtime code; map the planned threat model and evidence traceability to future production owners.
- Separate retrieval, model-egress, publication, and retention decisions. Decide which model-dependent semantics are safe in one review pass and which require a native OCR contextual-adjudication API.
- Cover prompt/indirect injection, BOLA/confused deputy, service-identity mismatch, identity spoofing, poisoning, arbitrary URL/ID traversal, SSRF, malicious schemas/responses, oversized content, denial of wallet, TOCTOU/replay, omission, Unicode/Markdown deception, PII/secret bypass, output laundering, approval/suppression manipulation, existence oracles, and persistent-session leakage.
- Do not implement a second toolkit model loop or merge two independent model reviews.
- Define closed schemas for context records, completeness, receipts, independent budgets, and handles; keep retrieval, model egress, publication, and retention as four separate decisions.
- Use
.opencodereview/review-context-policy.json; absence means disabled. Read it only from the captured protected-target SHA so source changes cannot expand access. - Admit forge fields, provider-declared author classes, origins, tenants/resource classes, projections, budgets, and retention independently. Preserve complete-record and per-thread/age/count/byte/character/aggregate bounds. Keep names, email, avatars, and profile URLs out of model context; use run-local pseudonyms.
- Deterministic reference grammars create candidates only. Policy admission is not object authorization. Unknown identity, authorization, version, DLP, completeness, or OCR capability fails closed and remains visible.
- Keep context budgets independent of repository evidence. Retrieval, model egress, publication, and retention are four distinct policy decisions.
After B, C and D may proceed in parallel.
- Acquire bounded point-in-time snapshots with provider-declared user, automation/service, system, toolkit-bot, or unknown account class; thread/reply identity and order; edit/version/timestamp; anchor; resolution; stale/outdated state; pagination consistency; and visible partial/mutated/unavailable outcomes.
- Preserve existing lifecycle commands, discussion ownership, and suppression as separate consumers. Context projection cannot broaden any of them.
- Any admitted mutable discussion or external context makes automatic approval ineligible.
- Require adapter-owned object authorization for tenant, canonical object, operation, and fields before retrieval and before an opaque handle is minted.
- Normalize and DLP-project bounded responses, then commit them atomically to a run-local context store. Mint an unguessable handle only after successful storage.
- Bind the handle to run, adapter, tenant, canonical object, allowed field projection, version/ETag or digest, policy version, expiry, and stored record without exposing the upstream ID.
- Let the model list/read only minted handles. No generic search, arbitrary URL/ID fetch, redirect, traversal, recursion, or write is available.
- Qualify one synthetic issue or document adapter class in the pilot. Native APIs and arbitrary MCPs remain behind the broker and dedicated least-privilege credentials.
- Depends on B, C, and D. Run OCR with an isolated owner-only home, deterministic session cleanup, and independently bounded context arguments/results. If containment or cleanup fails, block publication unless an explicit secure-debug mode was agreed before the run.
- Project fixed toolkit-authored
context_listandcontext_getdescriptions and closed schemas through the existing toolkit MCP process. The M5 model loop has no upstream search, arbitrary ID/URL, external MCP schema, or external network traversal. - Minimize model egress to policy-admitted data. Publication validation/DLP is separate and cannot reverse disclosure that already reached the model; uncertainty blocks the affected projection.
- Preserve objective findings, calibrate or omit assumption-dependent findings only when OCR can do so safely, keep ambiguity visible, and make partial/mutated/unavailable context incapable of proving absence. Existing suppression,
/ocrcommands, and discussion ownership remain separate consumers. - External context cannot change policy, tools, permissions, lifecycle commands, suppression, posting authority, or approval. Use a native OCR dependency instead of a toolkit-driven second review pass if separate adjudication is required.
- Qualify installed wheel and sdist artifacts, real checksum-verified OCR, stdio protocol, persistence/cleanup, hostile providers, result/publication containment, package/docs, stable publication, and external release readback.
- Exercise an issue-tracker adapter, documentation/wiki adapter, and arbitrary read-only MCP adapter only behind the broker boundary.
- Prove expected automatic outcomes, not an imagined interactive author question, and retain explicit claim limits for model-dependent behavior.
Telemetry is intentionally outside M1 and M5. OCR owns token, cost, budget, provider-level review duration, request, and tool-call telemetry; the toolkit reuses those signals instead of adding another implementation. M6 audits remaining lifecycle, evidence/MCP, context-receipt, posting, and review-value gaps before proposing any provider-neutral toolkit telemetry.
- Status: parked
- Priority: medium
- Roadmap theme: M6 Profiles and quality measurement
- Dependencies: Established built-in MCP lifecycle and OCR per-run model/provider overrides; OCR 1.8.7 satisfies the capability dependency.
- Activation trigger: Not met. Activate only after repeated operations show direct settings are insufficient and the owner approves a closed model/provider matrix plus precedence contract.
- Upstream overlap: OCR 1.8.7 supplies direct run-level selection; OCR 1.9.0 per-file and 1.9.5 aggregate budgets remain explicit completeness controls, not profile defaults.
- Goal: If need appears, offer
economy,standard, andstrongaliases for one OCR run without hiding aggregate, per-file, or tool controls. - Scoped deliverables: Define an owner-approved closed matrix and precedence contract; map an alias to one OCR run; publish effective non-secret identity; validate compatibility and environment precedence.
- Acceptance criteria: One model remains active per run,
standardpreserves current behavior, explicit provider/model settings override profile aliases, aggregate/per-file/tool limits remain independent explicit operator inputs, secrets remain environment-only, and unavailable combinations fail before OCR execution. - Exclusions: Hidden budgets, per-file/per-tool routing, multi-agent orchestration, or full-repository scan profiles.
- Validation: Profile matrix, precedence, preflight, rendering, and compatibility tests.
- Release classification expectation:
release-required.
- Status: ready
- Priority: medium
- Roadmap theme: M6 Profiles and quality measurement
- Dependencies: Established result, discussion/fingerprint, coverage, posting, and MCP-use receipts. BL-016 is not required for the audit.
- Activation trigger: Met for an audit with current OCR telemetry and toolkit result-derived receipts.
- Upstream overlap: OCR remains authoritative for deterministic tool rendering, provider/model identity, session correlation, diff-review usage/budgets, and request/tool latency. Full-repository
scansignals do not widen toolkit scope. - Goal: Decide whether privacy-safe toolkit telemetry is needed before implementing metrics or routing.
- Scoped deliverables: Inventory OCR token, cost, budget, latency, request, tool-call, and provider/model identity alongside established review health, failed-file coverage, findings, suppression, omission, posting, MCP-use receipts, and M5 context receipts only if they exist. Document only genuinely missing lifecycle, evidence degradation, repeated-discussion, compatibility, or review-value gaps and their privacy/cardinality limits; conclude no-new-layer or create a separately scoped follow-up.
- Acceptance criteria: The audit maps every signal to its authoritative source, distinguishes derived from missing data, and reaches an explicit no-new-layer or separately scoped conclusion. OCR remains authoritative for token, cost, budget, request, latency, and tool-call telemetry; the audit adds no runtime, exporter, public schema, or second context telemetry implementation.
- Exclusions: User surveillance, developer ranking, automatic routing, or mandatory external telemetry.
- Validation: Representative result/discussion fixtures, privacy review, and source-to-signal matrix.
- Release classification expectation:
no-releasefor the audit.
- Status: conditional
- Priority: low
- Roadmap theme: M6 Profiles and quality measurement
- Dependencies: BL-016, BL-017, and an owner-approved quality/cost policy; M5 is not a dependency.
- Activation trigger: Representative metrics demonstrate a stable deterministic rule that improves an explicit objective without reducing safety.
- Goal: Select one run-level profile conservatively from trusted bounded inputs.
- Scoped deliverables: Document the decision rule, inputs, fallback, observability, and opt-out; implement only after replay evaluation and owner approval.
- Acceptance criteria: Routing is deterministic and explainable, never uses untrusted content as authority, cannot let merge-request-controlled inputs select below a repository-configured minimum profile, never selects
ocr scan, and falls back to the policy minimum (defaultstandard) on uncertainty. - Exclusions: Learned online routing, per-tool models, multiple agents, or silent policy changes.
- Validation: Offline replay, adversarial boundaries, fallback tests, and quality thresholds.
- Release classification expectation:
release-required.
- Status: conditional
- Priority: medium
- Roadmap theme: M7 Later and conditional work
- Dependencies: Established evidence, snapshot/delta, scoped-completeness, static-plugin, and built-in MCP contracts.
- Activation trigger: A real repository need identifies a missing ecosystem/framework and supplies safe synthetic fixtures and deterministic semantics.
- Goal: Extend coverage without accumulating shallow detectors.
- Scoped deliverables: Implement one coherent ecosystem or framework pack per activation, with provenance, bounds, source/target deltas, documentation, and public synthetic examples.
- Acceptance criteria: The use case and completion signal are documented before implementation; false-positive behavior and unsupported versions are explicit through the shared scoped coverage contract.
- Exclusions: Checkbox coverage, network resolution, runtime execution, or unrelated bundles.
- Validation: Pack fixtures plus common evidence/bootstrap/MCP contracts.
- Upstream overlap: OCR language allowlists and review rules are review-engine capabilities; a new reviewable language alone does not activate an evidence pack.
- Release classification expectation:
release-required.
- Status: parked
- Priority: medium
- Roadmap theme: M7 Later and conditional work
- Dependencies: Stable evidence/MCP parser interfaces from M1; M5 targets enter the inventory only after their contracts exist.
- Activation trigger: Not met: named targets, bounded CI resources, corpus ownership, and backend criteria across Python 3.12-3.14 are not agreed.
- Goal: Find crashes and invariant violations at untrusted evidence, MCP, result, GitLab payload, registry-metadata, and future M5 parser/protocol boundaries.
- Scoped deliverables: Candidate targets include current evidence/MCP/result/GitLab/registry parsers and future M5 policy parsers, recognizers, handle codec, broker schema, and hostile adapter responses. Select a bounded backend, synthetic seeds, corpus ownership, minimization, and regression policy before activation.
- Acceptance criteria: Targets are deterministic and bounded, minimized failures become tests, corpora contain no repository/provider secrets, and ownership is explicit.
- Exclusions: Unbounded CI, production data, blanket fuzzing, or a runtime dependency.
- Validation: Reproducible smoke campaign and minimized-corpus replay.
- Release classification expectation:
no-release, except user-visible fixes found by it.
- Status: parked
- Priority: low
- Roadmap theme: M7 Later and conditional work
- Dependencies: Established MCP/evidence schemas. M5 owns only
.opencodereview/review-context-policy.json; it neither activates nor depends on this general framework. - Activation trigger: Environment-only configuration is a demonstrated constraint and one coherent schema can cover affected non-secret settings.
- Goal: Improve maintainability without weakening precedence, validation, or secret handling.
- Scoped deliverables: Decide format/versioning, discovery/trust source, field-level environment precedence, migration, allowed non-secret fields, unknown/deprecated-key behavior, and redacted diagnostics before implementation.
- Acceptance criteria: Environment precedence is explicit; secrets are rejected; source files cannot self-authorize; paths remain rooted; unknown/deprecated keys and rollback are documented/tested.
- Exclusions: Credentials on disk, implicit configuration, overlapping formats, or implementation before design approval.
- Validation: Threat model, schema/precedence, migration, and secret rejection.
- Release classification expectation:
release-required.
- Status: conditional
- Priority: low
- Roadmap theme: M7 Later and conditional work
- Dependencies: Stable provider-neutral core contracts and a funded non-GitLab use case. GitLab-first M5 does not depend on it.
- Activation trigger: A named forge has an owner, synthetic fixtures, and explicit parity requirements for CI orchestration, positioning, deduplication, discussion ownership, and safe publication.
- Goal: Add one coherent host adapter without leaking forge semantics into evidence or core result handling.
- Scoped deliverables: The capability matrix covers authentication, diff positions, drafts, discussion acquisition, provider-declared account classification, thread/reply structure, edit/version identity, anchors, resolved/stale state, pagination/snapshot mutation, ambiguous writes, permissions, and idempotency.
- Acceptance criteria: Core remains provider-neutral, GitLab behavior does not regress, unsupported host capabilities fail or degrade explicitly rather than emulate unsafe parity, and the new host meets the approved lifecycle and security matrix.
- Exclusions: Repository ecosystem/framework detection, partial adapters, legacy namespace shims, or multi-host abstractions without a real second provider.
- Validation: Shared adapter contract suite, provider-specific synthetic integration tests, redaction/write-bound tests, and documentation validation.
- Release classification expectation:
release-required.