Current checkpoint, 2026-09-11
Canonical #73 implementation and admission checkpoint. PR #79 and PR #80 are merged, superseding the obsolete runner-implementation blockers. Source landing and synthetic proof are not venue admission or measured acceptance.
Dedicated venue execution evidence remains incomplete. Suitable real authentication/profile access, Gina MCP authorization, routing/reasoning evidence, concrete campaign bounds, UTC expiry, run approval and separate exact-content publication approval remain open. #73 stays open; #71 stays open and blocked by #73. M1-M16, all four profiles, all 34 goals, controlled Sol/medium and Spot-first sequencing remain settled.
Destination
A decision-complete evaluation specification for the dedicated Spot, Perps and Predictions MCPs. It must define two separate layers: deterministic MCP contract/safety checks and natural-language agent task evaluations, with domain coverage, grading evidence, safe execution boundaries, admitted runner profiles and acceptance criteria clear enough for implementation.
Notes
Confirmed scope
The owner selected all three dedicated MCPs, including research, account, trading and automation capabilities, and selected layered coverage plus agent quality. This is a planning map, not permission to implement, access accounts, execute transactions, run paid models or publish measured results. Covering a write intent does not authorize performing that write.
Use the public documentation index, Spot MCP features, Perps MCP features, Predictions MCP features and public askgina/plugins contracts as published references. Public notes must omit private source identities, paths, implementation details, credentials, account identifiers and raw captures. Document missing public evidence rather than presenting internal facts as publicly established.
Reuse, do not reopen
Planning terms
A venue MCP is one dedicated Spot, Perps or Predictions connection, not the combined Gina Read connection. A contract/safety check evaluates a declared tool/workflow invariant against controlled evidence. An agent task trial evaluates a natural-language user goal under an identified connection, account state and runner profile. Those are separate evidence layers, not interchangeable scores.
The evaluation vocabulary records these terms on the isolated research branch. It is a glossary, not a resolved grading or execution policy.
Decisions so far
-
Venue runner policy agreement: owner-confirmed profile/tool conditions and first-live task-rejection policy; numeric action caps return with concrete case proposals. Runner admission remains open for exact runtime/routing evidence and owner approval; measured acceptance remains blocked.
-
Cross-provider result identity and evidence policy: owner-confirmed model, reasoning, budget, package, routing and request-setting rules; native executable pins and exact upstream endpoint slugs return during runner admission for owner approval before credential-bearing execution. Policy closure is not runner admission or measured proof.
-
Spot workflow research: 23 native tools plus MCP bash, explicit intent/settlement distinctions, and candidate coverage with unresolved schemas and safety oracles.
-
Perps workflow research: venue-scoped bash workflows require identity and outcome evidence; public feature examples do not establish complete trading, funding or SQL receipts.
-
Predictions workflow research: event/outcome/time identity, execution state and analytics/scheduling evidence must be graded separately; public gaps remain explicit.
-
Reusable eval evidence research: current Gina Read routing/argument/completion contracts are reusable, but nested venue actions, result grounding, authorization and state outcomes need new admitted evidence.
-
Safe venue execution and account-state boundaries: owner-confirmed bounded, attended live trading in dedicated accounts; concrete campaign approval and separate runner-admission and measured-acceptance decisions remain required.
-
Separate contract and agent-task grading oracles: owner-confirmed outcome and evidence rules, durable started-trial accounting and Fail / Inconclusive / Unavailable / Pass precedence; contract, task and activation claims stay separate.
Venue research, safety and grading decisions are resolved. The shared authentication and identity policies are also resolved; runner admission and measured acceptance remain open.
Not yet specified
- Additional adversarial, recovery and long-running workflow cases that emerge from the first domain/evidence inventories.
- Statistical expansion, refresh cadence and historical comparability after representative cases and grading oracles are chosen.
- The final implementation sequence once the decision frontier is resolved.
Open questions already expressed as child tickets live in those tickets, not a second checklist here.
Out of scope
- Implementing new runners, changing MCP services, running model/MCP evaluations, provisioning credentials or touching live funds while charting.
- Reopening completed native-runner diagnostics, forking OMP to change retries, or silently replacing accepted comparison/auth decisions.
- Trading profitability, investment advice quality or financial-performance claims as a substitute for task correctness and safety.
- Automatic approval of public results, production rollout, consumer CI credential distribution or repository release work.
Current checkpoint, 2026-09-11
Canonical #73 implementation and admission checkpoint. PR #79 and PR #80 are merged, superseding the obsolete runner-implementation blockers. Source landing and synthetic proof are not venue admission or measured acceptance.
Dedicated venue execution evidence remains incomplete. Suitable real authentication/profile access, Gina MCP authorization, routing/reasoning evidence, concrete campaign bounds, UTC expiry, run approval and separate exact-content publication approval remain open. #73 stays open; #71 stays open and blocked by #73. M1-M16, all four profiles, all 34 goals, controlled Sol/medium and Spot-first sequencing remain settled.
Destination
A decision-complete evaluation specification for the dedicated Spot, Perps and Predictions MCPs. It must define two separate layers: deterministic MCP contract/safety checks and natural-language agent task evaluations, with domain coverage, grading evidence, safe execution boundaries, admitted runner profiles and acceptance criteria clear enough for implementation.
Notes
Confirmed scope
The owner selected all three dedicated MCPs, including research, account, trading and automation capabilities, and selected layered coverage plus agent quality. This is a planning map, not permission to implement, access accounts, execute transactions, run paid models or publish measured results. Covering a write intent does not authorize performing that write.
Use the public documentation index, Spot MCP features, Perps MCP features, Predictions MCP features and public
askgina/pluginscontracts as published references. Public notes must omit private source identities, paths, implementation details, credentials, account identifiers and raw captures. Document missing public evidence rather than presenting internal facts as publicly established.Reuse, do not reopen
Planning terms
A venue MCP is one dedicated Spot, Perps or Predictions connection, not the combined Gina Read connection. A contract/safety check evaluates a declared tool/workflow invariant against controlled evidence. An agent task trial evaluates a natural-language user goal under an identified connection, account state and runner profile. Those are separate evidence layers, not interchangeable scores.
The evaluation vocabulary records these terms on the isolated research branch. It is a glossary, not a resolved grading or execution policy.
Decisions so far
Venue runner policy agreement: owner-confirmed profile/tool conditions and first-live task-rejection policy; numeric action caps return with concrete case proposals. Runner admission remains open for exact runtime/routing evidence and owner approval; measured acceptance remains blocked.
Cross-provider result identity and evidence policy: owner-confirmed model, reasoning, budget, package, routing and request-setting rules; native executable pins and exact upstream endpoint slugs return during runner admission for owner approval before credential-bearing execution. Policy closure is not runner admission or measured proof.
Spot workflow research: 23 native tools plus MCP bash, explicit intent/settlement distinctions, and candidate coverage with unresolved schemas and safety oracles.
Perps workflow research: venue-scoped bash workflows require identity and outcome evidence; public feature examples do not establish complete trading, funding or SQL receipts.
Predictions workflow research: event/outcome/time identity, execution state and analytics/scheduling evidence must be graded separately; public gaps remain explicit.
Reusable eval evidence research: current Gina Read routing/argument/completion contracts are reusable, but nested venue actions, result grounding, authorization and state outcomes need new admitted evidence.
Safe venue execution and account-state boundaries: owner-confirmed bounded, attended live trading in dedicated accounts; concrete campaign approval and separate runner-admission and measured-acceptance decisions remain required.
Separate contract and agent-task grading oracles: owner-confirmed outcome and evidence rules, durable started-trial accounting and Fail / Inconclusive / Unavailable / Pass precedence; contract, task and activation claims stay separate.
Venue research, safety and grading decisions are resolved. The shared authentication and identity policies are also resolved; runner admission and measured acceptance remain open.
Not yet specified
Open questions already expressed as child tickets live in those tickets, not a second checklist here.
Out of scope