This project is an experimental market-making research engine, not a production trading system. The current goal is to test a reproducible loop from public top-of-book data to Rust paper execution and small learned policy controllers.
Can Python-trained policy controllers choose quoting behavior in a way that improves risk-adjusted paper-session utility when Rust executes the exported models?
The learned policies are considered interesting only if they improve over simple baselines after being exported to JSON and loaded back into the Rust paper engine.
static: uses the controller quote unchanged.adaptive: widens and skews quotes from observed spread, volatility, and inventory.hybrid: switches from static to adaptive on risk triggers.selector: weighted rule-based selector between static and adaptive behavior.learned_selector: Rust-executed logistic-regression gate trained by Python.linear_agent: Rust-executed multi-action ridge-regression utility model trained by Python.bandit_agent: Rust-executed LinUCB contextual-bandit policy trained by Python.
Run the gate once to produce window-level policy outcomes:
python3 research/policy_evaluation_gate.pyTrain the learned controllers from those outcomes:
python3 research/train_policy_selector.py
python3 research/train_linear_policy_agent.py
python3 research/train_contextual_bandit_agent.pyRerun the gate so Rust loads and evaluates the learned model:
python3 research/policy_evaluation_gate.py
python3 research/write_project_report.pyThe learned model artifacts are written to:
target/research/learned_policy_selector_model.json
target/research/linear_policy_agent_model.json
target/research/contextual_bandit_agent_model.json
The main utility is:
pnl - 2.0 * drawdown - 0.02 * mean_abs_inventory
Reports also track fills, fees, adaptive-step percentage, policy triggers, dataset wins, and window wins.
The policy gate evaluates each policy under:
configured: the fill model in the committed run configs.conservative_fill: a stricter touch-intensity fill assumption.liquid_fill: a more permissive touch-intensity fill assumption.
A result is not robust if it only works under one fill assumption.
A useful learned-policy result should satisfy most of these:
- Rust-executed
learned_selectorbeatsadaptiveunder configured assumptions. - Rust-executed
learned_selectoris competitive with or better than the hand-tunedselector. - The result does not collapse under conservative fill assumptions.
- Python leave-one-dataset-out validation does not contradict the Rust gate result.
- Trigger attribution shows the learned policy is not simply always-adaptive or never-adaptive.
- Linear and bandit agents are treated as proof-of-concept controllers unless they beat simpler policies in the gate.
These criteria are research checks, not proof of a trading edge.
The latest result is promising because the Rust-executed learned selector leads under configured assumptions after adding fresh quote datasets. The linear agent wins the liquid-fill sensitivity, and the contextual bandit is executable but not the best current policy. It is still a small-sample result, and fill realism remains the weakest point.
The project has also completed a live public-data paper demonstration with the learned selector. That demo is an operational check of the full loop, not evidence of a trading edge. See final_report.md for the current wrap-up result.