This project consists of two main components:
- Market Data Collector (Go) — Streams and stores market data frames
- Machine Learning Pipeline (Python) — Trains models on collected data
- Go (latest stable recommended)
- Python 3.13+
The collector gathers market data and stores it as sequential frames.
go run main.goData is written to:
data/frames.jsonl- Let the collector run for a sufficient amount of time to accumulate long sequences per ticker
- Longer sequences generally improve model training quality
Once you have collected enough data, you can train the model.
python ML/train.py --frames ./data/frames.jsonl --epochs 30 --arch {fast, full, auto}-
--frames
Path to the collected data file (.jsonl) -
--epochs
Number of training epochs (e.g.,30) -
--arch
Model architecture:fast— Lightweight, quicker trainingfull— More complex, higher performanceauto— Automatically selects architecture
-
Start the collector:
go run main.go
-
Let it run until sufficient data is collected
-
Train the model:
python ML/train.py --frames ./data/frames.jsonl --epochs 30 --arch {auto, fast, full}
- Ensure
frames.jsonlis non-trivial in size before training - Start with
--arch fastfor quick iteration - Use
--arch fullonce data volume is sufficient