cf-o200k-pretrainwas removed along with theo200ktraining pipeline (seetraining.md). The tracked YAML configurations below still describe the intended run shapes but currently need a driver to execute against.
The repository separates direct runner configurations, bounded experiments, and cluster plans.
cf-o200k-pretrain accepts YAML defaults:
cf-o200k-pretrain \
--config configs/run_configs/100m_o200k_tr_rocm_mi350x.yamlThe runner reads the top-level run mapping. CLI flags override YAML values:
cf-o200k-pretrain \
--config configs/run_configs/100m_o200k_tr_rocm_mi350x.yaml \
--steps 10 \
--run-name smoke-overrideUnknown keys fail immediately rather than being silently ignored.
run:
attention_type: gqa
mlp_type: token_routed
num_attention_heads: 8
num_key_value_heads: 2
shared_expert: true
routing_strategy: modulo_cyclic
top_k: 2run:
attention_type: mha
mlp_type: token_routed
num_attention_heads: 8
num_key_value_heads: 8
shared_expert: true
routing_strategy: modulo_cyclic
top_k: 2run:
attention_type: gqa # or mha
mlp_type: swiglu
shared_expert: falseMatch parameter counts explicitly by adjusting intermediate_size and
shared_intermediate_size.
| Value | Route source | Frequency artifact |
|---|---|---|
modulo_cyclic |
token ID | not required |
modulo |
token ID | legacy alias for modulo_cyclic |
round_robin |
token ID rank | not required |
random |
fixed seeded lexical partition | not required |
lsh_hidden |
hidden-state hash | not lexical routing |
zipf and modulo_balanced_secondary remain parser-compatible only for the
historical TMLR ablations. They are not used by canonical run configurations.
Older documentation mentioning zipf_token_class is obsolete.
The local runner defines 50m, 100m, 200m_o200k, 300m, 1b, and 8b
profiles.
Profile names are approximate. The realized parameter count depends on
vocabulary size and explicit overrides and is recorded at launch.
The matched 200,082,688-parameter Dense/TR protocol, frozen 4B-token dataset preparation, B200 launcher, artifact export, and server teardown checklist are documented in the 200M B200 runbook.
configs/run_configs/experiments_100m contains bounded MPS and architecture
experiments. Several files contain absolute local dataset paths. Treat them as
reproducibility records and update paths before launch.
configs/run_configs/review_h200 contains matched review and evidence runs.
configs/run_configs/ablations_100m contains routing controls.
Cluster plans with top-level model, parallel, and run sections are
validated with:
cf-plan-cluster \
--config configs/run_configs/8b_o200k_tr_32t_gb300_4608.yamlThe planner checks TP × PP × DP world size, global batch, steps, target tokens, and overshoot. It does not launch Slurm, Kubernetes, Ray, or vendor jobs.
Do not pass a cluster-plan YAML directly to cf-o200k-pretrain.
At first launch, the runner writes:
runs/<run-name>/run_config.json
On resume it rejects differences in training-critical arguments,
ModelConfig, and frozen token-shard identities. Operational fields such as
logging cadence, evaluation cadence, save cadence, and output names may change.
Use --force-resume only when the mismatch is understood and documented.