This repository assumes a three-model ladder:
- Flagship: highest quality candidate used for best cognition performance.
- Fallback: lower-cost model intended to preserve doctrine and bounded behavior.
- Eval: fast experimental model used for rapid benchmark iteration.
No model becomes canonical unless it improves reasoning and repair without unacceptable regression in calibration, compression retention, or prompt-independence.