Every meaningful run should leave behind structured evidence.
- Validate repo
- Check backend health
- Run model benchmarks
- Score outputs
- Attempt promotion
- Run quantization-retention report
- Record experiment metadata
- model_name
- adapter
- backend endpoint
- prompt_profile
- benchmark input file
- scored output file
- promotion decision
- quantization report path
- notes
No checkpoint should be described as improved unless an experiment record exists.