Finding
SWE Prime, arXiv 2608.27449, submitted 2026 08 27, reports that successful software engineering trajectories are not uniformly useful supervision. Its two stage selection ranks trajectories by process quality, result quality, and representativeness, then selects semantic segments by contribution, learnability, and risk. The originating team reports that a selected 10 percent subset beats training on the full resolved set by up to 12.2 percent relative on SWE Bench Pro and 24.2 percent relative on SWE Bench Verified.
Evidence class: originating team report. No Dream Machine reproduction yet.
RuV experiment
Reuse existing witness, anchored replay, security, and evaluator receipts rather than building another training subsystem. Keep full traces for provenance, but only selected traces and semantic segments may contribute to skill distillation, memory mutation, or supervised learning.
Compare all successful trajectories, top 10 percent by result only, and multi factor selection plus segment masking under identical model, tasks, budget, seeds, and evaluator.
Acceptance
Graduate only if multi factor selection beats the stronger baseline by at least 5 absolute percentage points on held out work, or preserves task success while reducing training tokens or GPU time by at least 50 percent. No selected segment may contain a known policy violation, evaluator mutation, or authority expansion. If result only selection matches the multi factor condition, reject the extra complexity.
A reproduction protocol is committed on an isolated branch.
Finding
SWE Prime, arXiv 2608.27449, submitted 2026 08 27, reports that successful software engineering trajectories are not uniformly useful supervision. Its two stage selection ranks trajectories by process quality, result quality, and representativeness, then selects semantic segments by contribution, learnability, and risk. The originating team reports that a selected 10 percent subset beats training on the full resolved set by up to 12.2 percent relative on SWE Bench Pro and 24.2 percent relative on SWE Bench Verified.
Evidence class: originating team report. No Dream Machine reproduction yet.
RuV experiment
Reuse existing witness, anchored replay, security, and evaluator receipts rather than building another training subsystem. Keep full traces for provenance, but only selected traces and semantic segments may contribute to skill distillation, memory mutation, or supervised learning.
Compare all successful trajectories, top 10 percent by result only, and multi factor selection plus segment masking under identical model, tasks, budget, seeds, and evaluator.
Acceptance
Graduate only if multi factor selection beats the stronger baseline by at least 5 absolute percentage points on held out work, or preserves task success while reducing training tokens or GPU time by at least 50 percent. No selected segment may contain a known policy violation, evaluator mutation, or authority expansion. If result only selection matches the multi factor condition, reject the extra complexity.
A reproduction protocol is committed on an isolated branch.