Skip to content

Experiment: SWE Prime quality gated trajectory and segment selection #51

Description

@ruvnet

Finding

SWE Prime, arXiv 2608.27449, submitted 2026 08 27, reports that successful software engineering trajectories are not uniformly useful supervision. Its two stage selection ranks trajectories by process quality, result quality, and representativeness, then selects semantic segments by contribution, learnability, and risk. The originating team reports that a selected 10 percent subset beats training on the full resolved set by up to 12.2 percent relative on SWE Bench Pro and 24.2 percent relative on SWE Bench Verified.

Evidence class: originating team report. No Dream Machine reproduction yet.

RuV experiment

Reuse existing witness, anchored replay, security, and evaluator receipts rather than building another training subsystem. Keep full traces for provenance, but only selected traces and semantic segments may contribute to skill distillation, memory mutation, or supervised learning.

Compare all successful trajectories, top 10 percent by result only, and multi factor selection plus segment masking under identical model, tasks, budget, seeds, and evaluator.

Acceptance

Graduate only if multi factor selection beats the stronger baseline by at least 5 absolute percentage points on held out work, or preserves task success while reducing training tokens or GPU time by at least 50 percent. No selected segment may contain a known policy violation, evaluator mutation, or authority expansion. If result only selection matches the multi factor condition, reject the extra complexity.

A reproduction protocol is committed on an isolated branch.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions