Skip to content

release: publish a current-source 14-arm S30/H600 replication dataset #8018

Description

@ll7

Archetype Metadata

archetype: benchmark-campaign
evidence_tier: blocked
linked_policy:
  - docs/benchmark_release_protocol.md
  - docs/benchmark_release_reproducibility.md
  - docs/maintainer_values.md
  - docs/context/artifact_evidence_vocabulary.md

Goal / Problem

Publish the next full Robot SF benchmark-data release as a current-source replication and
terminal-outcome data-quality release
. Use the unchanged comparable 14-arm × 48-scenario ×
30-seed, H600, differential-drive contract (20,160 exact episodes), while enforcing a strict,
mutually exclusive partition among route-complete, collision, and timeout outcomes.

The previous release is complete and immutable under #7742. This successor must use a new issue,
campaign, source freeze, source-derived tag, GitHub Release, and fresh benchmark-only Zenodo
concept/version. It must not mutate or reuse the August tag or DOI.

Hypothesis / Claim Boundary

The next frozen source reproduces the established raw/component outcome surface while eliminating
the terminal-state contradiction corrected after the August release. The comparison is paired by
planner/scenario/seed against DOI 10.5281/zenodo.22077448.

This is not a new planner-superiority study. Do not add RecurrentPPO, force-coupled potential field,
or any other smoke-only planner. Social Navigation Quality Index (SNQI) remains advisory and has
no ranking authority when calibration is invalid.

Frozen Scientific Contract

  • Planner arms: the existing canonical 14-arm S30/H600 roster.
  • Scenarios: configs/scenarios/classic_interactions_francis2023.yaml (48).
  • Seeds: paper_eval_s30, 111–140 (30).
  • Horizon: 600 steps; dt=0.1; differential-drive kinematics.
  • Expected identities: 20,160, each exactly once under one source SHA.
  • Authoritative interpretation: raw and component metrics plus paired outcome deltas.
  • Forbidden evidence: fallback, degraded, unavailable, failed, partial, duplicate, unexpected,
    provenance-invalid, or terminal-contradictory rows.

The actual release SHA is frozen only after the public/private rehearsal tooling and any software
release-contract changes merge. Prefer the same exact source candidate as v0.0.6 when those
changes are runtime-neutral and all benchmark gates pass; otherwise record distinct source
identities explicitly.

Blocking Dependencies

  1. Public PR feat(release): add hash-bound no-campaign rehearsal mode #7967: refresh to current main, obtain canonical independent domain approval, merge,
    and run a real no-campaign rehearsal.
  2. Private issue ll7/robot_sf_ll7-private-ops#225 / PR feat(139): extract visualization & formatting helpers #226: consume public rehearsal, close
    receipt/side-effect/export-control findings, rebase, fully test, independently review, and merge.
  3. Implement and prove a satisfiable future source_sha freeze/resolved-manifest construction;
    do not use a self-referential tracked manifest.
  4. Create new v0.2 release/config/metadata identities and a final-source-derived tag. Reserve a
    fresh benchmark-only Zenodo concept/version before freeze without exposing credentials.
  5. Freeze and stage all five checkpoint records with submit_safe=true.

Execution

  1. Freeze exact source, release manifest, matrix/config/scenario/SNQI assets, planner configs,
    checkpoint digests, environment, route, private packet, and queue identity.
  2. Require exact-source CI and CodeQL plus the no-campaign public/private rehearsal.
  3. Run at the same SHA:
    • 14-arm one-scenario/one-seed runtime smoke;
    • 70-cell H600 hybrid stress smoke, including both francis2023_leave_group -> orca branch
      witnesses;
    • guarded-PPO × francis2023_parallel_traffic × seed 132 horizon-boundary regression.
  4. After zero forbidden markers and a passing release doctor, submit one fresh one-attempt Slurm
    campaign through private ops.
  5. Accept only the exact full matrix; preserve source and derived trees in two independently
    read-back failure domains; build and checksum the publication archive.
  6. Create GitHub and Zenodo drafts with byte-identical archive content, cold-download both into
    empty directories, verify all member checksums/metadata/source/tag/DOI/cardinality, then perform
    the protected publish operation and rerun the credential-free public audit.

Runtime / Storage Estimate

  • Prior observed full campaign: 2:43:05 at 2.06 episodes/s on 36 CPUs, 256 GiB, one allocated L40S
    with CPU-only benchmark workers.
  • Request: proven one-node route, 36 CPUs, 256 GiB, one L40S, 8-hour wall clock.
  • Expected execution: 3–4 hours plus scheduler wait; stop/review after the 8-hour bound.
  • Working storage: reserve 3 GiB; stop and re-plan above 5 GiB.
  • Publication archive baseline: 54.2 MB compressed / 714.9 MB expanded; retain two independent
    expanded copies plus receipts and cold extraction.

Definition of Done

  • Public and private no-allocation rehearsal paths are merged, reviewed, and pass at the exact
    release source.
  • Future source identity construction is executable and non-self-referential.
  • New manifest/config/metadata/tag/DOI/campaign identities are frozen and collision-free.
  • Five of five checkpoints are staged, checksummed, and submit_safe=true.
  • Runtime smoke, hybrid stress, and the terminal-boundary regression pass with zero forbidden
    markers.
  • Full run has 14/14 arms and 20,160/20,160 unique exact identities under one SHA.
  • Route-complete, collision, and timeout outcomes are mutually exclusive for every row.
  • Two-copy preservation and cold readback reproduce every checksum.
  • GitHub and Zenodo archives are byte-identical and independently retrievable.
  • Publication notes preserve the paired-replication boundary and advisory SNQI status.

Stop / Restart Rules

  • Stop before allocation on any source/config/receipt/route/queue/CI/tag/DOI drift.
  • Resume an identity only for a receipt-proven infrastructure interruption with unchanged inputs.
  • Code, configuration, dependency, checkpoint, or semantic defects require a new commit and fresh
    campaign identity. Never mix attempts.
  • Unexpected paired drift triggers domain review, not selective reruns or changed thresholds.
  • Validator-only post-execution failures preserve original job truth and may use a separately
    reviewed derived revalidation; never relabel the scheduler attempt.

Validation / Testing

  • Focused release-protocol, rehearsal, source-freeze, checkpoint, adaptive-branch, terminal-state,
    preservation, publication, and credential-free audit suites.
  • BASE_REF=origin/main PR_READY_MODE=final scripts/dev/pr_ready_check.sh for the final candidate.
  • Exact private queue/packet/route/ledger/preflight tests and sbatch --test-only before the single
    submission.
  • Independent cold-start verification from public tag plus DOI after publication.

Effort / Confidence

  • After prerequisites merge: same-day execution is plausible—about 3–4 hours compute and 2–5
    hours active validation/preservation/publication, plus external queue/service delays.
  • Tooling/PR prerequisites: 1–3 engineering days if no new defects are found.
  • Scientific-scope confidence: 91%; the user explicitly requested another full benchmark, so this
    issue selects the comparable replication option and does not expand the matrix.

Project Metadata

  • Priority: highest benchmark release lane
  • Effort: 12–30 active hours plus 3–4 hours observed compute and external waits
  • Reviewed: 2026-08-30 audit at 28fa8acb5474a728b2f3ce0c75dba92a2e2006a4
schema: goal_autopilot_preparation.v1
repository: ll7/robot_sf_ll7
issue: 8018
source_body_sha256: 5cc15adf7b3a8a404e6110943de19d15c9850fd8bd50d24751188a4cdbf27625
source_comments_sha256: 
audit_schema: open_issue_contract_audit.v1
audit_digest: 1c191411793929640684e0827d763f3a2009a8682ab98dcefa8c2ded3e54f4ea
audit_classification: needs_compute
next_action: route_to_compute_authority
authority: compute_authority
execution_mode: compute
preferred_worker: MaxRunner
expected_pr_runner_label: runner:max
implementation_admitted: False
state_ready_change_proposed: False
mutation_batch: open-issues-20260830-04

This packet is preparation evidence only. It never overrides live labels, exact claim state, branch state, typed dependencies, domain gates, compute authority, release authority, or scientific evidence rules.

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkBenchmark-related workpriority:0P0: correctness, safety, or CI-critical workreleaseRelease relevant issuesresource:slurmSLURM or worker execution requiredslurmRequires or is running on SLURM/Auxme infrastructurestate:runningExternal run is currently activetype:benchmarkBenchmark campaign task

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions