Systematic verification that vla-eval reproduces published VLA model scores across codebases and benchmarks.
| Codebase | Model | LIBERO | CALVIN | SE WidowX VM | SE GR VM | SE GR VA |
|---|---|---|---|---|---|---|
| openvla/openvla | OpenVLA | ✅ 76.2% / 76.5% |
· | · | · | · |
| Physical-Intelligence/openpi | π₀.₅ | ✅ 97.7% / 96.9% |
· | · | · | · |
| microsoft/CogACT | CogACT | · | · | ⬜ 51.3% |
⬜ 74.8% |
⬜ 61.3% |
| moojink/openvla-oft | OpenVLA-OFT | ✅ 96.7% / 97.1% |
· | · | · | · |
| NVIDIA/Isaac-GR00T | GR00T N1 | 🟡 94.9% / 97.0%¹ |
· | 🔧 29.6% / 57.1% |
🟡 59.7% / 67.7% |
· |
| baaivision/UniVLA | UniVLA | ⬜ 95.5% |
⬜ 4.63 |
⬜ 69.8% |
· | · |
| 2toinf/X-VLA | X-VLA | ✅ 97.4% / 98.1% |
✅ 4.30 / 4.43 |
✅ 94.8% / 95.8% |
✅ 100% / 98.3% |
🟡 80.8% / 84.0%⁴ |
| Dexmal/dexbotic | DB-CogACT | ✅ 94.7% / 94.9% |
✅ 4.02 / 4.06 |
🟡 63.5% / 69.5% |
· | · |
| DB-π₀ | ⬜ 93.9% |
· | · | · | · | |
| DB-OFT | · | ⬜ 3.54 |
⬜ 76.4% |
· | · | |
| DB-MemVLA | ⬜ 97.0% |
· | ⬜ 84.4% |
· | · | |
| DB-GR00TN1 | ⬜ 94.8%† |
· | · | · | · | |
| starVLA/starVLA | Q2.5-FAST | ⬜ 95.2% |
· | ✅ 64.6% / 58.6% |
· | · |
| Qwen3-FAST | ⬜ 95.4%† |
· | ⬜ 31.6%† |
· | · | |
| Qwen3-OFT | ✅ 96.8% / 97.8%² |
· | ⬜ 42.7% |
· | · | |
| Qwen3-PI | ⬜ 95.7% |
· | ⬜ 60.9%† |
· | · | |
| Qwen3-GR00T | ⬜ 96.5%† |
⬜ 3.76† |
✅ 66.7% / 65.3% |
· | · | |
| DravenALG/VLANeXt | VLANeXt | ✅ 96.9% / 97.4%³ |
· | · | · | · |
| huggingface/lerobot | π₀.₅ (LeRobot) | ✅ 100% / 99.0%⁵ |
· | · | · | · |
| GR00T N1.7 | ✅ 99.0% / 81%⁵ |
· | · | · | · | |
| MolmoAct2 | ✅ 97.0% / 98.0%⁵ |
· | · | · | · | |
| VLA-JEPA | ✅ 96.0% / 93.0%⁵ |
· | · | · | · |
SE = SimplerEnv. SE GR = Google Robot VM.
¹ Community checkpoint (not official NVIDIA). ² Spatial suite only (reported 97.8%); 4-suite avg is 96.6%. ³ 4-suite LIBERO average (Spatial 98.2, Object 98.8, Goal 97.0, Long 93.6). ⁴ Self-reported best rollout. Move Near 10 variants × 60 eps. ⁵ Single suite, 10 tasks × 10 eps (π₀.₅/GR00T: Object, MolmoAct2: Goal, VLA-JEPA: Long); see lerobot.md. † Checkpoint not publicly available on HuggingFace.
| Benchmark | Model | Reproduced | Reported | Details |
|---|---|---|---|---|
| RoboTwin | X-VLA | ⬜ | 70/39% | xvla.md |
| RoboTwin | DB-CogACT | ⬜ | 58.5% | dexbotic.md |
| RoboTwin | starVLA Qwen3-OFT | ⬜ | 50.4% | starvla.md |
| RoboMME | MME-VLA π₀.5 | ✅ 25.5% | 22.7%¹ | robomme.md |
| RoboMME | MME-VLA FrameSamp | ⬜ | 44.5% | robomme.md |
| ManiSkill2 | DB-CogACT | ⬜ | 58% | dexbotic.md |
| ManiSkill2 | DB-π₀ | ⬜ | 65% | dexbotic.md |
| ManiSkill2 | DB-OFT | ⬜ | 63% | dexbotic.md |
| Kinetix | RTC | ⬜ | ckpt | rtc.md |
| VLABench | X-VLA | ⬜ | 51.1% | xvla.md |
| MolmoSpaces-Bench | MolmoBot (F=2) | ✅ 57.0% | 57.7%² | molmobot.md |
| RoboDojo | π₀.₅ | 🔧 **6.4 / 5.6%**³ | 5.78 / 4.56% | robodojo.md |
¹ Counting suite only (4/16 tasks). Full 4-suite evaluation pending.
² Pick-and-Place only on procthor-objaverse/FrankaPickandPlaceHardBench (200 ep). Other task types not yet reproduced.
³ Score / SR, Memory dimension only (1 of 5), 15/50 episodes, seed 0 of 3; 2 of 6 tasks blocked by an upstream layout bug (counted as 0).
Cell format: status / [reported%](HF checkpoint link). Bold = our reproduced score. Per-codebase reproduction details live in the per-codebase docs linked below; rows without a per-codebase doc (currently VLANeXt) hyperlink the bolded score to the landing PR instead.
Status: ✅ within 95% CI · 🟡 outside CI but ≤5pp · 🔧 in progress or >5pp with known cause · ⬜ not attempted · · no score / no checkpoint. CI is binomial at p=0.95 (±1.9pp for 500 episodes).
Integrated in vla-eval: RLBench, RoboCasa, Mikasa, RoboCerebra, LIBERO-90, LIBERO-Pro, BEHAVIOR-1K (details; needs an R1Pro-compatible model server).
- openvla/openvla: openvla.md
- Physical-Intelligence/openpi: openpi.md
- microsoft/CogACT: cogact.md
- moojink/openvla-oft: oft.md
- NVIDIA/Isaac-GR00T: groot.md
- Physical-Intelligence/rtc: rtc.md
- 2toinf/X-VLA: xvla.md
- Dexmal/dexbotic: dexbotic.md
- starVLA/starVLA: starvla.md
- RoboMME/robomme_policy_learning: robomme.md
- DravenALG/VLANeXt: PR #34
- allenai/MolmoBot: molmobot.md
- RoboDojo-Benchmark/RoboDojo (π₀.₅ via XPolicyLab): robodojo.md
| File | Contents |
|---|---|
| common-pitfalls.md | Reproduction pitfalls taxonomy |
| running-guide.md | How to run evaluations + supply/demand data |
data/ |
Raw result JSONs per codebase×benchmark |