Skip to content

Latest commit

 

History

History

README.md

Reproductions

Systematic verification that vla-eval reproduces published VLA model scores across codebases and benchmarks.

Core Benchmarks

Codebase Model LIBERO CALVIN SE WidowX VM SE GR VM SE GR VA
openvla/openvla OpenVLA
76.2% / 76.5%
· · · ·
Physical-Intelligence/openpi π₀.₅
97.7% / 96.9%
· · · ·
microsoft/CogACT CogACT · ·
51.3%

74.8%

61.3%
moojink/openvla-oft OpenVLA-OFT
96.7% / 97.1%
· · · ·
NVIDIA/Isaac-GR00T GR00T N1 🟡
94.9% / 97.0%¹
· 🔧
29.6% / 57.1%
🟡
59.7% / 67.7%
·
baaivision/UniVLA UniVLA
95.5%

4.63

69.8%
· ·
2toinf/X-VLA X-VLA
97.4% / 98.1%

4.30 / 4.43

94.8% / 95.8%

100% / 98.3%
🟡
80.8% / 84.0%⁴
Dexmal/dexbotic DB-CogACT
94.7% / 94.9%

4.02 / 4.06
🟡
63.5% / 69.5%
· ·
DB-π₀
93.9%
· · · ·
DB-OFT ·
3.54

76.4%
· ·
DB-MemVLA
97.0%
·
84.4%
· ·
DB-GR00TN1
94.8%†
· · · ·
starVLA/starVLA Q2.5-FAST
95.2%
·
64.6% / 58.6%
· ·
Qwen3-FAST
95.4%†
·
31.6%†
· ·
Qwen3-OFT
96.8% / 97.8%²
·
42.7%
· ·
Qwen3-PI
95.7%
·
60.9%†
· ·
Qwen3-GR00T
96.5%†

3.76†

66.7% / 65.3%
· ·
DravenALG/VLANeXt VLANeXt
96.9% / 97.4%³
· · · ·
huggingface/lerobot π₀.₅ (LeRobot)
100% / 99.0%
· · · ·
GR00T N1.7
99.0% / 81%
· · · ·
MolmoAct2
97.0% / 98.0%
· · · ·
VLA-JEPA
96.0% / 93.0%
· · · ·

SE = SimplerEnv. SE GR = Google Robot VM.

¹ Community checkpoint (not official NVIDIA). ² Spatial suite only (reported 97.8%); 4-suite avg is 96.6%. ³ 4-suite LIBERO average (Spatial 98.2, Object 98.8, Goal 97.0, Long 93.6). ⁴ Self-reported best rollout. Move Near 10 variants × 60 eps. ⁵ Single suite, 10 tasks × 10 eps (π₀.₅/GR00T: Object, MolmoAct2: Goal, VLA-JEPA: Long); see lerobot.md. † Checkpoint not publicly available on HuggingFace.

Other Benchmarks

Benchmark Model Reproduced Reported Details
RoboTwin X-VLA 70/39% xvla.md
RoboTwin DB-CogACT 58.5% dexbotic.md
RoboTwin starVLA Qwen3-OFT 50.4% starvla.md
RoboMME MME-VLA π₀.5 25.5% 22.7%¹ robomme.md
RoboMME MME-VLA FrameSamp 44.5% robomme.md
ManiSkill2 DB-CogACT 58% dexbotic.md
ManiSkill2 DB-π₀ 65% dexbotic.md
ManiSkill2 DB-OFT 63% dexbotic.md
Kinetix RTC ckpt rtc.md
VLABench X-VLA 51.1% xvla.md
MolmoSpaces-Bench MolmoBot (F=2) 57.0% 57.7%² molmobot.md
RoboDojo π₀.₅ 🔧 **6.4 / 5.6%**³ 5.78 / 4.56% robodojo.md

¹ Counting suite only (4/16 tasks). Full 4-suite evaluation pending. ² Pick-and-Place only on procthor-objaverse/FrankaPickandPlaceHardBench (200 ep). Other task types not yet reproduced. ³ Score / SR, Memory dimension only (1 of 5), 15/50 episodes, seed 0 of 3; 2 of 6 tasks blocked by an upstream layout bug (counted as 0).

Cell format: status / [reported%](HF checkpoint link). Bold = our reproduced score. Per-codebase reproduction details live in the per-codebase docs linked below; rows without a per-codebase doc (currently VLANeXt) hyperlink the bolded score to the landing PR instead.

Status: ✅ within 95% CI · 🟡 outside CI but ≤5pp · 🔧 in progress or >5pp with known cause · ⬜ not attempted · · no score / no checkpoint. CI is binomial at p=0.95 (±1.9pp for 500 episodes).

Benchmarks with No Model Coverage Yet

Integrated in vla-eval: RLBench, RoboCasa, Mikasa, RoboCerebra, LIBERO-90, LIBERO-Pro, BEHAVIOR-1K (details; needs an R1Pro-compatible model server).

Per-Codebase Details

Files

File Contents
common-pitfalls.md Reproduction pitfalls taxonomy
running-guide.md How to run evaluations + supply/demand data
data/ Raw result JSONs per codebase×benchmark