benchmark · benchmark--vlabench

VLABench

Language-conditioned simulation benchmark for VLA manipulation.

Identity and state

PublicationReleased
DeploymentResearch
Canonical claims16
Last reviewed2026-09-04

Reviewed relationships

StateRelated entityRelationScopeBoundary
VerifiedMuJoCouses simulatorbenchmark executionThe benchmark states MuJoCo plus dm_control.

Canonical evidence claims

Claims are projected from authored tables. A dash remains an explicit missing value; it is never filled from a neighboring product or name match.

MetricValueClass
Asset library2,164 objects across 163 categories [benchmarks-r32]vla-manipulation-eval
Claim boundary100 task categories are not RoboTwin’s bimanual tasks; performance tables stay in the paper [benchmarks-r32]vla-manipulation-eval
Embodiment in the sourcestandard evaluation: 7-DoF Franka Emika Panda with parallel gripper; framework also supports other forms [benchmarks-r32]vla-manipulation-eval
Engine / rendererMuJoCo + dm_control [benchmarks-r32]vla-manipulation-eval
Evaluation scope6 capability dimensions: mesh / texture, spatial, common sense / knowledge, semantic instruction, physical laws, and long-horizon reasoning [benchmarks-r32]vla-manipulation-eval
First public datearXiv 2024-12-24 [benchmarks-r32]vla-manipulation-eval
Kindlanguage-conditioned manipulation benchmark for VLAs, foundation-model workflows, and VLMs [benchmarks-r32]vla-manipulation-eval
Language-conditionedyes — natural instructions include implicit intent, common sense, and multi-step reasoning [benchmarks-r32]vla-manipulation-eval
Mediumsimulation [benchmarks-r32]vla-manipulation-eval
PaperarXiv:2412.18194; ICCV 2025 [benchmarks-r32] [benchmarks-r33]vla-manipulation-eval
Primary metric familyinteractive task success + graduated Progress Score; non-interactive VLM action-sequence matching [benchmarks-r32]vla-manipulation-eval
Public codeyes — MIT-licensed official repository [benchmarks-r33]vla-manipulation-eval
Public datayes — evaluation episodes and fine-tuning datasets are linked from the official project [benchmarks-r33]vla-manipulation-eval
Public leaderboardincomplete — the official repository still lists standard VLA / VLM leaderboard integration as open [benchmarks-r33]vla-manipulation-eval
Released trajectoriesstandard evaluation episodes and primitive-task fine-tuning data are public; no one suite-wide trajectory count is stated [benchmarks-r33]vla-manipulation-eval
Task structure100 task categories: 60 primitive + 40 composite [benchmarks-r32]vla-manipulation-eval

Source ledger

What is not inferred

No fuzzy entity matching, adjacent-column borrowing, universal compatibility, commercial maturity, or checkpoint-to-product identity. Absence of a relation means unreviewed or unsupported here—not proven incompatibility.