VLABench
Language-conditioned simulation benchmark for VLA manipulation.
Identity and state
- No openness, availability, deployment, or interoperability facet is published for this identity.
Reviewed relationships
| State | Related entity | Relation | Scope | Boundary |
|---|---|---|---|---|
| Verified | → MuJoCo | uses simulator | benchmark execution | The benchmark states MuJoCo plus dm_control. |
Canonical evidence claims
Claims are projected from authored tables. A dash remains an explicit missing value; it is never filled from a neighboring product or name match.
| Metric | Value | Class |
|---|---|---|
| Asset library | 2,164 objects across 163 categories [benchmarks-r32] | vla-manipulation-eval |
| Claim boundary | 100 task categories are not RoboTwin’s bimanual tasks; performance tables stay in the paper [benchmarks-r32] | vla-manipulation-eval |
| Embodiment in the source | standard evaluation: 7-DoF Franka Emika Panda with parallel gripper; framework also supports other forms [benchmarks-r32] | vla-manipulation-eval |
| Engine / renderer | MuJoCo + dm_control [benchmarks-r32] | vla-manipulation-eval |
| Evaluation scope | 6 capability dimensions: mesh / texture, spatial, common sense / knowledge, semantic instruction, physical laws, and long-horizon reasoning [benchmarks-r32] | vla-manipulation-eval |
| First public date | arXiv 2024-12-24 [benchmarks-r32] | vla-manipulation-eval |
| Kind | language-conditioned manipulation benchmark for VLAs, foundation-model workflows, and VLMs [benchmarks-r32] | vla-manipulation-eval |
| Language-conditioned | yes — natural instructions include implicit intent, common sense, and multi-step reasoning [benchmarks-r32] | vla-manipulation-eval |
| Medium | simulation [benchmarks-r32] | vla-manipulation-eval |
| Paper | arXiv:2412.18194; ICCV 2025 [benchmarks-r32] [benchmarks-r33] | vla-manipulation-eval |
| Primary metric family | interactive task success + graduated Progress Score; non-interactive VLM action-sequence matching [benchmarks-r32] | vla-manipulation-eval |
| Public code | yes — MIT-licensed official repository [benchmarks-r33] | vla-manipulation-eval |
| Public data | yes — evaluation episodes and fine-tuning datasets are linked from the official project [benchmarks-r33] | vla-manipulation-eval |
| Public leaderboard | incomplete — the official repository still lists standard VLA / VLM leaderboard integration as open [benchmarks-r33] | vla-manipulation-eval |
| Released trajectories | standard evaluation episodes and primitive-task fine-tuning data are public; no one suite-wide trajectory count is stated [benchmarks-r33] | vla-manipulation-eval |
| Task structure | 100 task categories: 60 primitive + 40 composite [benchmarks-r32] | vla-manipulation-eval |
Source ledger
- Research VLABench — “VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks”. Type: research paper. ID: arXiv:2412.18194; ICCV 2025. Retrieved: 2026-09-04. reviewed 2026-09-04
- Research VLABench — official project site and linked MIT-licensed repository. Type: research project page. Retrieved: 2026-09-04. reviewed 2026-09-04
What is not inferred
No fuzzy entity matching, adjacent-column borrowing, universal compatibility, commercial maturity, or checkpoint-to-product identity. Absence of a relation means unreviewed or unsupported here—not proven incompatibility.