Robotics benchmarks

This page catalogues evaluation-suite identity: what a named benchmark or corpus is, when it appeared, and the scale the paper actually publishes. It is not a robot ranking, not a method leaderboard, and not a hardware sheet. Incomparable units stay in separate tables. Empty cells stay empty.

Class definition

A column here is a named evaluation suite, dataset, or training corpus. A row is an identity or scale field with one meaning. Policy success rates, hardware payload, and IFR installation counts live elsewhere. Galbot G1, Unitree G1, and Unitree H1 are different bodies; a suite that names one of them does not move the others. Announced robots without a sheet do not enter these tables.

One class per table

Closed-loop sim manipulation, VLA-focused manipulation, real-to-sim eval, grasp synthesis, VLN, tracking, spatial VQA, embodied QA, humanoid tracking, humanoid whole-body sim, and training corpora do not share a “task count” column.

Paper scale, not a ranking

Published counts stay in the paper’s unit: tasks, episodes, grasps, hours, frames. Derived charts on this page count named columns, not performance.

Evidence grade R

Suite identity is a research claim. Official competition pages are manuals for status only. means the inspected source did not publish a comparable figure.

Not a table row. Galbot ET1 has no public eval sheet. AstraBrain-Agent is a launch architecture. LDA-1B, WAM-TTT, LATENT, and Humanoid-GPT explain research lineage; they are not copied into these suite columns as product checkpoints.

Coverage on this page

These totals are derived from the tables below. They are not hub data-stat counters and not a quality score. “Task” is not a shared unit across tables.

Identity tables
16
One data-class each
Named suite columns
28
Eval suites and corpora, not robots
Role groups
8
Charts count columns, not scores
Training corpora
5
Hours ≠ trajectories ≠ frames

Retrieved from the papers and official project pages cited in each caption.

Suite map

Physical form still controls hardware comparability on other pages. Here the split is evaluation role. Cards jump to the owning table.

Simulation columns: the six closed-loop manipulation suites, VLABench, RoboTwin 2.0, BEHAVIOR-1K, ALFRED, EVT-Bench, and HumanoidBench. Real or real-imagery: SIMPLER’s two setups, FurnitureBench, Open6DOR, Room-to-Room, and OpenEQA. Dataset / mocap / corpus: two DexGraspNet columns, GAPartNet, HumanTracker, and the five training corpora. OmniSpatial is image QA.

Manipulation evals

Closed-loop simulation, VLA-focused simulation, paired real-to-sim, and real furniture assembly stay in four tables. Do not add their task counts.

Closed-loop simulated manipulation

Sources: RLBench [R1]; RLBench site [R2]; CALVIN [R3]; CALVIN code [R4]; LIBERO [R5]; LIBERO site [R6]; ManiSkill2 [R7]; ManiSkill2 site [R8]; Meta-World [R9]; Meta-World site [R10]; RoboCasa [R11]; RoboCasa site [R12]. A “task” here is each paper’s own unit. Do not join these columns.
MetricRLBenchCALVINLIBEROManiSkill2Meta-WorldRoboCasa
Identity
Kind#closed-loop simulated manipulation benchmark [R1]language-conditioned simulated manipulation benchmark [R3]lifelong language-conditioned simulated manipulation benchmark [R5]simulated manipulation benchmark across embodiments and object sets [R7]multi-task / meta-learning simulated manipulation benchmark [R9]kitchen simulation framework and evaluation task set [R11]
First public date#arXiv 2019-09-26 [R1]arXiv 2021-12-06 [R3]arXiv 2023-06-05 [R5]arXiv 2023-02-09 [R7]arXiv 2019-10-24 [R9]arXiv 2024-06-04 [R11]
Paper#arXiv:1909.12271 [R1]arXiv:2112.03227 [R3]arXiv:2306.03310 [R5]arXiv:2302.04659 [R7]arXiv:1910.10897 [R9]arXiv:2406.02523 [R11]
Scale
Published scale#100 unique hand-designed tasks [R1]4 environments; 34 subtasks versus 18 prior; ~24 h play data; 20K language directives [R3]130 language-conditioned tasks in 4 suites (Spatial, Object, Goal: 10 tasks each; LIBERO-100) [R5]20 task families; 2000+ objects; 4 million+ demonstration frames [R7]50 distinct manipulation tasks [R9]100 tasks (25 atomic + 75 composite); 120 kitchen scenes; 2,509 objects; 153 categories; 100K trajectories [R11]
Medium#simulation [R1]simulation [R3]simulation [R5]simulation [R7]simulation [R9]kitchen simulation [R11]
Embodiment in the source#Franka Panda [R1]Franka Panda 7-DOF [R3]tabletop manipulator in the paper’s suites [R5]stationary or mobile-base; single-arm or dual-arm; rigid or soft-body [R7]Sawyer [R9]kitchen robots in the paper, including mobile platforms and humanoid robots [R11]
Engine / renderer#V-REP (CoppeliaSim) [R1]PyBullet [R3]SAPIEN [R7]MuJoCo [R9]simulation framework on the project site [R12]
Primary metric family#task success on the paper’s 100-task set [R1]multi-step language instruction success in the paper’s protocol [R3]lifelong / sequential transfer success on the four suites [R5]task-family success in the paper’s protocol [R7]multi-task and meta-learning success on the 50-task set [R9]atomic and composite kitchen-task success in the paper’s protocol [R11]
Language-conditioned#task names and descriptions; not a CALVIN-style language-conditioned suite [R1]yes — 20K language directives [R3]yes — 130 language-conditioned tasks [R5]not published as a language-instruction suite [R7]not published as a language-instruction suite [R9]yes — language goals in the paper’s BC policy; composite tasks from LLM guidance [R11]
Access
Public code#yes — official project publishes code [R2]yes — official project publishes code [R4]yes — official project publishes code [R6]yes — official project publishes code [R8]yes — official project publishes code [R10]yes — official project publishes code [R12]
Public leaderboard#
Claim boundary#identity and published scale only; not a method leaderboard [R1]identity and published scale only; not a method leaderboard [R3]paper training protocol is 50 demonstrations per task; that is a protocol, not a second task count, and not a 65k-demo figure from third-party pages [R5]identity and published scale only; “task family” is not the same unit as RLBench’s 100 tasks [R7]identity and published scale only; Sawyer in simulation is not a product SKU column [R9]body figures 2,509 objects and 153 categories; abstract rounds to 2,500+ / 150+. Not a success-rate table [R11]

VLA-focused manipulation

These suites target different failure surfaces: VLABench emphasizes language, world knowledge, and long-horizon reasoning on a standard single arm; RoboTwin 2.0 emphasizes domain randomization, bimanual data generation, and sim-to-real robustness. Their success rates are not joined.

Sources: VLABench paper [R32]; VLABench project [R33]; RoboTwin 2.0 paper [R34]; RoboTwin 2.0 project [R35]. Identity and protocol only; method scores remain in the papers and official leaderboards.
MetricVLABenchRoboTwin 2.0
Identity
Kind#language-conditioned manipulation benchmark for VLAs, foundation-model workflows, and VLMs [R32]scalable bimanual data generator and policy-generalization benchmark [R34]
First public date#arXiv 2024-12-24 [R32]arXiv 2025-06-22 [R34]
Paper#arXiv:2412.18194; ICCV 2025 [R32] [R33]arXiv:2506.18088 [R34]
Scale and protocol
Task structure#100 task categories: 60 primitive + 40 composite [R32]50 dual-arm tasks in the paper; the project describes over 50 supported tasks [R34] [R35]
Asset library#2,164 objects across 163 categories [R32]RoboTwin-OD: 731 objects across 147 categories [R34]
Released trajectories#standard evaluation episodes and primitive-task fine-tuning data are public; no one suite-wide trajectory count is stated [R33]100,000+ pre-collected expert trajectories across 50 tasks and 5 embodiments [R34]
Evaluation scope#6 capability dimensions: mesh / texture, spatial, common sense / knowledge, semantic instruction, physical laws, and long-horizon reasoning [R32]5 randomization axes: clutter, lighting, background, tabletop height, and language; clean versus randomized policy evaluation [R34]
Medium#simulation [R32]simulation benchmark and data generator with a separate sim-to-real study [R34]
Embodiment in the source#standard evaluation: 7-DoF Franka Emika Panda with parallel gripper; framework also supports other forms [R32]dual-arm pairings of Aloha-AgileX, ARX-X5, Piper, Franka, and UR5 manipulators [R34]
Engine / renderer#MuJoCo + dm_control [R32]
Primary metric family#interactive task success + graduated Progress Score; non-interactive VLM action-sequence matching [R32]task success under clean and hard randomized conditions on the full 50-task protocol [R34]
Language-conditioned#yes — natural instructions include implicit intent, common sense, and multi-step reasoning [R32]yes — trajectory-level language is one randomization axis [R34]
Access
Public code#yes — MIT-licensed official repository [R33]yes — MIT-licensed official repository [R35]
Public data#yes — evaluation episodes and fine-tuning datasets are linked from the official project [R33]yes — object library and pre-collected trajectories are linked from the official project [R35]
Public leaderboard#incomplete — the official repository still lists standard VLA / VLM leaderboard integration as open [R33]yes — official leaderboard linked from the project [R35]
Claim boundary#100 task categories are not RoboTwin’s bimanual tasks; performance tables stay in the paper [R32]50 benchmark tasks are distinct from the 10-task code-generation study and 4-task real-world study; those method scores stay in the paper [R34]

Real-to-sim (SIMPLER)

Sources: SIMPLER paper [R13]; SIMPLER site [R14]. Columns are the two real-robot setups named in the paper. No policy success rates are copied.
MetricGoogle RobotWidowX
Identity
Kind#real-to-sim evaluation environment for RT-series setups [R13]real-to-sim evaluation environment for the BridgeData V2 WidowX setup [R13]
First public date#arXiv 2024-05-09 [R13]arXiv 2024-05-09 [R13]
Paper#arXiv:2405.05941 [R13]arXiv:2405.05941 [R13]
Scale
Published scale#~1,500 paired sim-and-real evaluation episodes across both embodiments in the paper, not a per-robot task count [R13]~1,500 paired sim-and-real evaluation episodes across both embodiments in the paper, not a per-robot task count [R13]
Medium#paired real and simulated evaluation [R13]paired real and simulated evaluation [R13]
Embodiment in the source#Google Robot (RT-series evaluation setup) [R13]WidowX / BridgeData V2 setup [R13]
Engine / renderer#open-sourced SIMPLER environments (SAPIEN in the default paper setup) [R13]open-sourced SIMPLER environments [R13]
Primary metric family#correlation of simulated vs real policy success in the paper; not copied here [R13]correlation of simulated vs real policy success in the paper; not copied here [R13]
Language-conditioned#yes — language-conditioned Google Robot tasks in the paper [R13]yes — BridgeData V2 WidowX tasks in the paper [R13]
Access
Public code#yes — official project publishes code [R14]yes — official project publishes code [R14]
Public leaderboard#
Claim boundary#setup identity only. Do not read the paper’s policy tables into this atlas [R13]setup identity only. WidowX here is the Bridge evaluation body, not a product SKU [R13]

Real furniture assembly

Source: FurnitureBench [R15]. Real furniture assembly is not a simulated kitchen task and not SIMPLER.
MetricFurnitureBench
Identity
Kind#real furniture-assembly benchmark with teleoperation demonstrations [R15]
First public date#arXiv 2023-05-22 [R15]
Paper#arXiv:2305.12821 [R15]
Scale
Published scale#8 furniture models; 5,000+ teleoperation demonstrations; 200+ hours [R15]
Medium#real [R15]
Embodiment in the source#real robot furniture assembly as described in the paper [R15]
Engine / renderer#
Primary metric family#assembly success on the paper’s furniture set [R15]
Language-conditioned#
Access
Public code#yes — paper points to public code [R15]
Public leaderboard#
Claim boundary#identity and published scale only; not a simulated kitchen suite [R15]

Everyday activity sim

BEHAVIOR-1K publishes activities, scenes, and objects. That is not CALVIN’s language-directive count.

Sources: BEHAVIOR-1K [R16]; BEHAVIOR site [R17]. An “activity” here is not a ManiSkill2 task family.
MetricBEHAVIOR-1K
Identity
Kind#everyday household-activity simulation benchmark [R16]
First public date#arXiv 2024-03-14 [R16]
Paper#arXiv:2403.09227 [R16]
Scale
Published scale#1,000 activities; 50 scenes; 9,000+ objects [R16]
Medium#simulation [R16]
Embodiment in the source#simulated household agents in the paper’s scenes [R16]
Engine / renderer#OmniGibson [R16]
Primary metric family#activity success in the paper’s protocol [R16]
Language-conditioned#activities are named household tasks; not mixed into CALVIN’s 20K directives [R16]
Access
Public code#yes — official project publishes code [R17]
Public leaderboard#
Claim boundary#identity and published scale only; 1,000 activities are not 100 RLBench tasks [R16]

Grasp and articulated parts

Synthesis datasets are not closed-loop policy suites. DexGraspNet 2.0 uses a LEAP hand in the paper; that is not a robot payload cell.

Sources: DexGraspNet [R39]; DexGraspNet 2.0 [R40]. Grasp counts are dataset scale, not policy success.
MetricDexGraspNetDexGraspNet 2.0
Identity
Kind#large-scale synthetic dexterous grasp dataset [R39]generative dexterous grasp dataset in synthetic clutter [R40]
First public date#arXiv 2022-10-06 [R39]arXiv 2024-10-30 [R40]
Paper#arXiv:2210.02697 [R39]arXiv:2410.23004 [R40]
Scale
Published scale#1.32 million grasps; 5,355 objects [R39]1,319 objects; 8,270 scenes; 427 million grasps [R40]
Medium#synthetic [R39]synthetic clutter [R40]
Embodiment in the source#ShadowHand [R39]LEAP hand [R40]
Engine / renderer#
Primary metric family#grasp synthesis dataset; not a closed-loop policy suite [R39]grasp synthesis dataset; not a closed-loop policy suite [R40]
Language-conditioned#
Access
Public code#yes — paper and project publish code [R39]yes — paper and project publish code [R40]
Public leaderboard#
Claim boundary#dataset identity only. Author method success rates are not this table [R39]LEAP hand here is the dataset’s synthesis hand, not a robot payload cell [R40]

Articulated parts

Source: GAPartNet [R41]. Part instances are not DexGraspNet grasps.
MetricGAPartNet
Identity
Kind#cross-category articulated-part perception and manipulation dataset [R41]
First public date#arXiv 2022-11-09 [R41]
Paper#arXiv:2211.05272 [R41]
Scale
Published scale#9 GAPart classes; 27 object categories; 8,489 part instances; 1,166 objects [R41]
Medium#dataset of articulated objects [R41]
Embodiment in the source#
Engine / renderer#
Primary metric family#part-class perception / actionable-part evaluation in the paper [R41]
Language-conditioned#
Access
Public code#yes — paper and project publish code [R41]
Public leaderboard#
Claim boundary#part instances are not grasp counts and not a robot SKU [R41]

Instruction, rearrangement, navigation

ALFRED is household instruction in AI2-THOR. Open6DOR is open-instruction 6-DoF rearrangement with no public scale here. Room-to-Room is VLN on real indoor scans.

Household instruction

Source: ALFRED [R18]. AI2-THOR simulation. Not Open6DOR real rearrangement and not R2R navigation.
MetricALFRED
Identity
Kind#household instruction-following benchmark in interactive visual simulation [R18]
First public date#arXiv 2019-12-03 [R18]
Paper#arXiv:1912.01734 [R18]
Scale
Published scale#25,743 English language directives; 8,055 expert demonstrations; 120 indoor scenes [R18]
Medium#simulation (AI2-THOR) [R18]
Embodiment in the source#egocentric household agent in AI2-THOR 2.0 [R18]
Engine / renderer#AI2-THOR 2.0 [R18]
Primary metric family#goal and step-instruction success in the paper’s protocol [R18]
Language-conditioned#yes — high-level goals and low-level instructions [R18]
Access
Public code#yes — paper publishes the dataset and evaluation [R18]
Public leaderboard#
Claim boundary#body count is 25,743 directives; abstract says 25k. Simulated household instruction, not real 6-DoF rearrangement [R18]

Open-instruction rearrangement

Source: Open6DOR project [R38]. No public arXiv was posted on the lab homepage. Unpublished scale stays empty.
MetricOpen6DOR
Identity
Kind#open-instruction 6-DoF object rearrangement benchmark [R38]
First public date#IROS 2024 Oral [R38]
Paper#
Scale
Published scale#
Medium#real rearrangement evaluation as described on the project page [R38]
Embodiment in the source#
Engine / renderer#
Primary metric family#open-instruction placement, not pick-only grasping [R38]
Language-conditioned#yes — open language instructions [R38]
Access
Public code#project page [R38]
Public leaderboard#
Claim boundary#no arXiv on the homepage; scale cells stay empty rather than guessed [R38]

Vision-and-language navigation

Sources: Room-to-Room [R19]; R2R site [R20]. Real Matterport3D imagery. Not Habitat the platform, and not ALFRED.
MetricRoom-to-Room
Identity
Kind#vision-and-language navigation on real indoor scans [R19]
First public date#arXiv 2017-11-20 [R19]
Paper#arXiv:1711.07280 [R19]
Scale
Published scale#21,567 instructions; average 29 words [R19]
Medium#real Matterport3D indoor imagery [R19]
Embodiment in the source#navigation agent on Matterport3D scans [R19]
Engine / renderer#Matterport3D [R19]
Primary metric family#instruction-following navigation success in the paper’s protocol [R19]
Language-conditioned#yes — English navigation instructions [R19]
Access
Public code#yes — official project publishes the dataset [R20]
Public leaderboard#
Claim boundary#R2R is a VLN dataset on real scans. Habitat is a later simulator stack and is not this column [R19]

Tracking, spatial VQA, embodied QA

EVT-Bench episode counts are not TrackVLA’s 1.7 million training samples. OmniSpatial is not OpenEQA. ERQA’s 400 questions stay on the foundation-models page.

Embodied visual tracking

Source: TrackVLA [R19]. Episode counts are EVT-Bench. 1.7 million is TrackVLA’s training mixture, not the eval suite.
MetricEVT-Bench
Identity
Kind#embodied visual tracking benchmark introduced with TrackVLA [R19]
First public date#arXiv 2025-05-29 [R19]
Paper#arXiv:2505.23189 [R19]
Scale
Published scale#25,986 episodes; 100 avatars; 804 scenes (HM3D + MP3D); train 21,771 / test 4,215; three sub-tasks [R19]
Medium#simulation (HM3D and MP3D scenes) [R19]
Embodiment in the source#simulated tracking agent; 100 humanoid avatars in the benchmark [R19]
Engine / renderer#
Primary metric family#success rate, episode length, tracking rate, and collision rate in the paper [R19]
Language-conditioned#yes — language descriptions of the target [R19]
Access
Public code#paper states EVT-Bench would be made public [R19]
Public leaderboard#
Claim boundary#1.7 million samples in the TrackVLA abstract is the training mixture (recognition + tracking), not the EVT-Bench episode count [R19]

Spatial VLM benchmark

Source: OmniSpatial [R37]. VLM spatial QA. Not OpenEQA and not a motor-control suite.
MetricOmniSpatial
Identity
Kind#spatial-reasoning benchmark for vision-language models [R37]
First public date#arXiv 2025-06-03 [R37]
Paper#arXiv:2506.03135 [R37]
Scale
Published scale#4 categories; 50 subcategories; 8.4K+ QA pairs [R37]
Medium#image / VQA benchmark [R37]
Embodiment in the source#
Engine / renderer#
Primary metric family#spatial QA accuracy in the paper [R37]
Language-conditioned#yes — language questions over images [R37]
Access
Public code#yes — paper and project publish code [R37]
Public leaderboard#
Claim boundary#not an embodied motor suite and not OpenEQA [R37]

Embodied question answering

Source: OpenEQA project [R21]. CVPR 2024. The project bibtex has no arXiv id; none is invented.
MetricOpenEQA
Identity
Kind#open-vocabulary embodied question answering benchmark [R21]
First public date#CVPR 2024 [R21]
Paper#
Scale
Published scale#1,600+ human-generated questions; 180+ real-world environments [R21]
Medium#real environments (episodic memory and active exploration settings) [R21]
Embodiment in the source#EQA agent as defined on the project page [R21]
Engine / renderer#
Primary metric family#LLM-Match agreement with human answers on the project protocol [R21]
Language-conditioned#yes — untemplated natural-language questions [R21]
Access
Public code#project page [R21]
Public leaderboard#
Claim boundary#no arXiv id in the project bibtex. Not OmniSpatial and not ERQA’s 400 Gemini questions [R21]

Humanoid motion

HumanTracker scores trackers on mocap. HumanoidBench is a MuJoCo whole-body task suite whose primary agent is Unitree H1 with two Shadow Hands. They do not share a task column.

HumanTracker

Source: HumanTracker [R36]. Motion-tracking evaluation. Not HumanoidBench locomotion/manipulation tasks.
MetricHumanTracker
Identity
Kind#humanoid motion-tracking benchmark with a preference-aligned metric [R36]
First public date#arXiv 2026-08-24 [R36]
Paper#arXiv:2608.13555 [R36]
Scale
Published scale#~153 hours optical mocap; 4 motion families; ~25K clips; HumanScore trained on 12K pairs / 24K motions [R36]
Medium#optical motion capture retargeted to a simulated humanoid [R36]
Embodiment in the source#29-DoF humanoid qpos representation in the paper’s evaluator [R36]
Engine / renderer#MuJoCo evaluation entry point in the paper [R36]
Primary metric family#completion, MPJPE, and HumanScore (0–100) in the paper [R36]
Language-conditioned#text labels on clips; not a household instruction suite [R36]
Access
Public code#yes — paper lists official code [R36]
Public leaderboard#
Claim boundary#Humanoid-GPT is one evaluated tracker, not the benchmark body. This is not HumanoidBench and not a Galbot product column [R36]

HumanoidBench

Sources: HumanoidBench [R22]; HumanoidBench site [R23]. Primary agent in the paper is Unitree H1 with two Shadow Hands. Not Galbot G1 and not Unitree G1 as the primary body.
MetricHumanoidBench
Identity
Kind#simulated whole-body locomotion and manipulation benchmark [R22]
First public date#arXiv 2024-03-15 [R22]
Paper#arXiv:2403.10506 [R22]
Scale
Published scale#15 manipulation + 12 locomotion tasks (27 tasks) [R22]
Medium#simulation [R22]
Embodiment in the source#Unitree H1 with two Shadow Hands as the primary agent [R22]
Engine / renderer#MuJoCo [R22]
Primary metric family#RL task return / success on the paper’s 27-task set [R22]
Language-conditioned#
Access
Public code#yes — official project publishes code [R23]
Public leaderboard#
Claim boundary#optional other simulator models in the paper (including a Unitree G1 model) are not this column and are not Galbot G1. Not HumanTracker [R22]

Training corpora

These rows are pretraining mixtures, not evaluation leaderboards. Do not add hours to trajectories or frames.

Sources: Open X-Embodiment [R24]; OXE site [R25]; DROID [R26]; DROID site [R27]; BridgeData V2 [R28]; BridgeData site [R29]; LDA-1B / EI-30k [R9]; GraspVLA / SynGrasp-1B [R17]. Hours, trajectories, and frames are different units. Do not add them.
MetricOpen X-EmbodimentDROIDBridgeData V2EI-30kSynGrasp-1B
Identity
Kind#multi-embodiment robot-learning dataset pool [R24]in-the-wild robot manipulation dataset [R26]multi-environment robot demonstration dataset [R28]mixed human and robot embodied pretraining corpus used with LDA-1B [R9]synthetic VLA pretraining frames used with GraspVLA [R17]
First public date#arXiv 2023-10-13 [R24]arXiv 2024-03-19 [R26]arXiv 2023-08-24 [R28]arXiv 2026-02-16 [R9]arXiv 2025-05-06 [R17]
Paper#arXiv:2310.08864 [R24]arXiv:2403.12945 [R26]arXiv:2308.12952 [R28]arXiv:2602.12215 [R9]arXiv:2505.03233 [R17]
Scale
Published scale#1M+ trajectories; 22 embodiments; 60 datasets; 34 labs; 527 skills; 160266 tasks (body). Abstract: 22 robots, 21 institutions, 527 skills, 160266 tasks [R24]76k trajectories; 350 hours; 564 scenes; 86 tasks [R26]60,096 trajectories; 24 environments [R28]30k+ hours mixed-quality human and robot data [R9]1 billion synthetic VLA frames [R17]
Medium#real robot trajectories pooled from many datasets [R24]real teleoperated trajectories [R26]real demonstrations [R28]mixed real, sim, labeled, and unlabeled human and robot data [R9]synthetic [R17]
Embodiment in the source#22 embodiments: single arms, bi-manual robots, and quadrupeds [R24]Franka Panda stack [R26]WidowX-class Bridge robots in the paper [R28]human video plus robot trajectories; LDA-1B real figures use Galbot G1 and Unitree G1 variants [R9]synthetic grasping VLA frames; not a named robot SKU [R17]
Engine / renderer#RLDS / OXE distribution [R25]DROID distribution [R27]BridgeData V2 distribution [R29]LeRobot-format corpus described in the LDA paper [R9]
Primary metric family#pretraining corpus; not an eval leaderboard [R24]pretraining corpus; not an eval leaderboard [R26]pretraining corpus; not an eval leaderboard [R28]pretraining corpus for a world-action model [R9]pretraining frames for GraspVLA [R17]
Language-conditioned#language annotations vary by constituent dataset [R24]language-annotated tasks in the paper [R26]language-annotated demonstrations in the paper [R28]mixed labeled and unlabeled [R9]synthetic VLA frames with action data in the paper [R17]
Access
Public code#yes — official project publishes the dataset [R25]yes — official project publishes the dataset [R27]yes — official project publishes the dataset [R29]described with LDA-1B; public dump as stated in that paper [R9]promised with GraspVLA; check the project page for release [R17]
Public leaderboard#
Claim boundary#abstract uses 21 institutions; body uses 34 labs / 60 datasets. Both stay attributed. Not an evaluation suite [R24]76k trajectories and 350 hours are the same corpus in two units, not two datasets. Abstract collector counts are not reconciled here [R26]not SIMPLER’s WidowX eval column [R28]hours are not trajectories. Not EVT-Bench [R9]frames are not hours or trajectories. Not DexGraspNet’s 1.32 million grasps [R17]

Lab-owned special-env pins

Scale for HumanTracker, OmniSpatial, Open6DOR, DexGraspNet, GAPartNet, EVT-Bench, and EI-30k lives in the tables above. These pins stay on Research so existing permalinks keep working. UrbanVLA’s outdoor eval is the paper’s own suite, not a separately named public benchmark.

Simulation platforms

A renderer is not a VLN dataset. Habitat numbers stay on this card so they are not mixed into Room-to-Room.

Habitat-Sim

Platform paper: Habitat-Sim is a high-performance 3D simulator. On a Matterport3D scene it reports several thousand frames per second single-threaded, and over 10,000 fps multi-process on one GPU. Habitat-API defines tasks such as navigation, instruction following, and question answering on top of that simulator. That is not the Room-to-Room instruction count. [R30] [R31]

Competitions and field trials

Official organisation pages only. Status and identity, not scores, payloads, or invented participant counts.

RoboCup

Ongoing robot-soccer and related leagues run by the RoboCup Federation. [M1]

DARPA Robotics Challenge

Completed disaster-response trial series. Historical programme identity only. [M2]

ANA Avatar XPRIZE

Completed avatar / telepresence prize. [M3]

CYBATHLON

ETH Zurich assistive-device championship. [M4]

EUROBENCH

EU benchmarking framework for biped systems; organisation page, not a spec table. [M5]

How this atlas connects

Search indexes this page once the landscape export is rebuilt. Compare can open a single table; it must not join “task” rows across classes. Research keeps media pins. Foundation-model pages keep policy scores.