Robot foundation models

This page is for vision-language-action and world-action policies — parameter counts, licenses, and public artifacts. It is not the 3D / URDF catalogue and not the SDK / simulator page. Most cells are declared gaps. That is the finding: a press release with no model card is published as an empty cell, not as an invented spec.

Class definition

A row here is a named policy or “robot brain,” not a robot body. Hardware numbers stay on the class pages. A company demo is not a benchmark table. A blog parameter count is press unless a paper or model card repeats it. Weights without a license, or a license that conflicts with the backbone, stay written as such.

Identity. Genesis AI (the company; GENE-26.5 / Eno / Genesis World 1.0) is not the Apache-licensed Genesis-Embodied-AI/genesis-world physics engine, even though the company now supports that repository. GENE-26.5 is not a public checkpoint identity for Galbot GraspVLA, LDA-1B, or GroceryVLA. GR00T N1 evaluations named in the paper use Fourier GR-1 and thank 1X; they do not name Galbot bodies or Genesis Eno.

Genesis AI — the lab

Stealth exit

Genesis AI emerged from stealth on 2025-07-01 with a $105M seed co-led by Eclipse and Khosla Ventures. Participating names in contemporaneous coverage include Bpifrance, HSG, Eric Schmidt, and Xavier Niel. CEO Zhou Xian (CMU robotics Ph.D., 2024); co-founder Théophile Gervet (ex-Mistral; press also names Skild AI). Stated goal: a universal robotics foundation model plus a horizontal platform. [P1]

GENE-26.5

Announced 2026-05-06 as the company’s VLA / “robot brain.” Company materials describe a human-scale five-finger hand and a tactile e-skin data-collection glove with a claimed 1:1:1 glove→human-hand→robot-hand mapping, “100× cheaper” than typical teleop rigs, and up to 5× data-collection efficiency (company / internal testing). The company blog assigns Genesis Hand 1.0 “20 active, back-drivable degrees of freedom.” That figure is a hand claim, not a robot-body DoF, and is not copied into any hardware table. [P2] [P3]

Public unknowns for GENE-26.5. No parameter count, success rates, benchmark table, API, model card, or released weights were public when checked. Demonstrations only: 20-step meal prep (including one-handed egg crack), pipetting / liquid transfer, wire harnessing, Rubik’s cube, piano, and a four-object one-handed grasp. Eno is the later general-purpose body and lives on the announced page — no numeric column.

Policy models — what is actually public

One comparability class: named policy models. Empty cells are intentional. Helix architecture numbers come from Figure’s company post, not a paper. Skild’s S1 post describes an in-context learner and does not publish a parameter count or weight dump. Later GR00T N1.6 / N1.7 labels exist; this column is the N1-2B paper checkpoint only.

Sources: TechCrunch stealth [P1]; GENE-26.5 press [P2]; GENE-26.5 blog [P3]; GR00T N1 [R1]; Isaac-GR00T [M1]; OpenVLA [R2]; OpenVLA project [M2]; π0.5 [R3]; openpi [M3]; Skild [P4]; Figure Helix [P5].
Metric GENE-26.5 GR00T N1-2B OpenVLA 7B π0.5 Skild S1 Figure Helix
Identity
Kind#company VLA / “robot brain” [P2]open VLA; dual-system VLM + DiT flow matching [R1]open VLA on Llama 2 + SigLIP/DINOv2 [R2]VLA; PaliGemma backbone + action expert [R3]company “omni-bodied” foundation model; S1 is an in-context learner [P4]company VLA; System 2 VLM + System 1 visuomotor policy [P5]
Public date#2026-05-06 [P2]arXiv 2025-03-18 [R1]arXiv 2024-06-13; CoRL 2025 [R2]arXiv 2025-04-22; CoRL 2025 Oral [R3]S1 post August 2026 [P4]2025-02-20 [P5]
Paper#arXiv:2503.14734 [R1]arXiv:2406.09246 [R2]arXiv:2504.16054 [R3]
Scale (as published)
Parameter count#2.2B total; 1.34B in the VLM [R1]7B [R2]2B VLM (PaliGemma) + 300M action expert [R3]S2 7B VLM; S1 80M [P5]
Stated training data#real-robot trajectories, human video, synthetic; Fourier GR-1 collect named [R1]970k real-robot episodes from Open X-Embodiment [R2]web + heterogeneous robot data; experiment results only [R3]~500 hours teleop; hindsight language labels [P5]
Published latency / rate#16-action chunk 63.9 ms on L40, bf16 [R1]S2 7–9 Hz; S1 200 Hz (company post) [P5]
Public artifacts
Weights#released with Isaac-GR00T; NVIDIA Open Model License [M1]HF openvla/openvla-7b; backbone is Llama 2 (Community License) [M2] [R2]openpi checkpoints; not a product API [M3]
Code license#Apache-2.0 [M1]project and training code published; see repository [M2]Apache-2.0 (Physical-Intelligence/openpi) [M3]
Public API#
Public benchmark table#paper tables only [R1]paper tables only [R2]paper tables only [R3]
Stated embodiments in the source#company demos on Genesis hardware; Eno announced later [P2]Fourier GR-1 collect; 1X thanked for hardware [R1]Open X-Embodiment mixture (many arms) [R2]Physical Intelligence mobile-manipulation experiments [R3]company copy: many morphologies; no public eval table [P4]Figure humanoids (company post) [P5]

OpenVLA’s language backbone is Llama 2. A Hugging Face card that says “MIT” does not override that Community License for the inherited weights; both facts stay written. Helix 02 is a later Figure post and is not collapsed into this column. π0 (2024 blog) shares the PaliGemma + action-expert shape that π0.5 states explicitly; the parameter row cites the π0.5 paper only.

Smaller open VLAs and papers without public weights

These do not belong in the six-column family table. SmolVLA and SpatialVLA are later open checkpoints or papers. OpenVLA-OFT is a fine-tuning recipe on the existing OpenVLA 7B weights, not a new pretrain. Gemini Robotics is a paper (and company blog) with no parameter count and no public weights. Empty cells stay empty.

Sources: SmolVLA paper [R7]; SmolVLA Hugging Face card [M6]; SpatialVLA paper [R8]; SpatialVLA project [M7]; OpenVLA-OFT paper [R9]; OpenVLA-OFT project [M8]; Gemini Robotics paper [R10]; DeepMind Gemini Robotics blog [P7].
Metric SmolVLA SpatialVLA OpenVLA-OFT Gemini Robotics
Kind#compact VLA on LeRobot; SmolVLM-2 plus a flow-matching action expert [R7]3D spatial VLA with Ego3D positional encoding and Adaptive Action Grids; Paligemma2 backbone [R8]fine-tuning recipe on OpenVLA 7B, not a new pretrain [R9]Gemini Robotics VLA plus Gemini Robotics-ER, built on Gemini 2.0 [R10]
Parameter count#0.5B on the Hugging Face card; paper says less than 0.5 billion [M6] [R7]3.5B [R8]uses the existing OpenVLA 7B weights [R9]
Stated training data#fewer than 30k community episodes [R7]1.1 million real-robot episodes (Open X-Embodiment subset plus RH20T) [R8]
Named result#trains on a single GPU; paper also says consumer-GPU or CPU deploy [R7]LIBERO success 76.5% → 97.1%; 26× action-generation throughput vs the OpenVLA baseline in the paper [R9]Embodied Reasoning Question Answer (ERQA) set of 400 questions [R10]
Weights#HF lerobot/smolvla_base [M6]recipe on openvla/openvla-7b [M2]
Paper#arXiv:2506.01844 [R7]arXiv:2501.15830 [R8]arXiv:2502.19645 [R9]arXiv:2503.20020 [R10]
Public API#

A page of mostly cells is still the finding. Gemini Robotics has a paper and a blog, not a model card. SpatialVLA’s project page is cited for identity; no public weight dump is copied here. OpenVLA-OFT does not replace the OpenVLA 7B column above.

Related EPIC / Galbot models (already catalogued)

These already have pins and a paper→body map on the research page. They are not new columns here. GroceryVLA has no public arXiv.

GraspVLA

CoRL 2025 · arXiv:2505.03233. Synthetic-only grasping VLA; homepage product stack.

LDA-1B

RSS 2026 · arXiv:2602.12215. 1.6B latent world-action model; real figures on Galbot G1 and Unitree G1.

WAM-TTT

preprint · arXiv:2607.06988. Test-time steering of a frozen world-action model from human video.

GroceryVLA

Product model. No public arXiv. Not given an invented paper id.

Prior open policies — still the public baseline

These predate GENE-26.5 and are the papers a later company VLA is usually compared to. They are not Galbot product models and are not collapsed into the family table above. FAST is an action tokenizer (discrete cosine transform plus BPE), not a standalone VLA. The Physical Intelligence /blog/fast URL returned HTTP 404 from this checkout; the research page is cited instead. Empty cells remain empty.

Sources: Octo [R4]; Octo project [M4]; RT-2 [R5]; π0 blog [P6]; RDT-1B [R6]; RDT code [M5]; FAST paper [R11]; FAST research page [M9].
Metric Octo RT-2 π0 (2024 blog) RDT-1B FAST
Kind#open transformer generalist policy [R4]VLA; vision-language model used as a robot policy [R5]company VLA blog; PaliGemma + action expert shape later stated in π0.5 [P6]open diffusion Transformer for bimanual manipulation [R6]frequency-space action tokenizer (DCT + BPE); not a standalone VLA [R11]
Paper / post#arXiv:2405.12213 [R4]arXiv:2307.15818 [R5]company blog (not a paper) [P6]arXiv:2410.07864 [R6]arXiv:2501.09747 [R11]
Parameter count#released checkpoints 27M and 93M [R4]1B (plus a 170M ablation checkpoint) [R6]
Stated training data#800k trajectories from Open X-Embodiment [R4]web-scale VLM pretraining + robot data (paper) [R5]1M+ multi-robot episodes (card / paper) [R6]FAST+ trained on 1 million trajectories; π0-FAST trains on 10k hours [R11]
Named result#paper claims up to 5× faster training than diffusion VLAs when FAST is combined with π0; that is a training-time claim, not a policy parameter count [R11]
Weights#public checkpoints on the project page [M4]later openpi checkpoints belong to the π0.5 column above [P6]HF robotics-diffusion-transformer/rdt-1b; MIT [M5]research page; not a policy checkpoint [M9]
Public API#

Open X-Embodiment (arXiv:2310.08864) is a dataset paper, not a policy column. RT-2 weights were not released with the paper. π0’s parameter split is cited on the π0.5 row above, not invented here from the blog. RDT-1B is Tsinghua TSAIL research, not a Galbot or Genesis product model. Do not invent a FAST policy parameter count.