Research

Research

Each pin is one paper, product model, or benchmark. Grade R: a paper proves only the hardware, checkpoint, and evaluation it names. Galbot G1 is wheeled; Unitree G1 is biped; no paper here names Galbot ET1 as eval hardware.

Recent papers

A pin may embed a video, an official figure, or only the paper. Learned systems stay off hardware tables. GroceryVLA is a product model with no public arXiv. The catalog lists every homepage paper.

preprint 2026 · arXiv:2607.06988

WAM-TTT

Steering World-Action Models by Watching Human Play at Test Time. Test-time training on a frozen world-action model; unlabeled human video only — no robot actions or annotations at deploy.

CVPR 2026 · arXiv:2606.03985

Humanoid-GPT

Zero-shot humanoid motion tracking with a GPT-style causal Transformer trained on a 2B-frame motion corpus.

Real-robot experiments use a 29-DoF Unitree G1 — not Galbot G1 and not Galbot ET1. Nearest open paper to the AstraBrain-WBC 0.5 announcement; not a proven one-to-one checkpoint identity.

IROS 2026 · arXiv:2603.12686

LATENT

Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data. Whole-body returns on a racket-mounted humanoid.

Deployed body is a 29-DoF Unitree G1 — not Galbot G1 and not Galbot ET1. Quantitative table and extra project clips below.

RSS 2026 · arXiv:2602.12215

LDA-1B

Scaling a 1.6B-parameter latent world-action model on mixed-quality human and robot data (EI-30k: sim+real, labeled+unlabeled).

Real-world figures use Galbot G1 (parallel-jaw gripper and Sharpa 22-DoF hand) and Unitree G1 variants — not Galbot ET1.

Robot using non-prehensile extrinsic dexterity to manipulate objects in clutter

RSS 2026 · arXiv:2603.09882 · Best Paper Finalist

DAPL

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning. Non-prehensile pushing and trapping in clutter.

HDFlow hierarchical diffusion-flow planning pipeline diagram

ICML 2026 spotlight · arXiv:2605.04525

HDFlow

Hierarchical Diffusion-Flow Planning for long-horizon tasks.

UrbanVLA route-conditioned micromobility robot in a city scene

ICRA 2026 · arXiv:2510.23576

UrbanVLA

Vision-language-action model for urban micromobility and delivery. Special env: outdoor streets, route-conditioned driving.

TrackVLA++ embodied visual tracking with memory and reasoning

ICRA 2026 · arXiv:2510.07134

TrackVLA++

Reasoning and memory for embodied visual tracking on top of TrackVLA.

SoFar language-grounded object orientation on a real robot

NeurIPS 2025 spotlight · arXiv:2502.13143

SoFar

Language-grounded orientation that bridges spatial reasoning and object manipulation.

FoldNet closed-loop garment folding on a dual-arm robot

RA-L & IROS 2026 · arXiv:2505.09109

FoldNet

Generalizable closed-loop garment folding. Special env: deformable cloth.

Dexonomy taxonomy of dexterous grasp types on diverse objects

RSS 2025 · arXiv:2504.18829

Dexonomy

Synthesizing all dexterous grasp types in a grasp taxonomy.

Uni-NaVid video-based VLA unifying embodied navigation tasks

RSS 2025 · arXiv:2412.06224

Uni-NaVid

Video VLA that unifies embodied navigation tasks. Successor of NaVid; feeds NavFoM.

MM-Nav multi-view visual navigation in cluttered indoor scenes

ECCV 2026 · arXiv:2510.03142

MM-Nav

Multi-view VLA for robust visual navigation via multi-expert learning. Special env: wires, dynamic humans, cluttered corridors.

LIMMT motion-tracking pipeline diagram

ICML 2026 · arXiv:2606.06953

LIMMT

Less is More for Motion Tracking. Companion tracking paper to Humanoid-GPT / HumanTracker.

CL3R 3D reconstruction and contrastive learning for manipulation representations

ICRA 2026 · arXiv:2507.08262

CL3R

3D reconstruction and contrastive learning for robotic manipulation representations.

CoRL 2025 Oral · arXiv:2502.17894

FetchBot

Zero-shot sim2real object fetching in clutter. No official still is linked here; paper and project only.

ICCV 2025 highlight · arXiv:2507.02747

DexVLG

Dexterous vision-language-grasp model at scale. Text pin; project page is the source.

IROS 2026 · arXiv:2603.13832

GraspADMM

Improving dexterous grasp synthesis via ADMM. Follow-on to Dexonomy / BODex.

arXiv 2025 · arXiv:2504.17249

Berkeley Humanoid Lite

Open-source 3D-printed humanoid. Paper hardware numbers and the MuJoCo/URDF repository sit in the dedicated sheet below. Not a Unitree or Galbot body.

CoRL 2025 · arXiv:2502.00893

ToddlerBot

Open-source miniature humanoid for loco-manipulation data collection. Paper hardware numbers sit in the dedicated sheet below. Not a Unitree body.

ECCV 2026 · no arXiv on homepage

UniPart

Zero-shot language-grounded 3D part segmentation for embodied interaction. Not the CVPR 2026 UniPart generation paper by Xufan He.

DyWA non-prehensile manipulation of irregular objects on a table

ICCV 2025 · arXiv:2503.16806

DyWA

Dynamics-adaptive world-action model for generalizable non-prehensile manipulation. Lineage into LDA-1B.

D3RoMa material-agnostic depth sensing on transparent and specular objects

CoRL 2024 · arXiv:2409.14365

D3RoMa

Disparity-diffusion depth for material-agnostic manipulation. Special env: glass, metal, and other depth-hostile surfaces.

RoboHanger inserting hangers into diverse garments

RA-L · arXiv:2412.01083

RoboHanger

Generalizable hanger insertion for diverse garments. Special env: deformable clothing, not rigid pick-and-place.

ScissorBot paper-cutting skill pipeline

CoRL 2024 · arXiv:2409.13966

ScissorBot

Generalizable scissor skill for paper cutting. Special env: thin deformable sheets and a closing blade, not gripper pick.

Paper → body map

A paper proves only the robot it evaluates. A homepage “product model” is not a checkpoint identity. Galbot G1 and Unitree G1 stay in separate tables. Launch architecture names without a public checkpoint card sit outside tables.

Not a table row. AstraBrain-Agent is the named launch architecture for Galbot’s first biped announcement. Public sources do not establish that LDA-1B, WAM-TTT, LATENT, or Humanoid-GPT released checkpoints run on that body. Status belongs on the announced page. [P2]

Evaluated or listed on Galbot G1

Sources: paper/project pins on this page plus hughw19.github.io [R21] and Qiming founder profile [P3]. Vocabulary: arXiv is an id or no public arXiv; eval body is the robot the source names, or that it names none; claim grade is paper eval, homepage product, or not claimed.
SystemarXivWeights / codeEval body in the sourceEnd-effector in the sourceClaim grade
GraspVLA#arXiv:2505.03233 [R17]HF shengliangd/GraspVLA [R17]real-robot evals; robot not named in the paper [R17]homepage product — homepage and founder profile call it the generalist grasp stack on Galbot G1 [R21] [P3]
GroceryVLA#no public arXiv [R21]Galbot G1 [R21]homepage product — retail / pharmacy shelf-pick product model [P3]
TrackVLA / TrackVLA++#arXiv:2505.23189; arXiv:2510.07134 [R19] [R25]embodied visual tracking; Galbot G1 not named in the papers [R19]homepage product — profile lists product-level tracking for unmanned retail [P3]
NavFoM#arXiv:2509.12129 [R15]cross-embodiment: quadruped, drone, vehicle, wheeled platforms [R15]not claimed — no named Galbot G1 SKU [R21]
LDA-1B#arXiv:2602.12215 [R9]Galbot G1 [R9]parallel-jaw gripper; Sharpa 22-DoF hand [R9]paper eval — real-world figures in the paper, not a store SKU [R9]
WAM-TTT#arXiv:2607.06988 [R5]frozen world-action model (LDA line); robot not named in the source [R5]not claimed — no public proof of a named Galbot G1 product checkpoint [R5]
StereoVLA#arXiv:2512.21970 [R11]HF shengliangd/StereoVLA [R11]manipulation VLA; robot not named in the paper [R11]not claimed — weights public; not listed as a Galbot G1 SKU [R11]

Evaluated on Unitree G1

Sources: Humanoid-GPT arXiv:2606.03985 [R7]; LATENT arXiv:2603.12686 [R1]; LDA-1B arXiv:2602.12215 [R9]. These rows are not Galbot G1.
SystemarXivWeights / codeEval body in the sourceEnd-effector in the sourceClaim grade
Humanoid-GPT#arXiv:2606.03985 [R7]GalaxyGeneralRobotics code [R7]29-DoF Unitree G1 [R7]paper eval — nearest open paper to AstraBrain-WBC 0.5, not a proven one-to-one product checkpoint [P4]
LATENT#arXiv:2603.12686 [R1]official code on Galbot GitHub [R1]29-DoF Unitree G1 [R1]paper eval — tennis whole-body control; the announced Galbot biped’s tennis mention is not this paper’s eval [R1]
LDA-1B (BrainCo variant)#arXiv:2602.12215 [R9]Unitree G1 [R9]BrainCo 10-DoF hand [R9]paper eval — not a BrainCo Revo 2 product column [R9]

Lab pins for named suites

Lab-owned datasets and special-env suites. Scale, medium, and embodiment live on the Benchmarks atlas. A pin is not a robot SKU. Permalink ids stay stable.

HumanTracker humanoid motion-tracking benchmark teaser with diverse motions

ECCV 2026 · arXiv:2608.13555 benchmark

HumanTracker

August 2026 motion-tracking benchmark. HumanScore is preference-aligned so eval can match contact, skating, and unstable support rather than only per-frame pose error. Scale is on the Benchmarks atlas. Not HumanoidBench.

OmniSpatial spatial-reasoning task examples for vision-language models

ICLR 2026 · arXiv:2506.03135 benchmark

OmniSpatial

Spatial-reasoning benchmark for VLMs, not motor control. Published scale is on the Benchmarks atlas.

Open6DOR open-instruction 6-DoF object rearrangement scenes

IROS 2024 Oral benchmark

Open6DOR

Open-instruction 6-DoF rearrangement. Language-specified placement pose, not pick-only grasping. Identity sheet: Benchmarks atlas. Unpublished scale stays empty there.

DexGraspNet large-scale synthetic dexterous grasps on many objects

ICRA 2023 · arXiv:2210.02697 dataset

DexGraspNet

Dexterous grasp-synthesis dataset. Multifinger contact, not a parallel-jaw SKU. Grasp counts are on the Benchmarks atlas.

DexGraspNet 2.0 generative dexterous grasps in cluttered scenes

CoRL 2024 · arXiv:2410.23004 dataset

DexGraspNet 2.0

Generative dexterous grasping in synthetic clutter. Grasp counts are on the Benchmarks atlas. The paper’s LEAP hand is not a robot payload cell.

GAPartNet actionable parts on articulated objects such as drawers and doors

CVPR 2023 Highlight · arXiv:2211.05272 dataset

GAPartNet

Part-centric articulated objects: drawers, doors, lids — not rigid free objects. GAPartManip (ICRA 2025) extends the manipulation set. Part counts are on the Benchmarks atlas.

CoRL 2025 · with TrackVLA benchmark

EVT-Bench

Named evaluation suite in TrackVLA. 1.7 million is TrackVLA’s training mixture, not the suite size. Episode counts live on the Benchmarks atlas. Special env: occlusion, night, distractors, long-horizon follow.

UrbanVLA outdoor micromobility evaluation scenes

ICRA 2026 · UrbanVLA eval figure eval suite

UrbanVLA outdoor eval

Official project figure for UrbanVLA’s outdoor / route-conditioned evaluation. Not a separately named public benchmark; it is the paper’s own special-env suite.

DREDS depth restoration on transparent and specular objects

ECCV 2022 · arXiv:2208.03792 dataset

DREDS

RGB-D restoration for transparent and specular surfaces. Special env: depth cameras fail on glass and metal; this set is the usual sim-to-real depth test for those materials.

ASGrasp reconstructing and grasping transparent objects from active stereo RGB-D

ICRA 2024 · arXiv:2405.05648 special env

ASGrasp

Transparent-object reconstruction and 6-DoF grasp from RGB-D active stereo. Pair with DREDS and STOPNet when the scene is glass, not matte household objects.

STOPNet multiview suction detection on transparent objects

ICRA 2024 · arXiv:2310.05717 special env

STOPNet

Multiview 6-DoF suction detection for transparent objects. Special env: suction on glass, not parallel-jaw grasp on opaque parts.

BODex bilevel optimization synthesizing dexterous grasps at scale

ICRA 2025 · arXiv:2412.16490 dataset / method

BODex

Scalable dexterous grasp synthesis with bilevel optimization. Use with DexGraspNet / Dexonomy when the special env is multifinger contact, not parallel-jaw SKUs.

UniDexGrasp++ generalist dexterous grasping across object categories

ICCV 2023 Oral · arXiv:2304.00464 dataset / method

UniDexGrasp++

Generalist dexterous grasping in large-scale synthetic clutter. ICCV 2023 Best Paper Finalist. Lineage into GraspVLA / SynGrasp-1B.

Paper catalog

Homepage papers 2026–2023 plus selected earlier lineage. Pins above are a subset. Refresh from hughw19.github.io. GroceryVLA is a product model, not a paper.

2026

  • WAM-TTTpreprint · arXiv:2607.06988 · pin · paper
  • HumanTrackerECCV 2026 · arXiv:2608.13555 · pin · paper
  • UniPartECCV 2026 · no arXiv on homepage · pin
  • MM-NavECCV 2026 · arXiv:2510.03142 · pin · paper
  • GraspADMMIROS 2026 · arXiv:2603.13832 · pin · paper
  • LATENTIROS 2026 · arXiv:2603.12686 · pin · paper
  • HDFlowICML 2026 spotlight · arXiv:2605.04525 · pin · paper
  • LIMMTICML 2026 · arXiv:2606.06953 · pin · paper
  • LDA-1BRSS 2026 · arXiv:2602.12215 · pin · paper
  • DAPLRSS 2026 Best Paper Finalist · arXiv:2603.09882 · pin · paper
  • StereoVLARSS 2026 · arXiv:2512.21970 · pin · paper
  • Humanoid-GPTCVPR 2026 · arXiv:2606.03985 · pin · paper
  • Layered 4D-Rotor Gaussian SplattingCVPR 2026 · no arXiv on homepage · project
  • CL3RICRA 2026 · arXiv:2507.08262 · pin · paper
  • NavGSimICRA 2026 · arXiv:2603.15186 · paper
  • Robust Differentiable Collision DetectionICRA 2026 · arXiv:2511.06267 · paper
  • UrbanVLAICRA 2026 · arXiv:2510.23576 · pin · paper
  • TrackVLA++ICRA 2026 · arXiv:2510.07134 · pin · paper
  • OmniSpatialICLR 2026 · arXiv:2506.03135 · pin · paper
  • NavFoMICLR 2026 · arXiv:2509.12129 · pin · paper
  • FoldNetRA-L & IROS 2026 · arXiv:2505.09109 · pin · paper
  • DexNDMICLR 2026 · arXiv:2510.08556 · pin · paper

2025

  • SoFarNeurIPS 2025 spotlight · arXiv:2502.13143 · pin · paper
  • TrackVLA / EVT-BenchCoRL 2025 · arXiv:2505.23189 · pin · paper
  • FetchBotCoRL 2025 Oral · arXiv:2502.17894 · pin · paper
  • GraspVLACoRL 2025 · arXiv:2505.03233 · pin · paper
  • DexVLGICCV 2025 highlight · arXiv:2507.02747 · pin · paper
  • DyWAICCV 2025 · arXiv:2503.16806 · pin · paper
  • RoboHangerRA-L · arXiv:2412.01083 · pin · paper
  • DexonomyRSS 2025 · arXiv:2504.18829 · pin · paper
  • Uni-NaVidRSS 2025 · arXiv:2412.06224 · pin · paper
  • Code-as-MonitorCVPR 2025 · arXiv:2412.04455 · paper
  • GAPartManipICRA 2025 · arXiv:2411.18276 · GAPartNet pin · paper
  • BODexICRA 2025 · arXiv:2412.16490 · pin · paper
  • NaVid-4DICRA 2025 · pin · project
  • QuadWBGICRA 2025 · arXiv:2411.06782 · paper
  • Watch Less, Feel MoreICRA 2025 · arXiv:2502.14457 · paper
  • GroceryVLAproduct model · no arXiv · pin

2024

  • D3RoMaCoRL 2024 · arXiv:2409.14365 · pin · paper
  • DexGraspNet 2.0CoRL 2024 · arXiv:2410.23004 · pin · paper
  • ScissorBotCoRL 2024 · arXiv:2409.13966 · pin · paper
  • Task-Oriented Dexterous GraspIROS 2024 · arXiv:2309.13586 · paper
  • Open6DORIROS 2024 Oral · pin · project
  • NaVidRSS 2024 · arXiv:2402.15852 · paper
  • SAGERSS 2024 · arXiv:2312.01307 · paper
  • MaskClusteringCVPR 2024 · arXiv:2401.07745 · paper
  • STOPNetICRA 2024 · arXiv:2310.05717 · pin · paper
  • GAMMAICRA 2024 · arXiv:2309.15459 · paper
  • ASGraspICRA 2024 · arXiv:2405.05648 · pin · paper

2023 and selected earlier lineage

  • UniDexGrasp++ICCV 2023 Oral & Best Paper Finalist · arXiv:2304.00464 · pin · paper
  • GAPartNetCVPR 2023 Highlight · arXiv:2211.05272 · pin · paper
  • 3D-Aware Object Goal NavigationCVPR 2023 · arXiv:2212.00338 · paper
  • UniDexGraspCVPR 2023 · arXiv:2303.00938 · paper
  • PartManipCVPR 2023 · arXiv:2303.16958 · paper
  • Discrete Normalizing Flows on SO(3)CVPR 2023 · arXiv:2304.03937 · paper
  • DiGACVPR 2023 · arXiv:2304.02222 · paper
  • DexGraspNetICRA 2023 · arXiv:2210.02697 · pin · paper
  • GraspNeRFICRA 2023 · arXiv:2210.06575 · paper
  • DREDSECCV 2022 · arXiv:2208.03792 · pin · paper
  • HOI4DCVPR 2022 · arXiv:2203.01577 · paper
  • CAPTRAICCV 2021 Oral · arXiv:2104.03437 · paper
  • NOCSCVPR 2019 Oral · arXiv:1901.02970 · paper
  • EI-30kdataset with LDA-1B · pin

LATENT — autonomous humanoid tennis research

Deployed body: 29-DoF Unitree G1 with the right hand replaced by a racket on a 3D-printed wrist connector. Not Galbot G1 or Galbot ET1.

Official Galbot-hosted LATENT highlight. The system performs dynamic footwork, forehand/backhand returns, and multi-shot rallies.
Unitree G1 humanoid performing several tennis return poses on an indoor court
Whole-body tennis return sequence. Image supplied by Galbot to Beijing Daily; reproduced on the Beijing government portal.
Research evidence: LATENT, arXiv:2603.12686 [R1] and the official project page [R2]. Real-world success uses the paper's court-boundary criterion. Pose and ball state use external optical mocap; the task is ball return, not a rules-level match. Eval body is 29-DoF Unitree G1, not Galbot G1.
MetricPublished resultEvidence / condition
System#LATENT — Learning Athletic Humanoid TEnnis skills from imperfect human motioN daTa [R1]Paper §1
Robot body#29-DoF Unitree G1; right hand replaced by a racket on a 3D-printed wrist connector [R1]Paper §4.3 and §5
Motion data#5 amateur players; 5 h of unedited, unannotated primitive-skill motion; 3 × 5 m capture area [R1]Paper §3.1
Control / simulation#High-level planner and low-level controller at 50 Hz; training simulation at 2,000 Hz [R1]Paper §3
Training episode#8 incoming balls; one launch every 2 s [R1]Training task, paper §3.3.1
Real-world success#Forehand 90.90%; backhand 77.78%; forecourt 88.89%; backcourt 81.82% [R1]Paper Table 5; successful return lands inside the opponent's court
Real-world sensing#External optical motion capture estimates robot 6D pose and ball state using reflective markers [R1]Paper §4.3; not onboard-vision-only autonomy

Official project videos

Multi-shot rallyContinuous returns with a human player.
Reactive footworkLateral positioning before the return.
Different human playerOfficial real-world project clip.

Berkeley Humanoid Lite — open research hardware

Paper system-design section. Total DoF stays (no vendor-style total in the extract). Cost is “under $5,000 (U.S. market)”, not a later aggregator $15k. Not mixed into the Unitree family table.

Sources: Chi et al., arXiv:2504.17249 [R52]; project site [R53]; HybridRobotics/Berkeley-Humanoid-Lite [R54].
MetricBerkeley Humanoid Lite
Height#800 mm [R52]
Mass#16 kg [R52]
Total DoF#
Stated hardware cost#under $5,000 (U.S. market prices) [R52]
Compute#Intel N95 mini PC [R52]
Battery#6S 4000 mAh LiPo [R52]
Stated operation time#~30 min [R52]
Actuator bus#CAN 2.0 at 1 Mbps; 250 Hz to actuators and IMU [R52]
Code / asset licenses#MIT code; CC BY-SA 4.0 assets (repository) [R54]
Sim assets#URDF / MJCF / USD in the repository; no browser assembly on this site [R54]

The paper demonstrates zero-shot RL locomotion transfer and teleoperated manipulation. Those are experiment results, not a product payload. A 3D card is omitted until a commit is pinned and a render is produced here.

Research asset

ToddlerBot — open research hardware

CoRL 2025 system-design text. Height from 0.56 m. Active DoF excludes end-effectors. Payload and walking-endurance figures are experiment results, not a product rating. Project-page 2.0 extras unused. Not a Unitree family column.

Sources: Shi et al., arXiv:2502.00893 [R55]; project site [R56]; hshi74/toddlerbot [R57].
MetricToddlerBot
Height#560 mm [R55]
Mass#3.4 kg [R55]
Active DoF (excluding end-effectors)#30 — 7 per arm, 6 per leg, 2 neck, 2 waist [R55]
Stated hardware cost#under $6,000 [R55]
Compute#Jetson Orin NX 16GB [R55]
Experiment — lift#1484 g in the paper’s payload test [R55]
Experiment — walking endurance#19 min stepping in place on a walking RL policy [R55]
Code repository#hshi74/toddlerbot [R57]

“40% of body weight” is an experiment comment, not a second mass. No 3D card until a commit is pinned. Stanford-TML/toddlerbot returned HTTP 404 here.

Claim boundary

Claim boundary. CCTV News (2026-03-17) and Beijing Daily (2026-04-09) call LATENT the “world's first fully autonomous tennis humanoid robot.” [P1] [P2] That superlative is attributed press. GroceryVLA is a homepage product model with no arXiv.