Research

Each pin is one paper, product model, or benchmark. Media is optional: a video, an official project figure, or text only. Evidence grade R: a paper proves only the hardware, checkpoint, and evaluation it describes. Galbot G1 is a wheeled dual-arm; Unitree G1 is a biped; Galbot ET1 is an announced biped with no sheet here.

Recent papers

Four pins across on a wide desktop, then three, then two, then one. A pin may embed a video, an official figure, or only the paper. Learned systems stay off hardware tables. GroceryVLA is a product model with no public arXiv. The catalog lists every homepage paper.

preprint 2026 · arXiv:2607.06988

WAM-TTT

Steering World-Action Models by Watching Human Play at Test Time. Test-time training on a frozen world-action model; unlabeled human video only — no robot actions or annotations at deploy.

CVPR 2026 · arXiv:2606.03985

Humanoid-GPT

Zero-shot humanoid motion tracking with a GPT-style causal Transformer trained on a 2B-frame motion corpus.

Real-robot experiments use a 29-DoF Unitree G1 — not Galbot G1 and not Galbot ET1. Nearest open paper to the AstraBrain-WBC 0.5 announcement; not a proven one-to-one checkpoint identity.

IROS 2026 · arXiv:2603.12686

LATENT

Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data. Whole-body returns on a racket-mounted humanoid.

Deployed body is a 29-DoF Unitree G1 — not Galbot G1 and not Galbot ET1. Quantitative table and extra project clips below.

RSS 2026 · arXiv:2602.12215

LDA-1B

Scaling a 1.6B-parameter latent world-action model on mixed-quality human and robot data (EI-30k: sim+real, labeled+unlabeled).

Real-world figures use Galbot G1 (parallel-jaw gripper and Sharpa 22-DoF hand) and Unitree G1 variants — not Galbot ET1.

Robot using non-prehensile extrinsic dexterity to manipulate objects in clutter

RSS 2026 · arXiv:2603.09882 · Best Paper Finalist

DAPL

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning. Non-prehensile pushing and trapping in clutter.

HDFlow hierarchical diffusion-flow planning pipeline diagram

ICML 2026 spotlight · arXiv:2605.04525

HDFlow

Hierarchical Diffusion-Flow Planning for long-horizon tasks.

UrbanVLA route-conditioned micromobility robot in a city scene

ICRA 2026 · arXiv:2510.23576

UrbanVLA

Vision-language-action model for urban micromobility and delivery. Special env: outdoor streets, route-conditioned driving.

TrackVLA++ embodied visual tracking with memory and reasoning

ICRA 2026 · arXiv:2510.07134

TrackVLA++

Reasoning and memory for embodied visual tracking on top of TrackVLA.

SoFar language-grounded object orientation on a real robot

NeurIPS 2025 spotlight · arXiv:2502.13143

SoFar

Language-grounded orientation that bridges spatial reasoning and object manipulation.

FoldNet closed-loop garment folding on a dual-arm robot

RA-L & IROS 2026 · arXiv:2505.09109

FoldNet

Generalizable closed-loop garment folding. Special env: deformable cloth.

Dexonomy taxonomy of dexterous grasp types on diverse objects

RSS 2025 · arXiv:2504.18829

Dexonomy

Synthesizing all dexterous grasp types in a grasp taxonomy.

Uni-NaVid video-based VLA unifying embodied navigation tasks

RSS 2025 · arXiv:2412.06224

Uni-NaVid

Video VLA that unifies embodied navigation tasks. Successor of NaVid; feeds NavFoM.

MM-Nav multi-view visual navigation in cluttered indoor scenes

ECCV 2026 · arXiv:2510.03142

MM-Nav

Multi-view VLA for robust visual navigation via multi-expert learning. Special env: wires, dynamic humans, cluttered corridors.

LIMMT motion-tracking pipeline diagram

ICML 2026 · arXiv:2606.06953

LIMMT

Less is More for Motion Tracking. Companion tracking paper to Humanoid-GPT / HumanTracker.

CL3R 3D reconstruction and contrastive learning for manipulation representations

ICRA 2026 · arXiv:2507.08262

CL3R

3D reconstruction and contrastive learning for robotic manipulation representations.

CoRL 2025 Oral · arXiv:2502.17894

FetchBot

Zero-shot sim2real object fetching in clutter. No official still is linked here; paper and project only.

ICCV 2025 highlight · arXiv:2507.02747

DexVLG

Dexterous vision-language-grasp model at scale. Text pin; project page is the source.

IROS 2026 · arXiv:2603.13832

GraspADMM

Improving dexterous grasp synthesis via ADMM. Follow-on to Dexonomy / BODex.

ECCV 2026 · no arXiv on homepage

UniPart

Zero-shot language-grounded 3D part segmentation for embodied interaction. Not the CVPR 2026 UniPart generation paper by Xufan He.

DyWA non-prehensile manipulation of irregular objects on a table

ICCV 2025 · arXiv:2503.16806

DyWA

Dynamics-adaptive world-action model for generalizable non-prehensile manipulation. Lineage into LDA-1B.

D3RoMa material-agnostic depth sensing on transparent and specular objects

CoRL 2024 · arXiv:2409.14365

D3RoMa

Disparity-diffusion depth for material-agnostic manipulation. Special env: glass, metal, and other depth-hostile surfaces.

RoboHanger inserting hangers into diverse garments

RA-L · arXiv:2412.01083

RoboHanger

Generalizable hanger insertion for diverse garments. Special env: deformable clothing, not rigid pick-and-place.

ScissorBot paper-cutting skill pipeline

CoRL 2024 · arXiv:2409.13966

ScissorBot

Generalizable scissor skill for paper cutting. Special env: thin deformable sheets and a closing blade, not gripper pick.

Paper → body map

A paper proves only the robot it evaluates. A homepage “product model” is not a checkpoint identity. Galbot G1 and Unitree G1 stay in separate tables so form classes are not mixed. Launch architecture names that have no public checkpoint card sit outside tables. Cells use a fixed vocabulary rather than prose: arXiv is an id or no public arXiv; eval body is the exact robot the source names, or a statement that it names none; claim grade is one of paper eval (the paper evaluates on this body), homepage product (the company lists it on this body, without a paper eval), or not claimed (neither source puts the system on this body).

Not a table row. AstraBrain-Agent is the named launch architecture for Galbot’s first biped announcement. Public sources do not establish that LDA-1B, WAM-TTT, LATENT, or Humanoid-GPT released checkpoints run on that body. Status belongs on the announced page. [P2]

Evaluated or listed on Galbot G1

Sources: paper/project pins on this page plus hughw19.github.io [R21] and Qiming founder profile [P3].
SystemarXivWeights / codeEval body in the sourceEnd-effector in the sourceClaim grade
GraspVLAarXiv:2505.03233 [R17]HF shengliangd/GraspVLA [R17]real-robot evals; robot not named in the paper [R17]homepage product — homepage and founder profile call it the generalist grasp stack on Galbot G1 [R21] [P3]
GroceryVLAno public arXiv [R21]Galbot G1 [R21]homepage product — retail / pharmacy shelf-pick product model [P3]
TrackVLA / TrackVLA++arXiv:2505.23189; arXiv:2510.07134 [R19] [R25]embodied visual tracking; Galbot G1 not named in the papers [R19]homepage product — profile lists product-level tracking for unmanned retail [P3]
NavFoMarXiv:2509.12129 [R15]cross-embodiment: quadruped, drone, vehicle, wheeled platforms [R15]not claimed — no named Galbot G1 SKU [R21]
LDA-1BarXiv:2602.12215 [R9]Galbot G1 [R9]parallel-jaw gripper; Sharpa 22-DoF hand [R9]paper eval — real-world figures in the paper, not a store SKU [R9]
WAM-TTTarXiv:2607.06988 [R5]frozen world-action model (LDA line); robot not named in the source [R5]not claimed — no public proof of a named Galbot G1 product checkpoint [R5]
StereoVLAarXiv:2512.21970 [R11]HF shengliangd/StereoVLA [R11]manipulation VLA; robot not named in the paper [R11]not claimed — weights public; not listed as a Galbot G1 SKU [R11]

Evaluated on Unitree G1

Sources: Humanoid-GPT arXiv:2606.03985 [R7]; LATENT arXiv:2603.12686 [R1]; LDA-1B arXiv:2602.12215 [R9]. These rows are not Galbot G1.
SystemarXivWeights / codeEval body in the sourceEnd-effector in the sourceClaim grade
Humanoid-GPTarXiv:2606.03985 [R7]GalaxyGeneralRobotics code [R7]29-DoF Unitree G1 [R7]paper eval — nearest open paper to AstraBrain-WBC 0.5, not a proven one-to-one product checkpoint [P4]
LATENTarXiv:2603.12686 [R1]official code on Galbot GitHub [R1]29-DoF Unitree G1 [R1]paper eval — tennis whole-body control; the announced Galbot biped’s tennis mention is not this paper’s eval [R1]
LDA-1B (BrainCo variant)arXiv:2602.12215 [R9]Unitree G1 [R9]BrainCo 10-DoF hand [R9]paper eval — not a BrainCo Revo 2 product column [R9]

Benchmarks & special-env suites

Named datasets and evaluation suites from the same lab. A benchmark is not a robot SKU and does not enter a hardware table. Use these when a method paper’s “special env” needs a public test set: humanoid tracking, spatial VQA, open-instruction rearrangement, dex grasp, articulated parts, tracking-in-the-wild, urban micromobility.

HumanTracker humanoid motion-tracking benchmark teaser with diverse motions

ECCV 2026 · arXiv:2608.13555 benchmark

HumanTracker

August 2026 paper. ~153 h mocap, four motion families, and HumanScore (preference-aligned) so tracking eval matches contact, skating, and unstable support rather than only per-frame pose error. Special env: contact-rich, long-horizon humanoid motion.

OmniSpatial spatial-reasoning task examples for vision-language models

ICLR 2026 · arXiv:2506.03135 benchmark

OmniSpatial

Comprehensive spatial-reasoning benchmark for VLMs. Special env: language+vision spatial queries, not motor control.

Open6DOR open-instruction 6-DoF object rearrangement scenes

IROS 2024 Oral benchmark

Open6DOR

Open-instruction 6-DoF object rearrangement. Special env: language-specified placement pose, not pick-only grasping.

DexGraspNet large-scale synthetic dexterous grasps on many objects

ICRA 2023 · arXiv:2210.02697 dataset

DexGraspNet

Million-scale dexterous grasp dataset. Special env: multifinger grasp synthesis, not parallel-jaw SKUs.

DexGraspNet 2.0 generative dexterous grasps in cluttered scenes

CoRL 2024 · arXiv:2410.23004 dataset

DexGraspNet 2.0

Generative dexterous grasping in large-scale synthetic clutter. Special env: cluttered multifinger scenes.

GAPartNet actionable parts on articulated objects such as drawers and doors

CVPR 2023 Highlight · arXiv:2211.05272 dataset

GAPartNet

Part-centric articulated objects. Special env: drawers, doors, lids — not rigid free objects. GAPartManip (ICRA 2025) extends the manipulation set.

CoRL 2025 · with TrackVLA benchmark

EVT-Bench

Embodied Visual Tracking Benchmark named in TrackVLA (~1.7M samples). Special env: occlusion, night, distractors, long-horizon follow — not tabletop manipulation.

UrbanVLA outdoor micromobility evaluation scenes

ICRA 2026 · UrbanVLA eval figure eval suite

UrbanVLA outdoor eval

Official project figure for UrbanVLA’s outdoor / route-conditioned evaluation. Not a separately named public benchmark; it is the paper’s own special-env suite.

EI-30k mixed human and robot embodied data used to train LDA-1B

RSS 2026 · with LDA-1B dataset

EI-30k

30k+ hours of mixed-quality human and robot data (sim+real, labeled+unlabeled) used to scale LDA-1B. Special env: ingesting noisy human video next to robot trajectories, not a clean single-task demo set.

DREDS depth restoration on transparent and specular objects

ECCV 2022 · arXiv:2208.03792 dataset

DREDS

RGB-D restoration for transparent and specular surfaces. Special env: depth cameras fail on glass and metal; this set is the usual sim-to-real depth test for those materials.

ASGrasp reconstructing and grasping transparent objects from active stereo RGB-D

ICRA 2024 · arXiv:2405.05648 special env

ASGrasp

Transparent-object reconstruction and 6-DoF grasp from RGB-D active stereo. Pair with DREDS and STOPNet when the scene is glass, not matte household objects.

STOPNet multiview suction detection on transparent objects

ICRA 2024 · arXiv:2310.05717 special env

STOPNet

Multiview 6-DoF suction detection for transparent objects. Special env: suction on glass, not parallel-jaw grasp on opaque parts.

BODex bilevel optimization synthesizing dexterous grasps at scale

ICRA 2025 · arXiv:2412.16490 dataset / method

BODex

Scalable dexterous grasp synthesis with bilevel optimization. Use with DexGraspNet / Dexonomy when the special env is multifinger contact, not parallel-jaw SKUs.

UniDexGrasp++ generalist dexterous grasping across object categories

ICCV 2023 Oral · arXiv:2304.00464 dataset / method

UniDexGrasp++

Generalist dexterous grasping in large-scale synthetic clutter. ICCV 2023 Best Paper Finalist. Lineage into GraspVLA / SynGrasp-1B.

Paper catalog

One listing of the papers on the lab homepage (2026–2023 plus selected earlier lineage). This is the catalog intended as a single GitHub Issue; Issues cannot be opened from this environment, so the list lives here. Pins above are a subset. Refresh from hughw19.github.io before treating this as complete. GroceryVLA is a product model, not a paper.

2026

  • WAM-TTTpreprint · arXiv:2607.06988 · pin · paper
  • HumanTrackerECCV 2026 · arXiv:2608.13555 · pin · paper
  • UniPartECCV 2026 · no arXiv on homepage · pin
  • MM-NavECCV 2026 · arXiv:2510.03142 · pin · paper
  • GraspADMMIROS 2026 · arXiv:2603.13832 · pin · paper
  • LATENTIROS 2026 · arXiv:2603.12686 · pin · paper
  • HDFlowICML 2026 spotlight · arXiv:2605.04525 · pin · paper
  • LIMMTICML 2026 · arXiv:2606.06953 · pin · paper
  • LDA-1BRSS 2026 · arXiv:2602.12215 · pin · paper
  • DAPLRSS 2026 Best Paper Finalist · arXiv:2603.09882 · pin · paper
  • StereoVLARSS 2026 · arXiv:2512.21970 · pin · paper
  • Humanoid-GPTCVPR 2026 · arXiv:2606.03985 · pin · paper
  • Layered 4D-Rotor Gaussian SplattingCVPR 2026 · no arXiv on homepage · project
  • CL3RICRA 2026 · arXiv:2507.08262 · pin · paper
  • NavGSimICRA 2026 · arXiv:2603.15186 · paper
  • Robust Differentiable Collision DetectionICRA 2026 · arXiv:2511.06267 · paper
  • UrbanVLAICRA 2026 · arXiv:2510.23576 · pin · paper
  • TrackVLA++ICRA 2026 · arXiv:2510.07134 · pin · paper
  • OmniSpatialICLR 2026 · arXiv:2506.03135 · pin · paper
  • NavFoMICLR 2026 · arXiv:2509.12129 · pin · paper
  • FoldNetRA-L & IROS 2026 · arXiv:2505.09109 · pin · paper
  • DexNDMICLR 2026 · arXiv:2510.08556 · pin · paper

2025

  • SoFarNeurIPS 2025 spotlight · arXiv:2502.13143 · pin · paper
  • TrackVLA / EVT-BenchCoRL 2025 · arXiv:2505.23189 · pin · paper
  • FetchBotCoRL 2025 Oral · arXiv:2502.17894 · pin · paper
  • GraspVLACoRL 2025 · arXiv:2505.03233 · pin · paper
  • DexVLGICCV 2025 highlight · arXiv:2507.02747 · pin · paper
  • DyWAICCV 2025 · arXiv:2503.16806 · pin · paper
  • RoboHangerRA-L · arXiv:2412.01083 · pin · paper
  • DexonomyRSS 2025 · arXiv:2504.18829 · pin · paper
  • Uni-NaVidRSS 2025 · arXiv:2412.06224 · pin · paper
  • Code-as-MonitorCVPR 2025 · arXiv:2412.04455 · paper
  • GAPartManipICRA 2025 · arXiv:2411.18276 · GAPartNet pin · paper
  • BODexICRA 2025 · arXiv:2412.16490 · pin · paper
  • NaVid-4DICRA 2025 · pin · project
  • QuadWBGICRA 2025 · arXiv:2411.06782 · paper
  • Watch Less, Feel MoreICRA 2025 · arXiv:2502.14457 · paper
  • GroceryVLAproduct model · no arXiv · pin

2024

  • D3RoMaCoRL 2024 · arXiv:2409.14365 · pin · paper
  • DexGraspNet 2.0CoRL 2024 · arXiv:2410.23004 · pin · paper
  • ScissorBotCoRL 2024 · arXiv:2409.13966 · pin · paper
  • Task-Oriented Dexterous GraspIROS 2024 · arXiv:2309.13586 · paper
  • Open6DORIROS 2024 Oral · pin · project
  • NaVidRSS 2024 · arXiv:2402.15852 · paper
  • SAGERSS 2024 · arXiv:2312.01307 · paper
  • MaskClusteringCVPR 2024 · arXiv:2401.07745 · paper
  • STOPNetICRA 2024 · arXiv:2310.05717 · pin · paper
  • GAMMAICRA 2024 · arXiv:2309.15459 · paper
  • ASGraspICRA 2024 · arXiv:2405.05648 · pin · paper

2023 and selected earlier lineage

  • UniDexGrasp++ICCV 2023 Oral & Best Paper Finalist · arXiv:2304.00464 · pin · paper
  • GAPartNetCVPR 2023 Highlight · arXiv:2211.05272 · pin · paper
  • 3D-Aware Object Goal NavigationCVPR 2023 · arXiv:2212.00338 · paper
  • UniDexGraspCVPR 2023 · arXiv:2303.00938 · paper
  • PartManipCVPR 2023 · arXiv:2303.16958 · paper
  • Discrete Normalizing Flows on SO(3)CVPR 2023 · arXiv:2304.03937 · paper
  • DiGACVPR 2023 · arXiv:2304.02222 · paper
  • DexGraspNetICRA 2023 · arXiv:2210.02697 · pin · paper
  • GraspNeRFICRA 2023 · arXiv:2210.06575 · paper
  • DREDSECCV 2022 · arXiv:2208.03792 · pin · paper
  • HOI4DCVPR 2022 · arXiv:2203.01577 · paper
  • CAPTRAICCV 2021 Oral · arXiv:2104.03437 · paper
  • NOCSCVPR 2019 Oral · arXiv:1901.02970 · paper
  • EI-30kdataset with LDA-1B · pin

LATENT — autonomous humanoid tennis research

The deployed body is a 29-DoF Unitree G1 with the right hand replaced by a racket on a 3D-printed wrist connector. This is research on the Unitree G1 body; it is not attached to Galbot G1 or Galbot ET1.

Official Galbot-hosted LATENT highlight. The system performs dynamic footwork, forehand/backhand returns, and multi-shot rallies.
Unitree G1 humanoid performing several tennis return poses on an indoor court
Whole-body tennis return sequence. Image supplied by Galbot to Beijing Daily; reproduced on the Beijing government portal.
Research evidence: LATENT, arXiv:2603.12686 [R1] and the official project page [R2]. Real-world success uses the paper's court-boundary criterion.
MetricPublished resultEvidence / condition
SystemLATENT — Learning Athletic Humanoid TEnnis skills from imperfect human motioN daTa [R1]Paper §1
Robot body29-DoF Unitree G1; right hand replaced by a racket on a 3D-printed wrist connector [R1]Paper §4.3 and §5
Motion data5 amateur players; 5 h of unedited, unannotated primitive-skill motion; 3 × 5 m capture area [R1]Paper §3.1
Control / simulationHigh-level planner and low-level controller at 50 Hz; training simulation at 2,000 Hz [R1]Paper §3
Training episode8 incoming balls; one launch every 2 s [R1]Training task, paper §3.3.1
Real-world successForehand 90.90%; backhand 77.78%; forecourt 88.89%; backcourt 81.82% [R1]Paper Table 5; successful return lands inside the opponent's court
Real-world sensingExternal optical motion capture estimates robot 6D pose and ball state using reflective markers [R1]Paper §4.3; not onboard-vision-only autonomy

Official project videos

Multi-shot rallyContinuous returns with a human player.
Reactive footworkLateral positioning before the return.
Different human playerOfficial real-world project clip.

Research asset

Claim boundary

Claim boundary. CCTV News on 2026-03-17 and Beijing Daily on 2026-04-09 describe LATENT as the “world's first fully autonomous tennis humanoid robot.” [P1] [P2] That superlative remains an attributed media/company claim. The paper demonstrates learned, policy-controlled multi-shot returns without teleoperation, but its real-world setup uses external optical motion capture for the robot's global pose and tennis-ball state. The authors identify active vision as future work and state that the current task is ball return, not a complete rules-level two-player match. GroceryVLA is listed from the lab homepage as a product model; it has no public arXiv.