BEHAVIOR

BEHAVIOR source replay

Watch R1Pro turn on a radio in an official simulated human demonstration. Switch between its head and wrist cameras while recorded state and commands follow the displayed frame.

Open the recording · Reproduce the conversion · Source registry

What is available

Official dataset and selected source hashes and all-frame comparison.
CapabilityPublished evidence
Video and state replay1,956 LeRobot rows; head 720 × 720, both wrists 480 × 480; 30 fps; three 65.2-second H.264 clips.
Recorded channelsAll 61 observation-state values, 23 action values, three seven-value robot-relative camera poses and next-transition reward/termination/truncation fields are retained at their source row.
Task resultIndependent success: not established. The source next_terminated flag first becomes true at row 1,364 and remains true through row 1,955. A termination flag is not a locally verified task score.
Native execution and 3D worldNot available. No scene geometry or encrypted Data Bundle assets are included. This recording contributes zero complete 3D worlds and zero independently verified tasks.

Source identity and time

The selected demonstration is task turning_on_radio, LeRobot episode 0, raw episode id 10 and task instance 1. These are dataset identifiers, not a simulator seed. The dataset revision is 4f50b44796641a4d526a19d9aeadc8aa51e2f2c2.

Every source Float32 timestamp maps to frame i / 30 within four microseconds. The last row is at 65.166664 seconds; the final camera frame occupies the remainder of the 65.2-second clip. Playback follows the video’s presented frame. Metadata downloads retain the original row time and encoded-media time separately.

The reference writer pairs the observation before action i with that action and the next-transition fields after it. No one-frame shift is applied. The plot compares left arm J1 state[3] with action[7], and right arm J1 state[28] with action[15], in radians. Base velocity, gripper commands and finger positions have different meanings and are retained in exports without being mixed into this angle plot.

The source skill annotation covers 1,776 frames, while the recording contains 1,956 rows. Its move, pickup, press and place spans have no established mapping to the complete video timeline. They remain unsynchronized source material; no offset or stretched timeline is invented.

The v3.9.2 R1Pro reference layout and reference writer explain these channels. That August 24 code release postdates the August 5 dataset revision. The exact historical simulator build is not established.

Conversion and access

The public demonstration repository supplies an MIT licence. The selected ten-file source closure totals 706,407,125 bytes. The three derived H.264 clips total 26,718,034 bytes; no full dataset download is required.

Every published frame was compared with its original decoded source frame: 5,868 comparisons across three cameras, with matching frame counts, dimensions and presentation timestamps. The lowest frame PSNR was 43.69 dB. This checks the lossy video conversion; it does not establish simulator or task success.

Native OmniGibson requires a supported NVIDIA RTX GPU, which was unavailable on the Intel Arc desktop used here. The separate encrypted BEHAVIOR 3D Data Bundle has its own noncommercial, OmniGibson-only terms and restrictions on extraction and redistribution. No bundle, key, geometry or raw HDF5 file is included in this publication. Reference requirements and terms.

Downloads

Recording registry · Manifest and state chunk links · Source and conversion evidence