Policy DNA

Policy DNA

How each recorded controller moves, not only whether it succeeded. Every fingerprint is computed from the per-step actions and rewards in the site's own recorded episodes, so a reinforcement-learning policy, a VLA checkpoint, a human teleoperator, a scripted expert, a planner and random actions can be read side by side.

One fingerprint describes one recorded episode.

More

It is not a suite score, a policy ranking, or evidence that a controller generalizes. Action channels mean different things in different action spaces, so roughness and channel use are only compared inside one action family.

Class roster

One card per controller class. Medians are taken over that class's episodes that keep an action channel.

Smoothness field

One panel per action family. Left is smoother, right is rougher (log scale).

More

Higher means more of the action channels actually move. Each marker's shape and colour name its controller class; focus or hover a marker for its values.

DNA strips

Each strip is one episode. Rows are action channels and columns are time, so a channel's colour shows where it sits in its own range (pale is the low end, deep is the high end).

More

Grey rows never move. The bar underneath fills with cumulative reward where the recording declares one, and flags mark recorded events.

Method

Source
scripts/policy_dna.py reads every manifest and chunk under /worlds/episodes/ and writes /worlds/policy-dna.json, versioned by policy-dna.schema.json. The release build fails if the projection is stale.
Roughness
The mean, over interior control steps, of the root-mean-square across active channels of (a[t+1] − 2a[t] + a[t−1]) divided by that channel's range in the episode. It is unitless and per control step, so a faster controller takes smaller steps and reads smoother for the same motion.
Channel use
The share of action channels whose range in the episode is above 10−9 and above 10−6 of the channel's largest magnitude. A teleoperator who never touches the base, or a bimanual expert who holds one arm still, shows up here.
Reward timing (t50)
The fraction of reward steps after which cumulative reward first reaches half of a positive return. Only episodes whose manifest declares a reward overlay are timed.
Controller class and action family
Authored in the script from each manifest's label, notes and protocol. A new recording fails the build until it is classified.