Benchmark health

Benchmark health

Before trusting a leaderboard, check whether the benchmark can still tell policies apart. Each card shows the best published value and the headroom left, whether the leaders share a tier, how far scores fall under perturbation, and whether the simulator ranks policies the way real robots do. A saturated or unvalidated benchmark says so first, and a sim-against-real comparison over fewer than five policies is marked too few to judge.

Every number comes from the cited rows on Result intervals. Boards are never compared across benchmarks, and a missing trial count stays missing.

Health cards

    Loading benchmark health…

    Track record

    The cards above use the result ledger. This section adds dated scores from each suite’s literature, published perturbation studies, and published comparisons with real robots.

    More

    The HP bar is the headroom left above the best cited score; a dimmed tile is at or above the saturation line.

    Source
    catalogue/benchmark-history-ledger.json holds every cited score with its source URL, table or section, and a short quote. scripts/benchmark_history.py writes catalogue/benchmark-history.json, versioned by benchmark-history.schema.json; the release build fails when it is stale.

    Badge rules

    scripts/benchmark_health.py writes catalogue/benchmark-health.json from catalogue/result-intervals.json, versioned by benchmark-health.schema.json; the release build fails when it is stale.

    More

    MMRV is the mean maximum rank violation from the SIMPLER paper (Li et al., 2024): for each policy, the largest real-robot gap to any policy that simulation orders the other way, averaged. The full rules are in Methodology.