Result intervals

Result intervals

A success rate means little without the number of trials behind it. Every cited result here carries its trial count, a 95% interval, and a letter. Results on one board that share a letter cannot be told apart, however their headline numbers differ.

Each board holds one benchmark split, condition, metric and protocol.

More

Results on different boards are never compared. A result whose source does not state its trial count keeps its published value but gets no interval and no letter.

Evidence mix

Who ran each evaluation, where, and whether the trial count is published.

Boards

Each row is one cited result. The diamond is the published value and the bar is its 95% interval; a thicker bar means more trials.

More

Letters on the left are tiers: rows that share a letter are not separated at the board's corrected threshold.

Method

Source
catalogue/result-ledger.json holds each cited result with its source URL, table or section, and the trial count as the source states it. Results this site executed itself are read from worlds/maniskill.json and evidence/native-evaluation/. scripts/result_intervals.py writes catalogue/result-intervals.json, versioned by result-intervals.schema.json; the release build fails when it is stale.
Interval
Successes are recovered as k = round(value × n). The interval is the 95% equal-tailed range of a Beta(1 + k, 1 + n − k) posterior, the uniform-prior choice used in recent large policy evaluations. When every trial failed the interval starts at 0%, and when every trial succeeded it ends at 100%; the other end is then the one-sided 95% bound, so the interval always contains the published value. A result marked “k derived” is one where no whole count prints as the published rate, usually an average over seeds or tasks, so the interval is approximate. Details are in Methodology.
Letters
On each board, every pair of counted results is compared by the posterior probability that one success rate is higher. The probability is computed exactly, not on a grid. A pair is separated only when that probability clears a two-sided threshold of 0.05 divided by the number of pairs (Bonferroni). Letters follow the compact letter display: two results that share a letter are not separated. Past Z the letters continue AA, AB.
Evaluator
Self: the authors evaluated their own policy. Third-party: another group re-ran it. Challenge: an organiser evaluated submissions on a held-out set. Site: this site executed it.