Each board holds one benchmark split, condition, metric and protocol.
More
Results on different boards are never compared. A result whose source does not state its trial count keeps its published value but gets no interval and no letter.
Evidence mix
Who ran each evaluation, where, and whether the trial count is published.
Boards
Each row is one cited result. The diamond is the published value and the bar is its 95% interval; a thicker bar means more trials.
More
Letters on the left are tiers: rows that share a letter are not separated at the board's corrected threshold.
Method
- Source
catalogue/result-ledger.jsonholds each cited result with its source URL, table or section, and the trial count as the source states it. Results this site executed itself are read fromworlds/maniskill.jsonandevidence/native-evaluation/.scripts/result_intervals.pywrites catalogue/result-intervals.json, versioned by result-intervals.schema.json; the release build fails when it is stale.- Interval
- Successes are recovered as
k = round(value × n). The interval is the 95% equal-tailed range of a Beta(1 + k, 1 + n − k) posterior, the uniform-prior choice used in recent large policy evaluations. When every trial failed the interval starts at 0%, and when every trial succeeded it ends at 100%; the other end is then the one-sided 95% bound, so the interval always contains the published value. A result marked “k derived” is one where no whole count prints as the published rate, usually an average over seeds or tasks, so the interval is approximate. Details are in Methodology. - Letters
- On each board, every pair of counted results is compared by the posterior probability that one success rate is higher. The probability is computed exactly, not on a grid. A pair is separated only when that probability clears a two-sided threshold of 0.05 divided by the number of pairs (Bonferroni). Letters follow the compact letter display: two results that share a letter are not separated. Past Z the letters continue AA, AB.
- Evaluator
- Self: the authors evaluated their own policy. Third-party: another group re-ran it. Challenge: an organiser evaluated submissions on a held-out set. Site: this site executed it.