Leaderboard
mean@5worst→best@5
Scores across all 34 tasks. Each model runs 5 trials per task with a 20-hour budget.
#ModelScoreAvg CostTime
0%20%40%60%80%100%
Bar and value are mean@5, whiskers span worst@5–best@5. Green marks the top result. Cost and time are per-trial averages.
Analysis
Score for every model on every task; stronger color is higher. The "Best"column is the highest score any model achieved on that task. Click a cell for that pair's traces; click a Best cell to open the top-scoring run's trace.
0%100%
Best
Final score (mean@5)585732302625201615124