FrontierSWEV2

Benchmarking software engineering skill at the edge of human ability.

By

Leaderboard

mean@5worst→best@5

Scores across all 34 tasks. Each model runs 5 trials per task with a 20-hour budget.

#ModelScoreAvg CostTime

1

Claude Fable 5.1

proximus

56.3

%

±

11.1

$138.5511.6h

2

GPT-5.6

proximus

32.2

%

±

11.0

$179.648.6h

3

GLM-5.3

proximus

30.2

%

±

11.5

$97.2217.0h

4

Kimi K3

proximus

25.9

%

±

11.8

$109.7118.5h

5

Grok 4.6

proximus

25.3

%

±

12.2

$243.4313.8h

6

Gemini 3.7 Flash

proximus

20.3

%

±

10.5

$34.147.9h

7

Qwen3.8-Max

proximus

15.8

%

±

7.8

$55.1418.5h

8

DeepSeek V4 Flash Exp

proximus

14.8

%

±

9.7

$8.5714.9h

9

Muse Spark 1.2

proximus

12.0

%

±

5.8

$27.814.6h

10

Inkling

proximus

4.1

%

±

5.3

$9.1570m

0%20%40%60%80%100%

Bar and value are mean@5, whiskers span worst@5–best@5. Green marks the top result. Cost and time are per-trial averages.

Analysis

Score for every model on every task; stronger color is higher. The "Best"column is the highest score any model achieved on that task. Click a cell for that pair's traces; click a Best cell to open the top-scoring run's trace.

0%100%

Best

Final score (mean@5)585732302625201615124