Complete postmortem library · Updated 2026-08-26

Every run.
Its actual story.

What each model built, how its run unfolded, why the submitted artifact earned its result, and what that result cost. Every page separates observed evidence from interpretation and unresolved questions.

30Published result rows
38Valid seeds represented
3Official 3-seed estimates
#01

Grok 4.6 (xhigh)

A fast hybrid build, the strongest observed single-run score, and 18 million cached-context tokens.

26%Score
2.65Points / $
#02

Claude Opus 5 (xhigh)

An ambitious neural-symbolic system, perfect public calibration, and a valid hidden-exam result.

23%Score
6.45Points / $
#03

GPT-5.6 Sol (xhigh)

Two nearly identical 4-layer hybrids scored 23% and 22%, making Sol the strongest repeatable provisional result.

22.5%Score
2.64Points / $
#04

GPT-5.5 (xhigh)

The benchmark's official GPT-5.5 estimate: 18%, 20%, and 22% across three independently built hybrids.

20%Score
4.44Points / $
#05

DeepSeek V4 Flash 0731 (xhigh)

A full-hour deterministic hardening campaign reached 19% for $1.58, making DeepSeek one of the strongest value results.

19%Score
12.04Points / $
#05

Qwen3.8-27B Aggressive FastMTP Q3_K_P (xhigh)

Qwen3.8-27B Aggressive FastMTP Q3_K_P scored a provisional 19% hidden exact match with local inference and no API charge.

19%Score
Points / $
#07

GLM-5.3 (max)

A max-effort hybrid recovered from a split-source packaging trap and produced a provisional 18% result for $3.84.

18%Score
4.69Points / $
#08

GPT-5.6 Luna (xhigh)

One small hybrid, a 17% hidden score, and the leaderboard's strongest current score-per-dollar result.

17%Score
62.52Points / $
#08

Gemini 3.7 Flash (high)

Perfect public calibration in six minutes, a $0.90 Vertex run, and a 17% hidden score.

17%Score
18.9Points / $
#10

Grok 4.5 (high)

Three exhaustive hybrid builds scored 10%, 17%, and 21% after 239 candidate submissions and heavy context use.

16%Score
1.05Points / $
#10

Claude Opus 4.8 (xhigh)

Two trained checkpoints, a broad inference layer, and an early 16% finish—but at nearly eight dollars.

16%Score
2.01Points / $
#12

Kimi K3 (max)

An official 15.67% mean from three low-cost hybrid runs scoring 14%, 16%, and 17%.

15.67%Score
14.78Points / $
#13

Gemini 3.6 Flash (xhigh)

A six-minute rule-and-ngram hybrid reached perfect public calibration, then stopped with 14% hidden accuracy.

14%Score
11.9Points / $
#13

Qwen3.8-27B (default)

Qwen3.8-27B reached a provisional 14% hidden score with local inference and no API charge.

14%Score
Points / $
#15

GLM-5.2 (xhigh)

A tiny timeout-safe checkpoint and a much larger second build scored 14% and 13%, suggesting a method ceiling rather than a size ceiling.

13.5%Score
4.1Points / $
#16

GPT-5.5 (high)

A deep-model search that favored an eight-layer fallback, reached perfect calibration, and scored 11% hidden.

11%Score
4.36Points / $
#16

Gemini 3.5 Flash (xhigh)

A rapid architecture sweep reached 92% public calibration and 11% hidden accuracy before ending fifty minutes early.

11%Score
5.25Points / $
#18

GPT-5.5 (medium)

Five model scales, eighteen candidate submissions, and a 10% hidden score after a 28-minute run.

10%Score
4.39Points / $
#19

Claude Fable 5 (xhigh)

Five candidates, a 13.62 MB conventional model, and a costly public-to-hidden generalization gap.

7%Score
0.86Points / $
#20

GPT-5.5 (low)

A 19-minute sprint to a perfect public checkpoint that transferred to only 4% of the hidden exam.

4%Score
4.39Points / $
#20

DeepSeek V4 Pro 0813 (max)

DeepSeek Pro built a larger hybrid, trusted its local test over the submitted artifact, and scored 4% for $3.38.

4%Score
1.18Points / $
#20

GLM-5.3 Flash (max)

A low-cost max-effort run survived candidate-selection and package-path risks but generalized to just 4% hidden accuracy.

4%Score
75.33Points / $
#23

Kimi K2.7 Code (xhigh)

A deterministic 52-minute build reached 52% public calibration and 3% hidden accuracy after unusually heavy iteration.

3%Score
0.8Points / $
#24

Claude Sonnet 5 (xhigh)

A mostly neural 50-minute run with 4% public calibration and only one hidden exact match.

1%Score
0.5Points / $
#24

MiniMax M3 (xhigh)

A cheap, mostly neural artifact looked partially functional in public and then scored 1% on the hidden exam.

1%Score
2.63Points / $
#24

Qwen 3.8 Max (xhigh)

Qwen's rerun remained near the neural baseline: two valid seeds, weak public calibration, and a 1% mean.

1%Score
0.4Points / $
#24

Ox Alpha · OpenCode 1.18.21 (max)

OpenCode 1.18.21 completed cleanly at max, reached 52% public calibration, and scored 1% hidden.

1%Score
Points / $
#28

North Mini Code (xhigh)

Ten training attempts, seven candidates, no public calibration signal, and a valid 0% hidden result.

0%Score
Points / $
UNR

Qwen 3.7 Plus (xhigh)

A 96% public checkpoint reached grading, but its output distribution did not sum to one, leaving no valid seed.

InvalidScore
Points / $
#28

Inkling (max)

A max-effort hybrid trained for nearly an hour, submitted once, and scored a valid 0%.

0%Score
0Points / $

Rows with fewer than three valid seeds remain provisional. Invalid attempts are documented but unranked. Each generated page uses aggregate run evidence, final artifact structure, and a reviewed public narrative; existing hand-authored analyses remain canonical where available.