Head to head
Same brief, same rubric, same scale. Every model on the task side by side, or just the ones you pick: dimension scores, how many of their claims held up, and which issues each review raised.
It measured what everyone argued about.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model with no network tools, no skills, no plugins, no MCP servers and no subagents.
GPT-6 AstraReproduced, not asserted.
GPT-5.6 SolFixes the instrument first.
Claude Fable 5.1Counts the calls, misprices them.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model with no network tools, no skills, no plugins, no MCP servers and no subagents.
HY4 PreviewThe first one to measure.
GPT-5.6 LunaCareful, short, and never built.
UNIONALPHAAccurate to the byte. Hesitant about what stops decoding.
Grok 4.6Finds what erases the signal. Misses what stalls the voice.
Muse Spark 1.3 ContributorFinds both halves of the filter problem. Points the CPU fix the wrong way.
DeepSeek V4.1 FlashFinds the microphone and the filters. Waves de-emphasis through.
GLM 5.3Proves the real capture bug. Invents a second one.
Gemini 3.8 FlashA buildable plan, on an invented number.
Qwen3.8 MaxBest case yet for the microphone. Never looks at the CPU.
Qwen3.8 FlashRight about the filters. Wrong about the CPU.
GLM 5.3 FlashAudits the memory to the byte. Misreads the receive path.
DeepSeek V4 ProRight pivot. Wrong arithmetic.
These models ran under the same conditions, except where a card says otherwise. How the runs were set up →
The first to find all four, and the only one that measured them
Sixteen models have now reviewed this repository. The fourteen before this one all missed the vocoder CPU wall; this one profiled it in an emulator and tied it to the parser reset that ends the call. It also found both halves of the analog chain, the sample source and the frame-clock rule. GPT-6 Astra remains the best of the ranked audits and the one that found what this review skipped: the instruments.
From the Claude Opus 5 audit, 17 September 2026.
Weighted score
Out of 100, on the grade scale.
Earned and lost, by rubric slot
Each bar is a model’s 100-point grade split into the rubric’s weighted slots; filled is earned, empty is lost. The slots line up, so a gap in one bar against a full slot in another is exactly where the models part ways.
- 1Accuracy & evidence30 pts
- 2Coverage of decode problems25 pts
- 3Root cause & prioritisation15 pts
- 4Fix plan & acceptance gates15 pts
- 5Originality & attribution10 pts
- 6Clarity & calibration5 pts
Scores on each rubric dimension
One panel per model: its score on each rubric dimension and the weighted total over the grade zones, in its own colour, with every other model on this task as a grey dot for scale. Hover a row for every model’s score and the auditor’s reasoning.
- The panel’s model, joined down the dimensions
- Every other model on this task
Claude Opus 592 A
GPT-6 Astra89 B
GPT-5.6 Sol84 B
Claude Fable 5.179 C+
HY4 Preview77 C+
GPT-5.6 Luna75 C+
UNIONALPHA74 C
Grok 4.670 C
Muse Spark 1.3 Contributor68 C−
DeepSeek V4.1 Flash66 C−
GLM 5.365 C−
Gemini 3.8 Flash64 D
Qwen3.8 Max63 D
Qwen3.8 Flash59 D
GLM 5.3 Flash56 D
DeepSeek V4 Pro54 D
Scores and weighted points
| Dimension | Weight | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | HY4 Preview | GPT-5.6 Luna | UNIONALPHA | Grok 4.6 | Muse Spark 1.3 Contributor | DeepSeek V4.1 Flash | GLM 5.3 | Gemini 3.8 Flash | Qwen3.8 Max | Qwen3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Pro |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy & evidence | 30% | 94 · 28.2 pts | 92 · 27.6 pts | 91 · 27.3 pts | 85 · 25.5 pts | 85 · 25.5 pts | 82 · 24.6 pts | 90 · 27.0 pts | 74 · 22.2 pts | 74 · 22.2 pts | 73 · 21.9 pts | 68 · 20.4 pts | 65 · 19.5 pts | 70 · 21.0 pts | 62 · 18.6 pts | 64 · 19.2 pts | 61 · 18.3 pts |
| Coverage of decode problems | 25% | 86 · 21.5 pts | 84 · 21.0 pts | 70 · 17.5 pts | 62 · 15.5 pts | 56 · 14.0 pts | 58 · 14.5 pts | 65 · 16.3 pts | 59 · 14.8 pts | 54 · 13.5 pts | 49 · 12.3 pts | 52 · 13.0 pts | 46 · 11.5 pts | 39 · 9.8 pts | 40 · 10.0 pts | 34 · 8.5 pts | 33 · 8.3 pts |
| Root cause & prioritisation | 15% | 96 · 14.4 pts | 88 · 13.2 pts | 84 · 12.6 pts | 84 · 12.6 pts | 80 · 12.0 pts | 80 · 12.0 pts | 60 · 9.0 pts | 70 · 10.5 pts | 72 · 10.8 pts | 66 · 9.9 pts | 74 · 11.1 pts | 74 · 11.1 pts | 71 · 10.7 pts | 68 · 10.2 pts | 62 · 9.3 pts | 64 · 9.6 pts |
| Fix plan & acceptance gates | 15% | 93 · 14.0 pts | 92 · 13.8 pts | 90 · 13.5 pts | 86 · 12.9 pts | 84 · 12.6 pts | 80 · 12.0 pts | 70 · 10.5 pts | 75 · 11.3 pts | 71 · 10.7 pts | 70 · 10.5 pts | 68 · 10.2 pts | 70 · 10.5 pts | 71 · 10.7 pts | 70 · 10.5 pts | 62 · 9.3 pts | 60 · 9.0 pts |
| Originality & attribution | 10% | 95 · 9.5 pts | 90 · 9.0 pts | 84 · 8.4 pts | 84 · 8.4 pts | 82 · 8.2 pts | 72 · 7.2 pts | 76 · 7.6 pts | 72 · 7.2 pts | 70 · 7.0 pts | 74 · 7.4 pts | 70 · 7.0 pts | 70 · 7.0 pts | 66 · 6.6 pts | 64 · 6.4 pts | 58 · 5.8 pts | 62 · 6.2 pts |
| Clarity & calibration | 5% | 92 · 4.6 pts | 86 · 4.3 pts | 86 · 4.3 pts | 88 · 4.4 pts | 88 · 4.4 pts | 84 · 4.2 pts | 80 · 4.0 pts | 78 · 3.9 pts | 80 · 4.0 pts | 74 · 3.7 pts | 72 · 3.6 pts | 84 · 4.2 pts | 78 · 3.9 pts | 70 · 3.5 pts | 68 · 3.4 pts | 48 · 2.4 pts |
| Weighted total | 100% | 92 · A | 89 · B | 84 · B | 79 · C+ | 77 · C+ | 75 · C+ | 74 · C | 70 · C | 68 · C− | 66 · C− | 65 · C− | 64 · D | 63 · D | 59 · D | 56 · D | 54 · D |
Every claim, checked
One cell per claim in each audit’s claim table, checked at the lines it cites. Hover a cell for the claim and the auditor’s note. The reviews made different numbers of claims, so the bar under each grid gives the shares.
Claude Opus 5’s claim table · GPT-6 Astra’s claim table · GPT-5.6 Sol’s claim table · Claude Fable 5.1’s claim table · HY4 Preview’s claim table · GPT-5.6 Luna’s claim table · UNIONALPHA’s claim table · Grok 4.6’s claim table · Muse Spark 1.3 Contributor’s claim table · DeepSeek V4.1 Flash’s claim table · GLM 5.3’s claim table · Gemini 3.8 Flash’s claim table · Qwen3.8 Max’s claim table · Qwen3.8 Flash’s claim table · GLM 5.3 Flash’s claim table · DeepSeek V4 Pro’s claim table
- Holds
- Qualifiedoverstated, miscounted or doubtful
- Wrong
20 hold · 1 qualified · 0 wrong · 95% hold
27 hold · 2 qualified · 0 wrong · 93% hold
27 hold · 1 qualified · 0 wrong · 96% hold
15 hold · 4 qualified · 1 wrong · 75% hold
24 hold · 6 qualified · 0 wrong · 80% hold
23 hold · 0 qualified · 1 wrong · 96% hold
33 hold · 2 qualified · 0 wrong · 94% hold
21 hold · 7 qualified · 2 wrong · 70% hold
22 hold · 6 qualified · 2 wrong · 73% hold
28 hold · 4 qualified · 3 wrong · 80% hold
25 hold · 4 qualified · 2 wrong · 81% hold
23 hold · 8 qualified · 4 wrong · 66% hold
21 hold · 4 qualified · 2 wrong · 78% hold
21 hold · 6 qualified · 6 wrong · 64% hold
17 hold · 7 qualified · 4 wrong · 61% hold
17 hold · 4 qualified · 5 wrong · 65% hold
What each review raised
Claude Opus 5 found 4 of the 4 decode-critical issues; GPT-6 Astra found 3 of the 4 decode-critical issues; GPT-5.6 Sol found 2 of the 4 decode-critical issues; Claude Fable 5.1 found 2 of the 4 decode-critical issues; HY4 Preview found 2 of the 4 decode-critical issues; GPT-5.6 Luna found 2 of the 4 decode-critical issues; UNIONALPHA found 1 of the 4 decode-critical issues; Grok 4.6 found 3 of the 4 decode-critical issues; Muse Spark 1.3 Contributor found 2 of the 4 decode-critical issues; DeepSeek V4.1 Flash found 2 of the 4 decode-critical issues; GLM 5.3 found 2 of the 4 decode-critical issues; Gemini 3.8 Flash found 2 of the 4 decode-critical issues; Qwen3.8 Max found 2 of the 4 decode-critical issues; Qwen3.8 Flash found 2 of the 4 decode-critical issues; GLM 5.3 Flash found 1 of the 4 decode-critical issues; DeepSeek V4 Pro found 1 of the 4 decode-critical issues. Every tracked issue below, with the reviews that were already in the repository for reference.
Issue signatures
Each review as a trace across the tracked issues. Where traces part is where the reviews differ; hover a column for every review’s entry. Column numbers match the table below.
- raised
- partly
- absent
- Decode-critical
| # | Issue | Project docsEarlier repo review · in the repository | Grok 4.6Model · Sep 17 | DeepSeek V4.1 FlashModel · Sep 17 | Muse Spark 1.3 ContributorModel · Sep 17 | GLM 5.3 FlashModel · Sep 17 | UNIONALPHAModel · Sep 17 | Qwen3.8 FlashModel · Sep 17 | Qwen3.8 MaxModel · Sep 17 | GLM 5.3Model · Sep 17 | DeepSeek V4 ProModel · Sep 17 | Gemini 3.8 FlashModel · Sep 17 | HY4 PreviewModel · Sep 17 | GPT-5.6 SolModel · Sep 17 | GPT-5.6 LunaModel · Sep 17 | GPT-6 AstraModel · Sep 17 | Claude Opus 5Model · Sep 17 | Claude Fable 5.1Model · Sep 17 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Capture export stages the wrong I2S halfwords | absent | raised§1, first code change | absentits first experiment uses that capture | absenttrusts the capture | absentrelies on the capture | raisedF3, fixed and verified | raisedP1-2, tests miss the adapter | absentcalls the capture validated | raised§2, proven with a reproducer | absentcalls the capture module sound | absentcertifies the capture path | absentits plan depends on it | raisedderived from the stride | absentits first step relies on it | raisedF4, reproduced on the host | absent | absent |
| 2 | Capture parser certifies an incomplete stream | absent | absent | absent | absenttrusts the parser | absentrelies on the parser | absent | absent | absentrelies on the parser | absent | absent | absent | absent | partlymodule tests miss the stride | absent | partlyasks for tests at the callback | absent | absent |
| 3 | pcm_starve never increments | absentdocumented as working | absent | absent | absentrelies on it | absent | raisedF7 | raisedP1-7 | absent | raised§2 | absent | absentnamed only as a predicted symptom | absent | absent | raisedF7, first to find it | raisedF9, reproduced | absent | absent |
| 4 | No static RAM margin | partlymargins still to measure | absent | raisedM6, 0 bytes free | partly“nearly full”, from the docs | raisedF5, byte-exact map audit | raisedF6, byte-exact | partlycalls the reservation headroom | absent | partly“~zero headroom” | partlyfrom the docs, not the map | raised§3.5, exact map symbols | absent | raisedexact map figures | absentcould not read the map | raisedF8, from its own link map | raisedwith the overlay buffers named | raisedwith the objects to reclaim named |
| 5 | The I2S stream is most likely microphone audio Decode-critical | partlyopen, leaning sceptical | raised§1, 0x89 versus 0xC9 | raisedB1, four-value test | raisedP1-1 | partlyopen, leans towards RF | partlyblocking, but never says microphone | raisedP1-4, 0xE0 mic bit | raisedA, the clearest case yet | raised§1, three sources | raised§1–§2, with the tap as the fix | raised§1 and §3.1 | partlylisted unverified, tested first | raisedBlocker 1, source unproven | raisedF1, as the safe default | raisedF2, as far as the evidence goes | raisedproved from 0xE0=0x89 and sound.c | raisedtraced to the buffer’s only consumer |
| 6 | HR-C6000 de-emphasis on the capture path Decode-critical | absent | raised§2, bit 5 of 0x34 | absentcalls it benign | raisedP1-2, 0x34=0x3C | raisedF2, closes the eye | partlycited, called unmeasured | raisedP1-3, 0x34=0x1C | raisedC, with the 0x34 bit 5 fix | raised§1, tilts the eye | absent | raised§3.3, 0x34=0x3C | raisedP1-1, and the boot table too | raised0x34=0x3C, never replaced | raisedF2, without the fix | raisedF2, 0x34=0x3C | raisedmeasured at 48.9% SER | raisedwith the 300–3400 Hz bandpass |
| 7 | AT1846S FM filters, low-frequency bit, 25 kHz Decode-critical | partly“require characterization” | raised§2, register level | raisedB2, filter register | raisedP1-2, 0x58 filters | partly“voice filtering”; wrong bandwidth premise | partlycited, called unmeasured | raisedP1-3, DMR 0x58 probe | absent | absent | absent | raised§3.3, with the DMR fix | absent | raised25 kHz for a 12.5 kHz channel | absent | raisedF2, the FM settings table | raisedall three, with the DMR fix | raisedall three, FM against DMR |
| 8 | 0x10=0x6E hybrid state; 0x36 dual role | partlybring-up clock rules | partlymisses 0x6E and the 0x36 clock gate | partly“undocumented hybrid state” | partlyquiet-chip registers | raisedF4, 0x10=0x80 kills the clock | partlyF1, 0x6E against 0x80 | partly“hybrid I2S state” | partlyslot engine and eco only | partly§4.3, 0x10=0x80 hazard | partly“inconsistent hybrid state” | partly0x36 and 0xE0, not 0x10 | partly0x36, 0xE0, 0x26; not 0x10 | partlyFM mode writes, not 0x10 | partly0x36, 0xE0, 0x26 listed | partly0x36 and 0x10 via the squelch path | partly0x36 and 0xE0, not 0x10 | raisedboth, as the per-frame writes |
| 9 | Manual: I2S frame clock “must be 8KHz” Decode-criticalmissed or ruled out | absent | raised§3 | absentquotes the paragraph, not the rule | absent | absent | raisedF4 | absentquotes the formulas, not the rule | absentcites the section, not the rule | absent | absent | absent | raisedP1-2, with a test for it | absent | absent | raisedF7, with the divisor arithmetic | raisedH2 | absent |
| 10 | One-layer 4FSK test mode as a P25 tap missed or ruled out | partlystock BER-test block only | raisedGate D | raisedB5, exact recipe | absentdismissed | absent | partlyworth a bounded test | absentruled out at “9600 Bd” | raisedstep 2, a symbol source | absent“not a P25 symbol source” | absent“no raw modem mode” | absent“zero internal silicon capability” | partlycited, then dismissed | absent | absent | partlypoints at the layer architecture | raisedRoute A, with an acceptance test | absentruled out on DMR framing |
| 11 | ±10% health gate versus ±1% timing clamp | absent | absent | raisedB3, impact overstated | raisedP1-3 | raisedF3 | absent | absent | absent | absent | absent | absent | absent | raisedwith the resampler consequence | raisedF4, with the resampling consequence | raisedF7 | absent | absent |
| 12 | Fail-closed muting at LDU cadence | absent | absentlate-entry mute only | absent | absentcalls it an asset | absent | absent | partly“keep it” | absent | absentcalls it tested | absent | absentcertifies it as correct | absent | absentcertifies it as correct | raisedF8, weighed as a trade | raisedF6, the late-entry half | partlyvia the MFID path | partlylate entry mutes seven frames |
| 13 | Non-standard MFID mutes clear calls | absent | absent | absent | absent | absent | absent | absent | absent | absent | absent | absentcertifies it as correct | absent | absentcertifies it as correct | absent | absent | raisedM1, with the Motorola case | absent |
| 14 | Test waveform shares the receiver’s RRC filter | partly“synthetic RRC/AWGN” caveat | absent | partlytested it, says not to fix | partlysynthetic only, wants recordings | absentwould extend that model | raisedF5, unquantified | absentwould extend that model | partlysynthetic only, not the circularity | partlysynthetic only | partly“ideal RRC-shaped signal” | absent | partlyP1-5, circularity without the filter | raisednames the shared table | partlysynthetic, not circular | raisedF10, with why the tests missed F1/F4 | raisedtested, called harmless | partlyasks for a recorded waveform |
| 15 | MCU runs at 72 MHz | raised | raisedin passing | absent | raisedP1-5 | absent | raisedF7 | absent | absent | raised§5 | raised§7, with the OpenRTX precedent | raised§3.5, with the PLL settings | raisedP1-4, with the PLL lines | partlynamed as the target, never costed | raisedF5, as the real-time risk | raisedF9, the target for measurement | raisedand measured against it | raisedwith the PLL arithmetic |
| 16 | Vocoder needs 11–16× the 72 MHz CPU Decode-criticalmissed or ruled out | partlydecode timing unmeasured | absent“fine on a 1 ms tick” | absent“vocoder question settled” | partlyunmeasured; fix order backwards | absent“in good shape” | partlydeadlines “unproven” | absent“not the problem” | absent“not the problem” | partlyinline, unmeasured | partlycites 4.36M, calls it 60% | absentputs it at 15–18 ms per frame | partly87% measured, budget unresolved | absentleft as a later measurement | partlynamed, never sized | partlycounted the filter, not the vocoder | raisedmeasured 6.5× on the demo build | partly11,812 trig calls, priced at 60–120% |
| 17 | Direct discriminator tap (M17 mod) | raised | raiseduncredited | raisedpins, timer ADC, 48 kS/s | raisedfallback, pin 9 | raisedfallback | raisedfallback | raisedfallback | raisedstep 5, the likely answer | raised§1 pivot | raisedits central recommendation | raised§5.1, with the ADC and DMA design | raisedthe fallback if Gate 0 fails | raisedADC with bias and anti-alias | raisedF1, the fallback route | raisedF2, ADC or another interface | raisedRoute B, with the ADC channels checked | raisedwith DC coupling and the ADC plan |
| 18 | Phase 2 architecture and scope | partlynot implemented | partlymisplaces the AMBE+2 decoder | partlyvoice via mbelib AMBE+2 | raisedwith RF band limits | partlysays mbelib has no AMBE+2 | raisedmost accurate section | partlyDMR and X2-TDMA parts | raisedAMBE+2 present but uncalled | raisedwith a reference map | partlyno symbol rate or sync | raisedaccurate on rate and slots | partlyright conclusion, H-CPM mislabelled | raisedplus the linked-symbol check | raisedcareful about the vocoder | raisedthe most complete of the fourteen | raisedincluding the descrambler seed | raisedincluding the NET_STS_BCST seed |
| 19 | Two unverified SPI writes per decoded 20 ms frame | absent | absent | absent | absent | absenttreats them as protection | absent | absent | absent | absent | absent | partlythe SPI0inUse mechanism | partlysilent SPI0 failures | raisedradioSetAudioPath per frame | absent | partlythe sink is cited, not the writes | absent | raised0x36 and 0x10=0x6E |
| 20 | Capture sessions lack epochs | absent | absent | partlymeasured=0 only | absent | absent | raisedF9, 8 kHz under a 24 kHz header | absent | absent | absent | absent | absent | absent | partlyasks for source metadata | absent | raisedF4, segment on retune and clock change | absent | absent |
| 21 | Ring and tick real-time budget | partlydeadlines unproven | partlycalls it fine | absent | partlyoverruns look like weak RF | absent | raisedF7, 1 ms is a minimum | partlyregister stalls against the ring | absent | partly1 ms tick as a constraint | absent | absent | absent | raisedthe 2 s refresh stall | partlybuffer inventory only | raisedF9, all four figures exact | raisedthe 21 ms ring against a 55–190 ms stall | partlyper-window MAC cost only |
| 22 | Clock config 3 assumes 12,288 Hz; the codec formula gives 12,000 | absent | absent | raisedB3, clock model | absent | partly“guessed semantics” | absent | raisedP1-6, for a different reason | absent | partly“unvalidated on hardware” | absent | absent | absent | absent | absent | partlycomputes config 2 instead | absent | absent |
| 23 | Clock-config writes bypass the verified SPI writer | partlySPI retry note | absent | raisedM1 | absent | absent | raisedF9 | absent | absent | raised§4.2 | absent | absent | absent | raisedtraced through four files | raisedF6, both call sites | partlyasks that writes be verified | partlycalls the heal machinery fragile | absentcalls the machinery retirable |
| 24 | Stock squelch re-arms FM audio (0x10=0x80) during monitoring | absentassumes it can’t re-arm | absent | absent | absent | partlynames squelch logic as a risk | raisedF1, new | absent“fixed” by forcing squelch open | partlynames the squelch path as a writer | partlyhazard flagged, “unlikely” | absent | absent | absent | raisedsame call chain, independently | absentownership named in general | raisedF1, reproduced on the host | absent | absent |
| 25 | Stale clear-call state releases a new call’s first frames | absent | absent | absent | absent | absent | raisedF8, probe | absent | absent | absent | absent | absent | absent | absent | absent | partlyvia identity, not the 1,120 samples | absent | absent |
| 26 | No frequency tracking; ad-hoc timing loop gains | partlya code comment calls the DC estimate biased | absent | absent | absent | absent | absent | absent | absent | absent | raised§4, new | absentreads the loop as sound | absent | partlythe ±1% clamp only | partlyasks for a timing/AFC loop | absent | raisedDC fit once per window, with the fix | raisedwith the carrier-offset arithmetic |
| 27 | Unknown talkgroup opens audio (fail-open gating) | absent | absent | absent | absent | absent | absent | absent | absent | absent | absent | absent | absent | absent | absent | raisedF6, new and reproduced | absent | absent |
From the provenance table in the Claude Fable 5.1 audit (17 Sep 2026), the newest audit on this task. Shaded columns are reviews that were already in the repository.
How each run went
Descriptive, not graded, and in grade order rather than ranked: one run per model. Run time is wall-clock, so it includes the provider’s speed, rate-limit waits and retries, and token counts vary with each model’s tokenizer. Every run followed Clean Room 1.0, which is what makes them comparable.
| Model | Run time | Agent steps | Tool calls | Output tokens | Route | Launch |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 25 min | 120 | 54 | 185K | Claude Code CLI, subscription | one-shot |
| GPT-6 Astra | 15 min | 29 | 42 | 21K | OpenAI subscription | one-shot |
| GPT-5.6 Sol | 14 min | 35 | 142 | 28K | OpenAI subscription | one-shot |
| Claude Fable 5.1 | 12 min | 71 | 44 | 249K | Claude Code CLI, subscription | one-shot |
| HY4 Preview | 15 min | 40 | 62 | 27K | OpenRouter API | one-shot |
| GPT-5.6 Luna | 24 min | 26 | 37 | 21K | OpenAI subscription | one-shot |
| UNIONALPHAThird attempt; the first two stopped on provider rate limits | 36 min | 83 | 95 | 63K | OpenRouter API | one-shot |
| Grok 4.6 | 5 min | 16 | 61 | 17K | xAI subscription | one-shot |
| Muse Spark 1.3 Contributor | 5 min | 27 | 38 | 19K | OpenRouter API | one-shot |
| DeepSeek V4.1 Flash | 35 min | 88 | 146 | 99K | OpenRouter API | one-shot |
| GLM 5.3 | 2 h 57 min | 143 | 169 | 512K | Alibaba Cloud Token Plan | interactive |
| Gemini 3.8 Flash | 10 min | 75 | 77 | 40K | OpenRouter API | one-shot |
| Qwen3.8 Max | 5 h 10 min | 71 | 111 | 598K | Alibaba Cloud Token Plan | interactive |
| Qwen3.8 FlashSecond attempt; the first stopped on a provider rate limit | 57 min | 92 | 97 | 147K | OpenRouter API | one-shot |
| GLM 5.3 Flash | 30 min | 46 | 58 | 35K | OpenRouter API | one-shot |
| DeepSeek V4 Pro | 8 min | 16 | 41 | 17K | Alibaba Cloud Token Plan | interactive |
Every run figure
| Model | Started (UTC) | Output tokens | Reasoning tokens | Context tokens | Messages | Sessions | Retry budget |
|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 18 Sep 2026, 02:44 | 185,228 | 118,897 | 17,469,122 | – | 1 | – |
| GPT-6 Astra | 18 Sep 2026, 01:06 | 21,399 | 3,777 | 3,081,430 | 72 | 1 | 10 |
| GPT-5.6 Sol | 18 Sep 2026, 00:18 | 28,443 | 11,067 | 5,432,761 | 263 | 1 | 10 |
| Claude Fable 5.1 | 18 Sep 2026, 03:10 | 249,363 | 167,839 | 8,687,277 | – | 1 | – |
| HY4 Preview | 17 Sep 2026, 23:57 | 27,431 | 14,261 | 2,847,148 | 103 | 1 | 10 |
| GPT-5.6 Luna | 18 Sep 2026, 00:37 | 20,851 | 10,792 | 3,678,578 | 164 | 1 | 10 |
| UNIONALPHA | 17 Sep 2026, 12:57 | 63,278 | not reported | 9,988,183 | 179 | 1 | 10 |
| Grok 4.6 | 17 Sep 2026, 04:25 | 16,928 | 9,183 | 1,300,749 | 78 | 1 | 3 |
| Muse Spark 1.3 Contributor | 17 Sep 2026, 06:10 | 19,449 | 7,815 | 2,309,483 | 66 | 1 | 3 |
| DeepSeek V4.1 Flash | 17 Sep 2026, 05:22 | 99,482 | 68,350 | 17,785,897 | 235 | 1 | 3 |
| GLM 5.3 | 17 Sep 2026, 20:07 | 512,246 | 485,366 | 19,494,026 | 313 | 1 | 10 |
| Gemini 3.8 Flash | 17 Sep 2026, 23:29 | 40,460 | 26,385 | 7,273,099 | 153 | 1 | 10 |
| Qwen3.8 Max | 17 Sep 2026, 14:52 | 597,788 | 582,689 | 12,033,596 | 183 | 1 | 10 |
| Qwen3.8 Flash | 17 Sep 2026, 13:49 | 146,532 | 125,802 | 13,381,380 | 190 | 1 | 10 |
| GLM 5.3 Flash | 17 Sep 2026, 06:35 | 35,256 | 23,180 | 4,819,946 | 106 | 1 | 3 |
| DeepSeek V4 Pro | 17 Sep 2026, 23:13 | 17,430 | 2,862 | 1,283,246 | 58 | 1 | 10 |
Context tokens are input plus cache reads and writes: mostly the conversation re-sent at every step, so they grow with steps × context size. The retry budget is how many API retries the harness allowed.