Head to head
Same brief, same rubric, same scale. Every model on the task side by side, or just the ones you pick: dimension scores, how many of their claims held up, and which issues each review raised.
It measured what everyone argued about.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model with no network tools, no skills, no plugins, no MCP servers and no subagents.
GPT-6 AstraReproduced, not asserted.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model through an isolated agent harness with no network, skills or memory.
The first to find all four, and the only one that measured them
Sixteen models have now reviewed this repository. The fourteen before this one all missed the vocoder CPU wall; this one profiled it in an emulator and tied it to the parser reset that ends the call. It also found both halves of the analog chain, the sample source and the frame-clock rule. GPT-6 Astra remains the best of the ranked audits and the one that found what this review skipped: the instruments.
From the Claude Opus 5 audit, 17 September 2026.
Weighted score
Out of 100, on the grade scale.
Earned and lost, by rubric slot
Each bar is a model’s 100-point grade split into the rubric’s weighted slots; filled is earned, empty is lost. The slots line up, so a gap in one bar against a full slot in another is exactly where the models part ways.
- 1Accuracy & evidence30 pts
- 2Coverage of decode problems25 pts
- 3Root cause & prioritisation15 pts
- 4Fix plan & acceptance gates15 pts
- 5Originality & attribution10 pts
- 6Clarity & calibration5 pts
Scores on each rubric dimension
One panel per model: its score on each rubric dimension and the weighted total over the grade zones, in its own colour, with every other model on this task as a grey dot for scale. Hover a row for every model’s score and the auditor’s reasoning.
- The panel’s model, joined down the dimensions
- Every other model on this task
Claude Opus 592 A
GPT-6 Astra89 B
Scores and weighted points
| Dimension | Weight | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | HY4 Preview | GPT-5.6 Luna | UNIONALPHA | Grok 4.6 | Muse Spark 1.3 Contributor | DeepSeek V4.1 Flash | GLM 5.3 | Gemini 3.8 Flash | Qwen3.8 Max | Qwen3.8 Flash | GLM 5.3 Flash | DeepSeek V4 Pro | Difference |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy & evidence | 30% | 94 · 28.2 pts | 92 · 27.6 pts | 91 · 27.3 pts | 85 · 25.5 pts | 85 · 25.5 pts | 82 · 24.6 pts | 90 · 27.0 pts | 74 · 22.2 pts | 74 · 22.2 pts | 73 · 21.9 pts | 68 · 20.4 pts | 65 · 19.5 pts | 70 · 21.0 pts | 62 · 18.6 pts | 64 · 19.2 pts | 61 · 18.3 pts | +2 |
| Coverage of decode problems | 25% | 86 · 21.5 pts | 84 · 21.0 pts | 70 · 17.5 pts | 62 · 15.5 pts | 56 · 14.0 pts | 58 · 14.5 pts | 65 · 16.3 pts | 59 · 14.8 pts | 54 · 13.5 pts | 49 · 12.3 pts | 52 · 13.0 pts | 46 · 11.5 pts | 39 · 9.8 pts | 40 · 10.0 pts | 34 · 8.5 pts | 33 · 8.3 pts | +2 |
| Root cause & prioritisation | 15% | 96 · 14.4 pts | 88 · 13.2 pts | 84 · 12.6 pts | 84 · 12.6 pts | 80 · 12.0 pts | 80 · 12.0 pts | 60 · 9.0 pts | 70 · 10.5 pts | 72 · 10.8 pts | 66 · 9.9 pts | 74 · 11.1 pts | 74 · 11.1 pts | 71 · 10.7 pts | 68 · 10.2 pts | 62 · 9.3 pts | 64 · 9.6 pts | +8 |
| Fix plan & acceptance gates | 15% | 93 · 14.0 pts | 92 · 13.8 pts | 90 · 13.5 pts | 86 · 12.9 pts | 84 · 12.6 pts | 80 · 12.0 pts | 70 · 10.5 pts | 75 · 11.3 pts | 71 · 10.7 pts | 70 · 10.5 pts | 68 · 10.2 pts | 70 · 10.5 pts | 71 · 10.7 pts | 70 · 10.5 pts | 62 · 9.3 pts | 60 · 9.0 pts | +1 |
| Originality & attribution | 10% | 95 · 9.5 pts | 90 · 9.0 pts | 84 · 8.4 pts | 84 · 8.4 pts | 82 · 8.2 pts | 72 · 7.2 pts | 76 · 7.6 pts | 72 · 7.2 pts | 70 · 7.0 pts | 74 · 7.4 pts | 70 · 7.0 pts | 70 · 7.0 pts | 66 · 6.6 pts | 64 · 6.4 pts | 58 · 5.8 pts | 62 · 6.2 pts | +5 |
| Clarity & calibration | 5% | 92 · 4.6 pts | 86 · 4.3 pts | 86 · 4.3 pts | 88 · 4.4 pts | 88 · 4.4 pts | 84 · 4.2 pts | 80 · 4.0 pts | 78 · 3.9 pts | 80 · 4.0 pts | 74 · 3.7 pts | 72 · 3.6 pts | 84 · 4.2 pts | 78 · 3.9 pts | 70 · 3.5 pts | 68 · 3.4 pts | 48 · 2.4 pts | +6 |
| Weighted total | 100% | 92 · A | 89 · B | 84 · B | 79 · C+ | 77 · C+ | 75 · C+ | 74 · C | 70 · C | 68 · C− | 66 · C− | 65 · C− | 64 · D | 63 · D | 59 · D | 56 · D | 54 · D | +3 |
Every claim, checked
One cell per claim in each audit’s claim table, checked at the lines it cites. Hover a cell for the claim and the auditor’s note. The reviews made different numbers of claims, so the bar under each grid gives the shares.
- Holds
- Qualifiedoverstated, miscounted or doubtful
- Wrong
20 hold · 1 qualified · 0 wrong · 95% hold
27 hold · 2 qualified · 0 wrong · 93% hold
What each review raised
Claude Opus 5 found 4 of the 4 decode-critical issues; GPT-6 Astra found 3 of the 4 decode-critical issues. Every tracked issue below, with the reviews that were already in the repository for reference.
Issue signatures
Each review as a trace across the tracked issues. Where traces part is where the reviews differ; hover a column for every review’s entry. Column numbers match the table below.
- raised
- partly
- absent
- Decode-critical
| # | Issue | Project docsEarlier repo review · in the repository | GPT-6 AstraModel · Sep 17 | Claude Opus 5Model · Sep 17 |
|---|---|---|---|---|
| 1 | Capture export stages the wrong I2S halfwords | absent | raisedF4, reproduced on the host | absent |
| 2 | Capture parser certifies an incomplete stream | absent | partlyasks for tests at the callback | absent |
| 3 | pcm_starve never increments | absentdocumented as working | raisedF9, reproduced | absent |
| 4 | No static RAM margin | partlymargins still to measure | raisedF8, from its own link map | raisedwith the overlay buffers named |
| 5 | The I2S stream is most likely microphone audio Decode-critical | partlyopen, leaning sceptical | raisedF2, as far as the evidence goes | raisedproved from 0xE0=0x89 and sound.c |
| 6 | HR-C6000 de-emphasis on the capture path Decode-critical | absent | raisedF2, 0x34=0x3C | raisedmeasured at 48.9% SER |
| 7 | AT1846S FM filters, low-frequency bit, 25 kHz Decode-critical | partly“require characterization” | raisedF2, the FM settings table | raisedall three, with the DMR fix |
| 8 | 0x10=0x6E hybrid state; 0x36 dual role | partlybring-up clock rules | partly0x36 and 0x10 via the squelch path | partly0x36 and 0xE0, not 0x10 |
| 9 | Manual: I2S frame clock “must be 8KHz” Decode-criticalmissed or ruled out | absent | raisedF7, with the divisor arithmetic | raisedH2 |
| 10 | One-layer 4FSK test mode as a P25 tap missed or ruled out | partlystock BER-test block only | partlypoints at the layer architecture | raisedRoute A, with an acceptance test |
| 11 | ±10% health gate versus ±1% timing clamp | absent | raisedF7 | absent |
| 12 | Fail-closed muting at LDU cadence | absent | raisedF6, the late-entry half | partlyvia the MFID path |
| 13 | Non-standard MFID mutes clear calls | absent | absent | raisedM1, with the Motorola case |
| 14 | Test waveform shares the receiver’s RRC filter | partly“synthetic RRC/AWGN” caveat | raisedF10, with why the tests missed F1/F4 | raisedtested, called harmless |
| 15 | MCU runs at 72 MHz | raised | raisedF9, the target for measurement | raisedand measured against it |
| 16 | Vocoder needs 11–16× the 72 MHz CPU Decode-criticalmissed or ruled out | partlydecode timing unmeasured | partlycounted the filter, not the vocoder | raisedmeasured 6.5× on the demo build |
| 17 | Direct discriminator tap (M17 mod) | raised | raisedF2, ADC or another interface | raisedRoute B, with the ADC channels checked |
| 18 | Phase 2 architecture and scope | partlynot implemented | raisedthe most complete of the fourteen | raisedincluding the descrambler seed |
| 19 | Two unverified SPI writes per decoded 20 ms frame | absent | partlythe sink is cited, not the writes | absent |
| 20 | Capture sessions lack epochs | absent | raisedF4, segment on retune and clock change | absent |
| 21 | Ring and tick real-time budget | partlydeadlines unproven | raisedF9, all four figures exact | raisedthe 21 ms ring against a 55–190 ms stall |
| 22 | Clock config 3 assumes 12,288 Hz; the codec formula gives 12,000 | absent | partlycomputes config 2 instead | absent |
| 23 | Clock-config writes bypass the verified SPI writer | partlySPI retry note | partlyasks that writes be verified | partlycalls the heal machinery fragile |
| 24 | Stock squelch re-arms FM audio (0x10=0x80) during monitoring | absentassumes it can’t re-arm | raisedF1, reproduced on the host | absent |
| 25 | Stale clear-call state releases a new call’s first frames | absent | partlyvia identity, not the 1,120 samples | absent |
| 26 | No frequency tracking; ad-hoc timing loop gains | partlya code comment calls the DC estimate biased | absent | raisedDC fit once per window, with the fix |
| 27 | Unknown talkgroup opens audio (fail-open gating) | absent | raisedF6, new and reproduced | absent |
From the provenance table in the Claude Fable 5.1 audit (17 Sep 2026), the newest audit on this task. Shaded columns are reviews that were already in the repository.
How each run went
Descriptive, not graded, and in grade order rather than ranked: one run per model. Run time is wall-clock, so it includes the provider’s speed, rate-limit waits and retries, and token counts vary with each model’s tokenizer. Every run followed Clean Room 1.0, which is what makes them comparable.
| Model | Run time | Agent steps | Tool calls | Output tokens | Route | Launch |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 25 min | 120 | 54 | 185K | Claude Code CLI, subscription | one-shot |
| GPT-6 Astra | 15 min | 29 | 42 | 21K | OpenAI subscription | one-shot |
Every run figure
| Model | Started (UTC) | Output tokens | Reasoning tokens | Context tokens | Messages | Sessions | Retry budget |
|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 18 Sep 2026, 02:44 | 185,228 | 118,897 | 17,469,122 | – | 1 | – |
| GPT-6 Astra | 18 Sep 2026, 01:06 | 21,399 | 3,777 | 3,081,430 | 72 | 1 | 10 |
Context tokens are input plus cache reads and writes: mostly the conversation re-sent at every step, so they grow with steps × context size. The retry budget is how many API retries the harness allowed.