Head to head
Same brief, same rubric, same scale. Put up to three models side by side: dimension scores, how many of their claims held up, and which issues each review raised.
Audits the memory to the byte. Misreads the receive path.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model through an isolated agent harness with no network, skills or memory.
B Grok 4.6Finds what erases the signal. Misses what stalls the voice.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model through an isolated agent harness with no network, skills or memory.
Best on memory, last on the rest
All four clean-room audits miss the vocoder CPU wall. GLM 5.3 Flash is the only one to audit memory in detail, and one of three to name de-emphasis, but it finds one decode-critical issue where Grok 4.6 finds three and DeepSeek V4.1 Flash and Muse Spark 1.3 Contributor two each, and it makes more wrong claims about the receive path than any of them.
From the GLM 5.3 Flash audit, 17 September 2026.
Weighted score
Out of 100, on the grade scale.
Earned and lost, by rubric slot
Each bar is a model’s 100-point grade split into the rubric’s weighted slots; filled is earned, empty is lost. The slots line up, so a gap in one bar against a full slot in another is exactly where the models part ways.
- 1Accuracy & evidence30 pts
- 2Coverage of decode problems25 pts
- 3Root cause & prioritisation15 pts
- 4Fix plan & acceptance gates15 pts
- 5Originality & attribution10 pts
- 6Clarity & calibration5 pts
Scores on each rubric dimension
Each dimension out of 100, then the weighted total, over the grade zones. The compared models are the coloured dots, each joined down the dimensions by a thin line; the other models on this task are the small grey ones. Hover or focus a row for every score and the auditor’s reasoning.
- AGLM 5.3 Flash
- BGrok 4.6
- 2 other models on this task
Scores and weighted points
| Dimension | Weight | A · GLM 5.3 Flash | B · Grok 4.6 | Muse Spark 1.3 Contributor | DeepSeek V4.1 Flash | A − B |
|---|---|---|---|---|---|---|
| Accuracy & evidence | 30% | 64 · 19.2 pts | 74 · 22.2 pts | 74 · 22.2 pts | 73 · 21.9 pts | −10 |
| Coverage of decode problems | 25% | 34 · 8.5 pts | 59 · 14.8 pts | 54 · 13.5 pts | 49 · 12.3 pts | −25 |
| Root cause & prioritisation | 15% | 62 · 9.3 pts | 70 · 10.5 pts | 72 · 10.8 pts | 66 · 9.9 pts | −8 |
| Fix plan & acceptance gates | 15% | 62 · 9.3 pts | 75 · 11.3 pts | 71 · 10.7 pts | 70 · 10.5 pts | −13 |
| Originality & attribution | 10% | 58 · 5.8 pts | 72 · 7.2 pts | 70 · 7.0 pts | 74 · 7.4 pts | −14 |
| Clarity & calibration | 5% | 68 · 3.4 pts | 78 · 3.9 pts | 80 · 4.0 pts | 74 · 3.7 pts | −10 |
| Weighted total | 100% | 56 · D | 70 · C | 68 · C− | 66 · C− | −14 |
Every claim, checked
One cell per claim in each audit’s claim table, checked at the lines it cites. Hover a cell for the claim and the auditor’s note. The reviews made different numbers of claims, so the bar under each grid gives the shares.
- Holds
- Qualifiedoverstated, miscounted or doubtful
- Wrong
17 hold · 7 qualified · 4 wrong · 61% hold
21 hold · 7 qualified · 2 wrong · 70% hold
What each review raised
GLM 5.3 Flash found 1 of the 4 decode-critical issues; Grok 4.6 found 3 of the 4 decode-critical issues. Every tracked issue below, with the reviews that were already in the repository for reference.
Issue signatures
Each review as a trace across the tracked issues. Where traces part is where the reviews differ; hover a column for every review’s entry. Column numbers match the table below.
- raised
- partly
- absent
- Decode-critical
| # | Issue | Project docsEarlier repo review · in the repository | Grok 4.6Model · Sep 17 | GLM 5.3 FlashModel · Sep 17 |
|---|---|---|---|---|
| 1 | Capture export stages the wrong I2S halfwords | absent | raised§1, first code change | absentrelies on the capture |
| 2 | Capture parser certifies an incomplete stream | absent | absent | absentrelies on the parser |
| 3 | pcm_starve never increments | absentdocumented as working | absent | absent |
| 4 | No static RAM margin | partlymargins still to measure | absent | raisedF5, byte-exact map audit |
| 5 | The I2S stream is most likely microphone audio Decode-criticalmissed | partlyopen, leaning sceptical | raised§1, 0x89 versus 0xC9 | partlyopen, leans towards RF |
| 6 | HR-C6000 de-emphasis on the capture path Decode-critical | absent | raised§2, bit 5 of 0x34 | raisedF2, closes the eye |
| 7 | AT1846S FM filters, low-frequency bit, 25 kHz Decode-critical | partly“require characterization” | raised§2, register level | partly“voice filtering”; wrong bandwidth premise |
| 8 | 0x10=0x6E hybrid state; 0x36 dual role | partlybring-up clock rules | partlymisses 0x6E and the 0x36 clock gate | raisedF4, 0x10=0x80 kills the clock |
| 9 | Manual: I2S frame clock “must be 8KHz” Decode-criticalmissed | absent | raised§3 | absent |
| 10 | One-layer 4FSK test mode as a P25 tap | partlystock BER-test block only | raisedGate D | absent |
| 11 | ±10% health gate versus ±1% timing clamp | absent | absent | raisedF3 |
| 12 | Fail-closed muting at LDU cadence | absent | absentlate-entry mute only | absent |
| 13 | Non-standard MFID mutes clear calls | absent | absent | absent |
| 14 | Test waveform shares the receiver’s RRC filter | partly“synthetic RRC/AWGN” caveat | absent | absentwould extend that model |
| 15 | MCU runs at 72 MHz | raised | raisedin passing | absent |
| 16 | Vocoder needs 11–16× the 72 MHz CPU Decode-criticalmissed | partlydecode timing unmeasured | absent“fine on a 1 ms tick” | absent“in good shape” |
| 17 | Direct discriminator tap (M17 mod) | raised | raiseduncredited | raisedfallback |
| 18 | Phase 2 architecture and scope | partlynot implemented | partlymisplaces the AMBE+2 decoder | partlysays mbelib has no AMBE+2 |
| 19 | Two unverified SPI writes per decoded 20 ms frame | absent | absent | absenttreats them as protection |
| 20 | Capture sessions lack epochs | absent | absent | absent |
| 21 | Ring and tick real-time budget | partlydeadlines unproven | partlycalls it fine | absent |
| 22 | Clock config 3 assumes 12,288 Hz; the codec formula gives 12,000 | absent | absent | partly“guessed semantics” |
| 23 | Clock-config writes bypass the verified SPI writer | partlySPI retry note | absent | absent |
From the provenance table in the GLM 5.3 Flash audit (17 Sep 2026), the newest audit on this task. Shaded columns are reviews that were already in the repository.