Head to head
Same brief, same rubric, same scale. Put up to three models side by side: dimension scores, how many of their claims held up, and which issues each review raised.
Finds both halves of the filter problem. Points the CPU fix the wrong way.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model through an isolated agent harness with no network, skills or memory.
B Grok 4.6Finds what erases the signal. Misses what stalls the voice.
Clean room: a fresh copy of the repository (46eebda with its uncommitted changes, its own docs and the HR-C6000 manual) with no earlier reviews, reviewed by one model through an isolated agent harness with no network, skills or memory.
Four audits, four different blind spots
None of the four clean-room audits finds the vocoder CPU wall. Grok 4.6 adds de-emphasis, the 8 kHz rule and the capture bug; DeepSeek V4.1 Flash adds the clock model but clears de-emphasis; Muse Spark 1.3 Contributor adds de-emphasis and the gate mismatch, flags CPU as unmeasured, and has the best-calibrated prose of the three, but trusts the broken instruments and points its CPU fix at the wrong stage. GLM 5.3 Flash audits memory to the byte but misreads the receive path.
From the Muse Spark 1.3 Contributor audit, 17 September 2026.
Weighted score
Out of 100, on the grade scale.
Earned and lost, by rubric slot
Each bar is a model’s 100-point grade split into the rubric’s weighted slots; filled is earned, empty is lost. The slots line up, so a gap in one bar against a full slot in another is exactly where the models part ways.
- 1Accuracy & evidence30 pts
- 2Coverage of decode problems25 pts
- 3Root cause & prioritisation15 pts
- 4Fix plan & acceptance gates15 pts
- 5Originality & attribution10 pts
- 6Clarity & calibration5 pts
Scores on each rubric dimension
Each dimension out of 100, then the weighted total, over the grade zones. The compared models are the coloured dots, each joined down the dimensions by a thin line; the other models on this task are the small grey ones. Hover or focus a row for every score and the auditor’s reasoning.
- AMuse Spark 1.3 Contributor
- BGrok 4.6
- 2 other models on this task
Scores and weighted points
| Dimension | Weight | A · Muse Spark 1.3 Contributor | B · Grok 4.6 | DeepSeek V4.1 Flash | GLM 5.3 Flash | A − B |
|---|---|---|---|---|---|---|
| Accuracy & evidence | 30% | 74 · 22.2 pts | 74 · 22.2 pts | 73 · 21.9 pts | 64 · 19.2 pts | ±0 |
| Coverage of decode problems | 25% | 54 · 13.5 pts | 59 · 14.8 pts | 49 · 12.3 pts | 34 · 8.5 pts | −5 |
| Root cause & prioritisation | 15% | 72 · 10.8 pts | 70 · 10.5 pts | 66 · 9.9 pts | 62 · 9.3 pts | +2 |
| Fix plan & acceptance gates | 15% | 71 · 10.7 pts | 75 · 11.3 pts | 70 · 10.5 pts | 62 · 9.3 pts | −4 |
| Originality & attribution | 10% | 70 · 7.0 pts | 72 · 7.2 pts | 74 · 7.4 pts | 58 · 5.8 pts | −2 |
| Clarity & calibration | 5% | 80 · 4.0 pts | 78 · 3.9 pts | 74 · 3.7 pts | 68 · 3.4 pts | +2 |
| Weighted total | 100% | 68 · C− | 70 · C | 66 · C− | 56 · D | −2 |
Every claim, checked
One cell per claim in each audit’s claim table, checked at the lines it cites. Hover a cell for the claim and the auditor’s note. The reviews made different numbers of claims, so the bar under each grid gives the shares.
Muse Spark 1.3 Contributor’s claim table · Grok 4.6’s claim table
- Holds
- Qualifiedoverstated, miscounted or doubtful
- Wrong
22 hold · 6 qualified · 2 wrong · 73% hold
21 hold · 7 qualified · 2 wrong · 70% hold
What each review raised
Muse Spark 1.3 Contributor found 2 of the 4 decode-critical issues; Grok 4.6 found 3 of the 4 decode-critical issues. Every tracked issue below, with the reviews that were already in the repository for reference.
Issue signatures
Each review as a trace across the tracked issues. Where traces part is where the reviews differ; hover a column for every review’s entry. Column numbers match the table below.
- raised
- partly
- absent
- Decode-critical
| # | Issue | Project docsEarlier repo review · in the repository | Grok 4.6Model · Sep 17 | Muse Spark 1.3 ContributorModel · Sep 17 |
|---|---|---|---|---|
| 1 | Capture export stages the wrong I2S halfwords | absent | raised§1, first code change | absenttrusts the capture |
| 2 | Capture parser certifies an incomplete stream | absent | absent | absenttrusts the parser |
| 3 | pcm_starve never increments | absentdocumented as working | absent | absentrelies on it |
| 4 | No static RAM margin | partlymargins still to measure | absent | partly“nearly full”, from the docs |
| 5 | The I2S stream is most likely microphone audio Decode-criticalmissed | partlyopen, leaning sceptical | raised§1, 0x89 versus 0xC9 | raisedP1-1 |
| 6 | HR-C6000 de-emphasis on the capture path Decode-critical | absent | raised§2, bit 5 of 0x34 | raisedP1-2, 0x34=0x3C |
| 7 | AT1846S FM filters, low-frequency bit, 25 kHz Decode-critical | partly“require characterization” | raised§2, register level | raisedP1-2, 0x58 filters |
| 8 | 0x10=0x6E hybrid state; 0x36 dual role | partlybring-up clock rules | partlymisses 0x6E and the 0x36 clock gate | partlyquiet-chip registers |
| 9 | Manual: I2S frame clock “must be 8KHz” Decode-criticalmissed | absent | raised§3 | absent |
| 10 | One-layer 4FSK test mode as a P25 tap | partlystock BER-test block only | raisedGate D | absentdismissed |
| 11 | ±10% health gate versus ±1% timing clamp | absent | absent | raisedP1-3 |
| 12 | Fail-closed muting at LDU cadence | absent | absentlate-entry mute only | absentcalls it an asset |
| 13 | Non-standard MFID mutes clear calls | absent | absent | absent |
| 14 | Test waveform shares the receiver’s RRC filter | partly“synthetic RRC/AWGN” caveat | absent | partlysynthetic only, wants recordings |
| 15 | MCU runs at 72 MHz | raised | raisedin passing | raisedP1-5 |
| 16 | Vocoder needs 11–16× the 72 MHz CPU Decode-criticalmissed | partlydecode timing unmeasured | absent“fine on a 1 ms tick” | partlyunmeasured; fix order backwards |
| 17 | Direct discriminator tap (M17 mod) | raised | raiseduncredited | raisedfallback, pin 9 |
| 18 | Phase 2 architecture and scope | partlynot implemented | partlymisplaces the AMBE+2 decoder | raisedwith RF band limits |
| 19 | Two unverified SPI writes per decoded 20 ms frame | absent | absent | absent |
| 20 | Capture sessions lack epochs | absent | absent | absent |
| 21 | Ring and tick real-time budget | partlydeadlines unproven | partlycalls it fine | partlyoverruns look like weak RF |
| 22 | Clock config 3 assumes 12,288 Hz; the codec formula gives 12,000 | absent | absent | absent |
| 23 | Clock-config writes bypass the verified SPI writer | partlySPI retry note | absent | absent |
From the provenance table in the GLM 5.3 Flash audit (17 Sep 2026), the newest audit on this task. Shaded columns are reviews that were already in the repository.