Head to head
Same brief, same rubric, same scale. Every model on the task side by side, or just the ones you pick: dimension scores, how many of their claims held up, and which issues each review raised.
Accurate to the byte. Hesitant about what stops decoding.
Grok 4.6Finds what erases the signal. Misses what stalls the voice.
Muse Spark 1.3 ContributorFinds both halves of the filter problem. Points the CPU fix the wrong way.
DeepSeek V4.1 FlashFinds the microphone and the filters. Waves de-emphasis through.
Qwen3.8 MaxBest case yet for the microphone. Never looks at the CPU.
Qwen3.8 FlashRight about the filters. Wrong about the CPU.
GLM 5.3 FlashAudits the memory to the byte. Misreads the receive path.
Every model here ran under the same conditions. How the runs were set up →
Most accurate, least decisive
All five clean-room audits miss the vocoder CPU wall. UNIONALPHA makes no wrong claims where Grok 4.6, DeepSeek V4.1 Flash and Muse Spark 1.3 Contributor make two or three each, and it covers the most issues, but it fully finds one decode-critical issue where Grok 4.6 finds three.
From the UNIONALPHA audit, 17 September 2026.
Weighted score
Out of 100, on the grade scale.
Earned and lost, by rubric slot
Each bar is a model’s 100-point grade split into the rubric’s weighted slots; filled is earned, empty is lost. The slots line up, so a gap in one bar against a full slot in another is exactly where the models part ways.
- 1Accuracy & evidence30 pts
- 2Coverage of decode problems25 pts
- 3Root cause & prioritisation15 pts
- 4Fix plan & acceptance gates15 pts
- 5Originality & attribution10 pts
- 6Clarity & calibration5 pts
Scores on each rubric dimension
One panel per model: its score on each rubric dimension and the weighted total over the grade zones, in its own colour, with every other model on this task as a grey dot for scale. Hover a row for every model’s score and the auditor’s reasoning.
- The panel’s model, joined down the dimensions
- Every other model on this task
UNIONALPHA74 C
Grok 4.670 C
Muse Spark 1.3 Contributor68 C−
DeepSeek V4.1 Flash66 C−
Qwen3.8 Max63 D
Qwen3.8 Flash59 D
GLM 5.3 Flash56 D
Scores and weighted points
| Dimension | Weight | UNIONALPHA | Grok 4.6 | Muse Spark 1.3 Contributor | DeepSeek V4.1 Flash | Qwen3.8 Max | Qwen3.8 Flash | GLM 5.3 Flash |
|---|---|---|---|---|---|---|---|---|
| Accuracy & evidence | 30% | 90 · 27.0 pts | 74 · 22.2 pts | 74 · 22.2 pts | 73 · 21.9 pts | 70 · 21.0 pts | 62 · 18.6 pts | 64 · 19.2 pts |
| Coverage of decode problems | 25% | 65 · 16.3 pts | 59 · 14.8 pts | 54 · 13.5 pts | 49 · 12.3 pts | 39 · 9.8 pts | 40 · 10.0 pts | 34 · 8.5 pts |
| Root cause & prioritisation | 15% | 60 · 9.0 pts | 70 · 10.5 pts | 72 · 10.8 pts | 66 · 9.9 pts | 71 · 10.7 pts | 68 · 10.2 pts | 62 · 9.3 pts |
| Fix plan & acceptance gates | 15% | 70 · 10.5 pts | 75 · 11.3 pts | 71 · 10.7 pts | 70 · 10.5 pts | 71 · 10.7 pts | 70 · 10.5 pts | 62 · 9.3 pts |
| Originality & attribution | 10% | 76 · 7.6 pts | 72 · 7.2 pts | 70 · 7.0 pts | 74 · 7.4 pts | 66 · 6.6 pts | 64 · 6.4 pts | 58 · 5.8 pts |
| Clarity & calibration | 5% | 80 · 4.0 pts | 78 · 3.9 pts | 80 · 4.0 pts | 74 · 3.7 pts | 78 · 3.9 pts | 70 · 3.5 pts | 68 · 3.4 pts |
| Weighted total | 100% | 74 · C | 70 · C | 68 · C− | 66 · C− | 63 · D | 59 · D | 56 · D |
Every claim, checked
One cell per claim in each audit’s claim table, checked at the lines it cites. Hover a cell for the claim and the auditor’s note. The reviews made different numbers of claims, so the bar under each grid gives the shares.
UNIONALPHA’s claim table · Grok 4.6’s claim table · Muse Spark 1.3 Contributor’s claim table · DeepSeek V4.1 Flash’s claim table · Qwen3.8 Max’s claim table · Qwen3.8 Flash’s claim table · GLM 5.3 Flash’s claim table
- Holds
- Qualifiedoverstated, miscounted or doubtful
- Wrong
33 hold · 2 qualified · 0 wrong · 94% hold
21 hold · 7 qualified · 2 wrong · 70% hold
22 hold · 6 qualified · 2 wrong · 73% hold
28 hold · 4 qualified · 3 wrong · 80% hold
21 hold · 4 qualified · 2 wrong · 78% hold
21 hold · 6 qualified · 6 wrong · 64% hold
17 hold · 7 qualified · 4 wrong · 61% hold
What each review raised
UNIONALPHA found 1 of the 4 decode-critical issues; Grok 4.6 found 3 of the 4 decode-critical issues; Muse Spark 1.3 Contributor found 2 of the 4 decode-critical issues; DeepSeek V4.1 Flash found 2 of the 4 decode-critical issues; Qwen3.8 Max found 2 of the 4 decode-critical issues; Qwen3.8 Flash found 2 of the 4 decode-critical issues; GLM 5.3 Flash found 1 of the 4 decode-critical issues. Every tracked issue below, with the reviews that were already in the repository for reference.
Issue signatures
Each review as a trace across the tracked issues. Where traces part is where the reviews differ; hover a column for every review’s entry. Column numbers match the table below.
- raised
- partly
- absent
- Decode-critical
| # | Issue | Project docsEarlier repo review · in the repository | Grok 4.6Model · Sep 17 | DeepSeek V4.1 FlashModel · Sep 17 | Muse Spark 1.3 ContributorModel · Sep 17 | GLM 5.3 FlashModel · Sep 17 | UNIONALPHAModel · Sep 17 | Qwen3.8 FlashModel · Sep 17 | Qwen3.8 MaxModel · Sep 17 |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Capture export stages the wrong I2S halfwords | absent | raised§1, first code change | absentits first experiment uses that capture | absenttrusts the capture | absentrelies on the capture | raisedF3, fixed and verified | raisedP1-2, tests miss the adapter | absentcalls the capture validated |
| 2 | Capture parser certifies an incomplete stream | absent | absent | absent | absenttrusts the parser | absentrelies on the parser | absent | absent | absentrelies on the parser |
| 3 | pcm_starve never increments | absentdocumented as working | absent | absent | absentrelies on it | absent | raisedF7 | raisedP1-7 | absent |
| 4 | No static RAM margin | partlymargins still to measure | absent | raisedM6, 0 bytes free | partly“nearly full”, from the docs | raisedF5, byte-exact map audit | raisedF6, byte-exact | partlycalls the reservation headroom | absent |
| 5 | The I2S stream is most likely microphone audio Decode-critical | partlyopen, leaning sceptical | raised§1, 0x89 versus 0xC9 | raisedB1, four-value test | raisedP1-1 | partlyopen, leans towards RF | partlyblocking, but never says microphone | raisedP1-4, 0xE0 mic bit | raisedA, the clearest case yet |
| 6 | HR-C6000 de-emphasis on the capture path Decode-critical | absent | raised§2, bit 5 of 0x34 | absentcalls it benign | raisedP1-2, 0x34=0x3C | raisedF2, closes the eye | partlycited, called unmeasured | raisedP1-3, 0x34=0x1C | raisedC, with the 0x34 bit 5 fix |
| 7 | AT1846S FM filters, low-frequency bit, 25 kHz Decode-criticalmissed | partly“require characterization” | raised§2, register level | raisedB2, filter register | raisedP1-2, 0x58 filters | partly“voice filtering”; wrong bandwidth premise | partlycited, called unmeasured | raisedP1-3, DMR 0x58 probe | absent |
| 8 | 0x10=0x6E hybrid state; 0x36 dual role | partlybring-up clock rules | partlymisses 0x6E and the 0x36 clock gate | partly“undocumented hybrid state” | partlyquiet-chip registers | raisedF4, 0x10=0x80 kills the clock | partlyF1, 0x6E against 0x80 | partly“hybrid I2S state” | partlyslot engine and eco only |
| 9 | Manual: I2S frame clock “must be 8KHz” Decode-criticalmissed | absent | raised§3 | absentquotes the paragraph, not the rule | absent | absent | raisedF4 | absentquotes the formulas, not the rule | absentcites the section, not the rule |
| 10 | One-layer 4FSK test mode as a P25 tap | partlystock BER-test block only | raisedGate D | raisedB5, exact recipe | absentdismissed | absent | partlyworth a bounded test | absentruled out at “9600 Bd” | raisedstep 2, a symbol source |
| 11 | ±10% health gate versus ±1% timing clamp | absent | absent | raisedB3, impact overstated | raisedP1-3 | raisedF3 | absent | absent | absent |
| 12 | Fail-closed muting at LDU cadence | absent | absentlate-entry mute only | absent | absentcalls it an asset | absent | absent | partly“keep it” | absent |
| 13 | Non-standard MFID mutes clear calls | absent | absent | absent | absent | absent | absent | absent | absent |
| 14 | Test waveform shares the receiver’s RRC filter | partly“synthetic RRC/AWGN” caveat | absent | partlytested it, says not to fix | partlysynthetic only, wants recordings | absentwould extend that model | raisedF5, unquantified | absentwould extend that model | partlysynthetic only, not the circularity |
| 15 | MCU runs at 72 MHz | raised | raisedin passing | absent | raisedP1-5 | absent | raisedF7 | absent | absent |
| 16 | Vocoder needs 11–16× the 72 MHz CPU Decode-criticalmissed | partlydecode timing unmeasured | absent“fine on a 1 ms tick” | absent“vocoder question settled” | partlyunmeasured; fix order backwards | absent“in good shape” | partlydeadlines “unproven” | absent“not the problem” | absent“not the problem” |
| 17 | Direct discriminator tap (M17 mod) | raised | raiseduncredited | raisedpins, timer ADC, 48 kS/s | raisedfallback, pin 9 | raisedfallback | raisedfallback | raisedfallback | raisedstep 5, the likely answer |
| 18 | Phase 2 architecture and scope | partlynot implemented | partlymisplaces the AMBE+2 decoder | partlyvoice via mbelib AMBE+2 | raisedwith RF band limits | partlysays mbelib has no AMBE+2 | raisedmost accurate section | partlyDMR and X2-TDMA parts | raisedAMBE+2 present but uncalled |
| 19 | Two unverified SPI writes per decoded 20 ms frame | absent | absent | absent | absent | absenttreats them as protection | absent | absent | absent |
| 20 | Capture sessions lack epochs | absent | absent | partlymeasured=0 only | absent | absent | raisedF9, 8 kHz under a 24 kHz header | absent | absent |
| 21 | Ring and tick real-time budget | partlydeadlines unproven | partlycalls it fine | absent | partlyoverruns look like weak RF | absent | raisedF7, 1 ms is a minimum | partlyregister stalls against the ring | absent |
| 22 | Clock config 3 assumes 12,288 Hz; the codec formula gives 12,000 | absent | absent | raisedB3, clock model | absent | partly“guessed semantics” | absent | raisedP1-6, for a different reason | absent |
| 23 | Clock-config writes bypass the verified SPI writer | partlySPI retry note | absent | raisedM1 | absent | absent | raisedF9 | absent | absent |
| 24 | Stock squelch re-arms FM audio (0x10=0x80) during monitoring | absentassumes it can’t re-arm | absent | absent | absent | partlynames squelch logic as a risk | raisedF1, new | absent“fixed” by forcing squelch open | partlynames the squelch path as a writer |
| 25 | Stale clear-call state releases a new call’s first frames | absent | absent | absent | absent | absent | raisedF8, probe | absent | absent |
From the provenance table in the Qwen3.8 Max audit (17 Sep 2026), the newest audit on this task. Shaded columns are reviews that were already in the repository.
How each run went
Descriptive, not graded, and in grade order rather than ranked: one run per model. Run time is wall-clock, so it includes the provider’s speed, rate-limit waits and retries, and token counts vary with each model’s tokenizer. Every run followed Clean Room 1.0, which is what makes them comparable.
| Model | Run time | Agent steps | Tool calls | Output tokens | Route | Launch |
|---|---|---|---|---|---|---|
| UNIONALPHAThird attempt; the first two stopped on provider rate limits | 36 min | 83 | 95 | 63K | OpenRouter API | one-shot |
| Grok 4.6 | 5 min | 16 | 61 | 17K | xAI subscription | one-shot |
| Muse Spark 1.3 Contributor | 5 min | 27 | 38 | 19K | OpenRouter API | one-shot |
| DeepSeek V4.1 Flash | 35 min | 88 | 146 | 99K | OpenRouter API | one-shot |
| Qwen3.8 Max | 5 h 10 min | 71 | 111 | 598K | Alibaba Cloud Token Plan | interactive |
| Qwen3.8 FlashSecond attempt; the first stopped on a provider rate limit | 57 min | 92 | 97 | 147K | OpenRouter API | one-shot |
| GLM 5.3 Flash | 30 min | 46 | 58 | 35K | OpenRouter API | one-shot |
Every run figure
| Model | Started (UTC) | Output tokens | Reasoning tokens | Context tokens | Messages | Sessions | Retry budget |
|---|---|---|---|---|---|---|---|
| UNIONALPHA | 17 Sep 2026, 12:57 | 63,278 | not reported | 9,988,183 | 179 | 1 | 10 |
| Grok 4.6 | 17 Sep 2026, 04:25 | 16,928 | 9,183 | 1,300,749 | 78 | 1 | 3 |
| Muse Spark 1.3 Contributor | 17 Sep 2026, 06:10 | 19,449 | 7,815 | 2,309,483 | 66 | 1 | 3 |
| DeepSeek V4.1 Flash | 17 Sep 2026, 05:22 | 99,482 | 68,350 | 17,785,897 | 235 | 1 | 3 |
| Qwen3.8 Max | 17 Sep 2026, 14:52 | 597,788 | 582,689 | 12,033,596 | 183 | 1 | 10 |
| Qwen3.8 Flash | 17 Sep 2026, 13:49 | 146,532 | 125,802 | 13,381,380 | 190 | 1 | 10 |
| GLM 5.3 Flash | 17 Sep 2026, 06:35 | 35,256 | 23,180 | 4,819,946 | 106 | 1 | 3 |
Context tokens are input plus cache reads and writes: mostly the conversation re-sent at every step, so they grow with steps × context size. The retry budget is how many API retries the harness allowed.