Model review audits
Compare · P25 firmware review

Head to head

Same brief, same rubric, same scale. Every model on the task side by side, or just the ones you pick: dimension scores, how many of their claims held up, and which issues each review raised.

Models 6 of 6
Untick models to compare fewer.

Every model here ran under the same conditions. How the runs were set up →

From the audits

Most accurate, least decisive

All five clean-room audits miss the vocoder CPU wall. UNIONALPHA makes no wrong claims where Grok 4.6, DeepSeek V4.1 Flash and Muse Spark 1.3 Contributor make two or three each, and it covers the most issues, but it fully finds one decode-critical issue where Grok 4.6 finds three.

From the UNIONALPHA audit, 17 September 2026.

Weighted score

Out of 100, on the grade scale.

Where the points went

Earned and lost, by rubric slot

Each bar is a model’s 100-point grade split into the rubric’s weighted slots; filled is earned, empty is lost. The slots line up, so a gap in one bar against a full slot in another is exactly where the models part ways.

UNIONALPHA
74/ 100
Grok 4.6
70/ 100
Muse Spark 1.3 Contributor
68/ 100
DeepSeek V4.1 Flash
66/ 100
Qwen3.8 Flash
59/ 100
GLM 5.3 Flash
56/ 100
  • 1Accuracy & evidence30 pts
  • 2Coverage of decode problems25 pts
  • 3Root cause & prioritisation15 pts
  • 4Fix plan & acceptance gates15 pts
  • 5Originality & attribution10 pts
  • 6Clarity & calibration5 pts
Dimension profile

Scores on each rubric dimension

One panel per model: its score on each rubric dimension and the weighted total over the grade zones, in its own colour, with every other model on this task as a grey dot for scale. Hover a row for every model’s score and the auditor’s reasoning.

  • The panel’s model, joined down the dimensions
  • Every other model on this task

UNIONALPHA74 C

AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA

Grok 4.670 C

AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA

Muse Spark 1.3 Contributor68 C−

AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA

DeepSeek V4.1 Flash66 C−

AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA

Qwen3.8 Flash59 D

AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA

GLM 5.3 Flash56 D

AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA AccuracyCoverageRoot causeFix planOriginalityClarityTotalFDCBA
Scores and weighted points
Dimension Weight UNIONALPHAGrok 4.6Muse Spark 1.3 ContributorDeepSeek V4.1 FlashQwen3.8 FlashGLM 5.3 Flash
Accuracy & evidence 30% 90 · 27.0 pts 74 · 22.2 pts 74 · 22.2 pts 73 · 21.9 pts 62 · 18.6 pts 64 · 19.2 pts
Coverage of decode problems 25% 65 · 16.3 pts 59 · 14.8 pts 54 · 13.5 pts 49 · 12.3 pts 40 · 10.0 pts 34 · 8.5 pts
Root cause & prioritisation 15% 60 · 9.0 pts 70 · 10.5 pts 72 · 10.8 pts 66 · 9.9 pts 68 · 10.2 pts 62 · 9.3 pts
Fix plan & acceptance gates 15% 70 · 10.5 pts 75 · 11.3 pts 71 · 10.7 pts 70 · 10.5 pts 70 · 10.5 pts 62 · 9.3 pts
Originality & attribution 10% 76 · 7.6 pts 72 · 7.2 pts 70 · 7.0 pts 74 · 7.4 pts 64 · 6.4 pts 58 · 5.8 pts
Clarity & calibration 5% 80 · 4.0 pts 78 · 3.9 pts 80 · 4.0 pts 74 · 3.7 pts 70 · 3.5 pts 68 · 3.4 pts
Weighted total 100% 74 · C70 · C68 · C−66 · C−59 · D56 · D
Claim check

Every claim, checked

One cell per claim in each audit’s claim table, checked at the lines it cites. Hover a cell for the claim and the auditor’s note. The reviews made different numbers of claims, so the bar under each grid gives the shares.

UNIONALPHA’s claim table · Grok 4.6’s claim table · Muse Spark 1.3 Contributor’s claim table · DeepSeek V4.1 Flash’s claim table · Qwen3.8 Flash’s claim table · GLM 5.3 Flash’s claim table

  • Holds
  • Qualifiedoverstated, miscounted or doubtful
  • Wrong
UNIONALPHA35 claims checked

33 hold · 2 qualified · 0 wrong · 94% hold

Grok 4.630 claims checked

21 hold · 7 qualified · 2 wrong · 70% hold

Muse Spark 1.3 Contributor30 claims checked

22 hold · 6 qualified · 2 wrong · 73% hold

DeepSeek V4.1 Flash35 claims checked

28 hold · 4 qualified · 3 wrong · 80% hold

Qwen3.8 Flash33 claims checked

21 hold · 6 qualified · 6 wrong · 64% hold

GLM 5.3 Flash28 claims checked

17 hold · 7 qualified · 4 wrong · 61% hold

Issue coverage

What each review raised

UNIONALPHA found 1 of the 4 decode-critical issues; Grok 4.6 found 3 of the 4 decode-critical issues; Muse Spark 1.3 Contributor found 2 of the 4 decode-critical issues; DeepSeek V4.1 Flash found 2 of the 4 decode-critical issues; Qwen3.8 Flash found 2 of the 4 decode-critical issues; GLM 5.3 Flash found 1 of the 4 decode-critical issues. Every tracked issue below, with the reviews that were already in the repository for reference.

Issue signatures

Each review as a trace across the tracked issues. Where traces part is where the reviews differ; hover a column for every review’s entry. Column numbers match the table below.

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 Project docs earlier repo review · in the repository Grok 4.6 model · Sep 17 DeepSeek V4.1 Flash model · Sep 17 Muse Spark 1.3 Contri… model · Sep 17 GLM 5.3 Flash model · Sep 17 UNIONALPHA model · Sep 17 Qwen3.8 Flash model · Sep 17
  • raised
  • partly
  • absent
  • Decode-critical
Issues raised by each review of P25 voice decoding on the DM-1701
# Issue Project docsEarlier repo review · in the repository Grok 4.6Model · Sep 17 DeepSeek V4.1 FlashModel · Sep 17 Muse Spark 1.3 ContributorModel · Sep 17 GLM 5.3 FlashModel · Sep 17 UNIONALPHAModel · Sep 17 Qwen3.8 FlashModel · Sep 17
1 Capture export stages the wrong I2S halfwords absent raised§1, first code change absentits first experiment uses that capture absenttrusts the capture absentrelies on the capture raisedF3, fixed and verified raisedP1-2, tests miss the adapter
2 Capture parser certifies an incomplete stream absent absent absent absenttrusts the parser absentrelies on the parser absent absent
3 pcm_starve never increments absentdocumented as working absent absent absentrelies on it absent raisedF7 raisedP1-7
4 No static RAM margin partlymargins still to measure absent raisedM6, 0 bytes free partly“nearly full”, from the docs raisedF5, byte-exact map audit raisedF6, byte-exact partlycalls the reservation headroom
5 The I2S stream is most likely microphone audio Decode-critical partlyopen, leaning sceptical raised§1, 0x89 versus 0xC9 raisedB1, four-value test raisedP1-1 partlyopen, leans towards RF partlyblocking, but never says microphone raisedP1-4, 0xE0 mic bit
6 HR-C6000 de-emphasis on the capture path Decode-critical absent raised§2, bit 5 of 0x34 absentcalls it benign raisedP1-2, 0x34=0x3C raisedF2, closes the eye partlycited, called unmeasured raisedP1-3, 0x34=0x1C
7 AT1846S FM filters, low-frequency bit, 25 kHz Decode-critical partly“require characterization” raised§2, register level raisedB2, filter register raisedP1-2, 0x58 filters partly“voice filtering”; wrong bandwidth premise partlycited, called unmeasured raisedP1-3, DMR 0x58 probe
8 0x10=0x6E hybrid state; 0x36 dual role partlybring-up clock rules partlymisses 0x6E and the 0x36 clock gate partly“undocumented hybrid state” partlyquiet-chip registers raisedF4, 0x10=0x80 kills the clock partlyF1, 0x6E against 0x80 partly“hybrid I2S state”
9 Manual: I2S frame clock “must be 8KHz” Decode-criticalmissed absent raised§3 absentquotes the paragraph, not the rule absent absent raisedF4 absentquotes the formulas, not the rule
10 One-layer 4FSK test mode as a P25 tap partlystock BER-test block only raisedGate D raisedB5, exact recipe absentdismissed absent partlyworth a bounded test absentruled out at “9600 Bd”
11 ±10% health gate versus ±1% timing clamp absent absent raisedB3, impact overstated raisedP1-3 raisedF3 absent absent
12 Fail-closed muting at LDU cadence absent absentlate-entry mute only absent absentcalls it an asset absent absent partly“keep it”
13 Non-standard MFID mutes clear calls absent absent absent absent absent absent absent
14 Test waveform shares the receiver’s RRC filter partly“synthetic RRC/AWGN” caveat absent partlytested it, says not to fix partlysynthetic only, wants recordings absentwould extend that model raisedF5, unquantified absentwould extend that model
15 MCU runs at 72 MHz raised raisedin passing absent raisedP1-5 absent raisedF7 absent
16 Vocoder needs 11–16× the 72 MHz CPU Decode-criticalmissed partlydecode timing unmeasured absent“fine on a 1 ms tick” absent“vocoder question settled” partlyunmeasured; fix order backwards absent“in good shape” partlydeadlines “unproven” absent“not the problem”
17 Direct discriminator tap (M17 mod) raised raiseduncredited raisedpins, timer ADC, 48 kS/s raisedfallback, pin 9 raisedfallback raisedfallback raisedfallback
18 Phase 2 architecture and scope partlynot implemented partlymisplaces the AMBE+2 decoder partlyvoice via mbelib AMBE+2 raisedwith RF band limits partlysays mbelib has no AMBE+2 raisedmost accurate section partlyDMR and X2-TDMA parts
19 Two unverified SPI writes per decoded 20 ms frame absent absent absent absent absenttreats them as protection absent absent
20 Capture sessions lack epochs absent absent partlymeasured=0 only absent absent raisedF9, 8 kHz under a 24 kHz header absent
21 Ring and tick real-time budget partlydeadlines unproven partlycalls it fine absent partlyoverruns look like weak RF absent raisedF7, 1 ms is a minimum partlyregister stalls against the ring
22 Clock config 3 assumes 12,288 Hz; the codec formula gives 12,000 absent absent raisedB3, clock model absent partly“guessed semantics” absent raisedP1-6, for a different reason
23 Clock-config writes bypass the verified SPI writer partlySPI retry note absent raisedM1 absent absent raisedF9 absent
24 Stock squelch re-arms FM audio (0x10=0x80) during monitoring absentassumes it can’t re-arm absent absent absent partlynames squelch logic as a risk raisedF1, new absent“fixed” by forcing squelch open
25 Stale clear-call state releases a new call’s first frames absent absent absent absent absent raisedF8, probe absent

From the provenance table in the Qwen3.8 Flash audit (17 Sep 2026), the newest audit on this task. Shaded columns are reviews that were already in the repository.

Run facts

How each run went

Descriptive, not graded, and in grade order rather than ranked: one run per model. Run time is wall-clock, so it includes the provider’s speed, rate-limit waits and retries, and token counts vary with each model’s tokenizer.

Run facts for each compared model, in grade order. Descriptive only: not graded or ranked.
Model Run time Agent steps Tool calls Output tokens Route Launch
UNIONALPHAThird attempt; the first two stopped on provider rate limits 36 min 83 95 63K OpenRouter API one-shot
Grok 4.6 5 min 16 61 17K xAI subscription one-shot
Muse Spark 1.3 Contributor 5 min 27 38 19K OpenRouter API one-shot
DeepSeek V4.1 Flash 35 min 88 146 99K OpenRouter API one-shot
Qwen3.8 FlashSecond attempt; the first stopped on a provider rate limit 57 min 92 97 147K OpenRouter API one-shot
GLM 5.3 Flash 30 min 46 58 35K OpenRouter API one-shot
Every run figure
Model Started (UTC) Output tokens Reasoning tokens Context tokens Messages Sessions Retry budget
UNIONALPHA 17 Sep 2026, 12:57 63,278 not reported 9,988,183 179 1 10
Grok 4.6 17 Sep 2026, 04:25 16,928 9,183 1,300,749 78 1 3
Muse Spark 1.3 Contributor 17 Sep 2026, 06:10 19,449 7,815 2,309,483 66 1 3
DeepSeek V4.1 Flash 17 Sep 2026, 05:22 99,482 68,350 17,785,897 235 1 3
Qwen3.8 Flash 17 Sep 2026, 13:49 146,532 125,802 13,381,380 190 1 10
GLM 5.3 Flash 17 Sep 2026, 06:35 35,256 23,180 4,819,946 106 1 3

Context tokens are input plus cache reads and writes: mostly the conversation re-sent at every step, so they grow with steps × context size. The retry budget is how many API retries the harness allowed.