Right about the chip, wrong about the CPU.
- Holds upIt puts the chip path at the centre, and it’s right to
- ValidatedIts de-emphasis finding is decisive
- Critical errorIt clears the vocoder that cannot run in real time
Every model gets the same brief on a real codebase. Each claim in its review is checked against the code, its evidence is rerun, and what it missed is measured. One published rubric scores them all, so the grades compare directly.
The brief“Find why this firmware struggles to decode P25 Phase 1 and Phase 2 voice on the radio’s HR-C6000, and what to do about it.”
DM-1701 · HR-C6000 · AT1846S · STM32F405 · P25
Each bar sits on the grade scale: F below 50, then D, C and B, and A from 90. Hover or focus a model for its dimension scores; select it for the full scorecard.
| # | Model | Grade | Score | Claims that hold | Wrong claims | Decode-critical found | Graded | Read |
|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek | C− | 68 | 17 / 30 | 5 | 3 of 4 | 17 Sep 2026 | ScorecardFull review |
| 2 | UNIONALPHAStealth model | C− | 65 | 21 / 23 | 0 | 0 of 4 | 16 Sep 2026 | ScorecardFull review |
UNIONALPHA is the more careful review: almost every claim holds, but it spends them on tooling. DeepSeek finds the chip-side problems UNIONALPHA missed and proposes the best experiment either offers, but makes more mistakes, and one of them clears a real blocker. Neither found the vocoder CPU wall or the circular RRC test oracle.
From the DeepSeek V4.1 Flash audit, 17 September 2026.
Three headline findings from each audit. The scorecard has all of them; the full review has the evidence behind every one.
Right about the chip, wrong about the CPU.
Accurate to the line number. Silent on what stops the radio decoding.
Each rubric dimension is scored out of 100 and weighted into the grade. Same rubric and weights for every model on this task.
| Dimension | Weight | A · DeepSeek V4.1 Flash | B · UNIONALPHA | A − B |
|---|---|---|---|---|
| Accuracy & evidence | 30% | 68 · 20.4 pts | 90 · 27.0 pts | −22 |
| Coverage of decode problems | 25% | 66 · 16.5 pts | 45 · 11.2 pts | +21 |
| Root cause & prioritisation | 15% | 64 · 9.6 pts | 55 · 8.2 pts | +9 |
| Fix plan & acceptance gates | 15% | 74 · 11.1 pts | 74 · 11.1 pts | ±0 |
| Originality & attribution | 10% | 70 · 7.0 pts | 40 · 4.0 pts | +30 |
| Clarity & calibration | 5% | 68 · 3.4 pts | 76 · 3.8 pts | −8 |
| Weighted total | 100% | 68 · C− | 65 · C− | +3 |