Model review audits
Model review audits

AI models, graded on real engineering work.

Every model gets the same brief on a real codebase. Each claim in its review is checked against the code, its evidence is rerun, and what it missed is measured. One published rubric scores them all, so the grades compare directly.

Models graded
2
1 task, one rubric
Claims checked
53
each at the lines it cites
Top score
68C−
Latest audit
17 Sep 2026
Leaderboard · Firmware & RF

P25 voice decoding on the DM-1701

The brief“Find why this firmware struggles to decode P25 Phase 1 and Phase 2 voice on the radio’s HR-C6000, and what to do about it.”

DM-1701 · HR-C6000 · AT1846S · STM32F405 · P25

Weighted score out of 100

Each bar sits on the grade scale: F below 50, then D, C and B, and A from 90. Hover or focus a model for its dimension scores; select it for the full scorecard.

Leaderboard for P25 voice decoding on the DM-1701
# Model Grade Score Claims that hold Wrong claims Decode-critical found Graded Read
1 DeepSeek V4.1 FlashDeepSeek C− 68 17 / 30 5 3 of 4 17 Sep 2026
2 UNIONALPHAStealth model C− 65 21 / 23 0 0 of 4 16 Sep 2026
  • DeepSeek V4.1 Flash. Same brief, with UNIONALPHA’s findings handed over as a baseline, which the review acknowledges.
  • UNIONALPHA. Blind. The model’s identity was withheld from the auditor; the repository’s earlier reviews were in the working tree.
Head to head

More useful, less careful

UNIONALPHA is the more careful review: almost every claim holds, but it spends them on tooling. DeepSeek finds the chip-side problems UNIONALPHA missed and proposes the best experiment either offers, but makes more mistakes, and one of them clears a real blocker. Neither found the vocoder CPU wall or the circular RRC test oracle.

From the DeepSeek V4.1 Flash audit, 17 September 2026.

Headline findings

What each review got right, and what it missed

Three headline findings from each audit. The scorecard has all of them; the full review has the evidence behind every one.

Rank 2 of 2

UNIONALPHA

Stealth model
C−65 / 100

Accurate to the line number. Silent on what stops the radio decoding.

  • Holds upWhat it checked is correct
  • Critical missIt never mentions the analog-FM receive chain
  • Critical missThe vocoder cannot run in real time, and it never says so
By dimension

Where the points came from

Each rubric dimension is scored out of 100 and weighted into the grade. Same rubric and weights for every model on this task.

  • ADeepSeek V4.1 Flash
  • BUNIONALPHA
Accuracy & evidence30%
Coverage of decode problems25%
Root cause & prioritisation15%
Fix plan & acceptance gates15%
Originality & attribution10%
Clarity & calibration5%
Scores and weighted points
Dimension Weight A · DeepSeek V4.1 FlashB · UNIONALPHA A − B
Accuracy & evidence 30% 68 · 20.4 pts 90 · 27.0 pts −22
Coverage of decode problems 25% 66 · 16.5 pts 45 · 11.2 pts +21
Root cause & prioritisation 15% 64 · 9.6 pts 55 · 8.2 pts +9
Fix plan & acceptance gates 15% 74 · 11.1 pts 74 · 11.1 pts ±0
Originality & attribution 10% 70 · 7.0 pts 40 · 4.0 pts +30
Clarity & calibration 5% 68 · 3.4 pts 76 · 3.8 pts −8
Weighted total 100% 68 · C−65 · C− +3
Method

How a grade is made

Weights reflect the question asked: find what stops P25 decoding and say how to fix it. Accuracy carries the most weight because a wrong review does harm; coverage and root cause together outweigh it because an accurate review of the wrong things does not help.

The process, the rubric and the limits 

Rubric weights

Each dimension’s share of the 100-point grade.

  • 30%Accuracy & evidence
  • 25%Coverage of decode problems
  • 15%Root cause & prioritisation
  • 15%Fix plan & acceptance gates
  • 10%Originality & attribution
  • 5%Clarity & calibration