Model cards · No. 17 of 17
Mistral Small 4 119B
UD-Q4_K_XLRetiredRetired
61.7Local AI Index / 100
Finished every answer — and got many of them partly wrong.
The numbers
Measured, not quoted.
Mistral Small 4 119B (UD-Q4_K_XL) ranks #17 of 17 on the diego.technology Local AI Index with 61.7/100: 11.42/20 on twenty graded tasks (one run), 5.17/10 on the hard half, at 36.1 tokens/s on a 128 GB AMD Strix Halo machine.
- one run so far
- 11.42/ 20
- one run so far
- on the hard half
- 5.17/ 10
- on the hard half
- generation · 163 prompt
- 36.1tok/s
- generation · 163 prompt
- to load
- 30s
- to load
Where the 61.7 points came from
- Capability28.5 / 50
- Hard problems10.3 / 20
- Speed13.8 / 15
- Consistency9.0 / 15
Strengths
Where it wins, where it slips.
Twenty graded tasks in six families. Brighter is better. Hatched means a run ran out of output budget before answering — unfinished, not wrong.
- Code67%
- Math63%
- Data50%
- Tools100%
- Context25%
- Reasoning30%
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- 10
- 11
- 12
- 13
- 14
- 15
- 16
- 17
- 18
- 19
- 20
Tasks 11–20 are the hard half. Hover a cell for the task.
Every task, as a list
- 01 Python algorithm (interval merge)100%
- 02 Python bug fix (date arithmetic)38%
- 03 Compound-interest math0%
- 04 Tiered tax calculation100%
- 05 JavaScript ratio stress test38%
- 06 SQL: best value per group50%
- 07 Single tool call100%
- 08 Bilingual email → JSON100%
- 09 Long context (~18k tokens)50%
- 10 Predict program output50%
- 11 Payment schedule with prepayments50%
- 12 Two-formula penalty comparison100%
- 13 Expression parser, no eval()56%
- 14 JS async retry with backoff100%
- 15 Exact set cover40%
- 16 SQL window functions0%
- 17 Multi-file bug hunt71%
- 18 Multi-hop long context0%
- 19 Multi-step tool use100%
- 20 Logic puzzle0%
The record
One full run.
A second run is what moves a setup from provisional to verified.
- 2026-09-2711.42 / 20 · hard 5.17 / 10