Model cards · No. 3 of 17
Qwen3-Coder-Next
UD-Q4_K_XLContenderContender
91.1Local AI Index / 100
Fast and steady. Failed the logic puzzle the same way in all three runs.
The numbers
Measured, not quoted.
Qwen3-Coder-Next (UD-Q4_K_XL) ranks #3 of 17 on the diego.technology Local AI Index with 91.1/100: 18.63/20 on twenty graded tasks (mean of 3 runs), 8.8/10 on the hard half, at 46.8 tokens/s on a 128 GB AMD Strix Halo machine.
- mean of 3 runs
- 18.63/ 20
- mean of 3 runs
- on the hard half
- 8.8/ 10
- on the hard half
- generation · 265 prompt
- 46.8tok/s
- generation · 265 prompt
- to load
- 18s
- to load
Where the 91.1 points came from
- Capability46.6 / 50
- Hard problems17.6 / 20
- Speed14.8 / 15
- Consistency12.2 / 15
Strengths
Where it wins, where it slips.
Twenty graded tasks in six families. Brighter is better. Hatched means a run ran out of output budget before answering — unfinished, not wrong.
- Code93%
- Math100%
- Data100%
- Tools100%
- Context100%
- Reasoning68%
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- 10
- 11
- 12
- 13
- 14
- 15
- 16
- 17
- 18
- 19
- 20
Tasks 11–20 are the hard half. Hover a cell for the task.
Every task, as a list
- 01 Python algorithm (interval merge)100%
- 02 Python bug fix (date arithmetic)100%
- 03 Compound-interest math100%
- 04 Tiered tax calculation100%
- 05 JavaScript ratio stress test100%
- 06 SQL: best value per group100%
- 07 Single tool call100%
- 08 Bilingual email → JSON100%
- 09 Long context (~18k tokens)100%
- 10 Predict program output83%
- 11 Payment schedule with prepayments100%
- 12 Two-formula penalty comparison100%
- 13 Expression parser, no eval()69%
- 14 JS async retry with backoff100%
- 15 Exact set cover100%
- 16 SQL window functions100%
- 17 Multi-file bug hunt90%
- 18 Multi-hop long context100%
- 19 Multi-step tool use100%
- 20 Logic puzzle20%
As a coding agent
Six real repo tasks, four harnesses.
Hidden tests decide the score. Minutes are wall-clock for all six tasks.
- Pi
- 5.83/ 6
- 4 min
- OpenCode
- —/ 6
- Not re-run after the permissions fix
- dsh
- 5.83/ 6
- 6 min
- Hermes
- 5.83/ 6
- 8 min
The record
3 full runs.
Repeat runs are what separate a daily driver from a lucky score.
- 2026-09-2618.56 / 20 · hard 8.56 / 10
- 2026-09-2819.04 / 20 · hard 9.04 / 10
- 2026-09-2918.29 / 20 · hard 8.79 / 10