Model cards · No. 3 of 17

Qwen3-Coder-Next

UD-Q4_K_XLContenderContender

91.1Local AI Index / 100

Fast and steady. Failed the logic puzzle the same way in all three runs.

The numbers

Measured, not quoted.

Qwen3-Coder-Next (UD-Q4_K_XL) ranks #3 of 17 on the diego.technology Local AI Index with 91.1/100: 18.63/20 on twenty graded tasks (mean of 3 runs), 8.8/10 on the hard half, at 46.8 tokens/s on a 128 GB AMD Strix Halo machine.

mean of 3 runs
18.63/ 20
mean of 3 runs
on the hard half
8.8/ 10
on the hard half
generation · 265 prompt
46.8tok/s
generation · 265 prompt
to load
18s
to load

Where the 91.1 points came from

  • Capability46.6 / 50
  • Hard problems17.6 / 20
  • Speed14.8 / 15
  • Consistency12.2 / 15

Strengths

Where it wins, where it slips.

Twenty graded tasks in six families. Brighter is better. Hatched means a run ran out of output budget before answering — unfinished, not wrong.

  • Code93%
  • Math100%
  • Data100%
  • Tools100%
  • Context100%
  • Reasoning68%
  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14
  15. 15
  16. 16
  17. 17
  18. 18
  19. 19
  20. 20

Tasks 11–20 are the hard half. Hover a cell for the task.

Every task, as a list
  • 01 Python algorithm (interval merge)100%
  • 02 Python bug fix (date arithmetic)100%
  • 03 Compound-interest math100%
  • 04 Tiered tax calculation100%
  • 05 JavaScript ratio stress test100%
  • 06 SQL: best value per group100%
  • 07 Single tool call100%
  • 08 Bilingual email → JSON100%
  • 09 Long context (~18k tokens)100%
  • 10 Predict program output83%
  • 11 Payment schedule with prepayments100%
  • 12 Two-formula penalty comparison100%
  • 13 Expression parser, no eval()69%
  • 14 JS async retry with backoff100%
  • 15 Exact set cover100%
  • 16 SQL window functions100%
  • 17 Multi-file bug hunt90%
  • 18 Multi-hop long context100%
  • 19 Multi-step tool use100%
  • 20 Logic puzzle20%

As a coding agent

Six real repo tasks, four harnesses.

Hidden tests decide the score. Minutes are wall-clock for all six tasks.

Pi
5.83/ 6
4 min
OpenCode
—/ 6
Not re-run after the permissions fix
dsh
5.83/ 6
6 min
Hermes
5.83/ 6
8 min

The record

3 full runs.

Repeat runs are what separate a daily driver from a lucky score.

  1. 2026-09-2618.56 / 20 · hard 8.56 / 10
  2. 2026-09-2819.04 / 20 · hard 9.04 / 10
  3. 2026-09-2918.29 / 20 · hard 8.79 / 10