Local AI Index · v1.0 · October 2, 2026

What’s worth running at home?

One machine. One set of graded tests. One score out of 100 for every open model I can run — measured, never copied from someone else’s leaderboard.

No. 1 right nowgpt-oss-120bMXFP4 · 48.5 tokens/s98.4

The ranking

17 setups, scored.

Each bar is built from four parts. Longer is better; the colours show where the points came from.

  1. 01gpt-oss-120bMXFP4 · 3 runs98.4
  2. 02Gemma 4 31B QATUD-Q4_K_XL · 3 runs94.4
  3. 03Qwen3-Coder-NextUD-Q4_K_XL · 3 runs91.1
  4. 04GLM-5.3-FlashUD-IQ3_XXS · 1 run89.2
  5. 05KAT-Coder v2.5Q8_0 · 1 run83.0
  6. 06Qwen3-Coder-NextQ6 · 1 run83.0
  7. 07Muse Glimmer 30BQ8_0 · 1 run82.7
  8. 08Tiel-Coder 35BQ8_0 · 1 run81.4
  9. 09Nemotron 3 Super 120BUD-Q4_K_XL · 3 runs81.0
  10. 10DeepSeek V4-FlashUD-IQ3_XXS · 1 run79.7
  11. 11Qwen3.8-Flash-NextIQ4_XS-PLE · fixed sampling · 1 run78.4
  12. 12Qwen3.8-Flash-NextUD-IQ4_XS · fixed sampling · 1 run77.7
  13. 13Gemma 4 26B-A4B QATUD-Q4_K_XL · 1 run71.0
  14. 14Devstral 2 24BQ8_0 · 1 run70.0
  15. 15Qwen3.8-27BQ8_0 · 1 run · settings under audit66.0
  16. 16Gemma 4 12B QATUD-Q4_K_XL · 1 run62.0
  17. 17Mistral Small 4 119BUD-Q4_K_XL · 1 run61.7
Score breakdown table
ModelCapabilityHard problemsSpeedConsistencyIndex
gpt-oss-120b MXFP449.719.814.914.198.4
Gemma 4 31B QAT UD-Q4_K_XL50.020.09.415.094.4
Qwen3-Coder-Next UD-Q4_K_XL46.617.614.812.291.1
GLM-5.3-Flash UD-IQ3_XXS50.020.010.29.089.2
KAT-Coder v2.5 Q8_043.915.115.09.083.0
Qwen3-Coder-Next Q643.416.314.39.083.0
Muse Glimmer 30B Q8_047.518.08.29.082.7
Tiel-Coder 35B Q8_042.315.614.59.081.4
Nemotron 3 Super 120B UD-Q4_K_XL45.816.711.07.581.0
DeepSeek V4-Flash UD-IQ3_XXS45.016.09.79.079.7
Qwen3.8-Flash-Next IQ4_XS-PLE · fixed sampling42.514.012.99.078.4
Qwen3.8-Flash-Next UD-IQ4_XS · fixed sampling42.514.012.29.077.7
Gemma 4 26B-A4B QAT UD-Q4_K_XL35.012.015.09.071.0
Devstral 2 24B Q8_036.915.58.69.070.0
Qwen3.8-27B Q8_036.912.08.29.066.0
Gemma 4 12B QAT UD-Q4_K_XL32.58.012.59.062.0
Mistral Small 4 119B UD-Q4_K_XL28.510.313.89.061.7

How the score works

Four ingredients. Nothing hidden.

The weights favour getting the answer right, then getting it right on hard problems, then not making you wait, then doing it again.

50%

Capability

Mean score on all twenty graded scenarios.

20%

Hard problems

Mean score on the ten hard scenarios — counted twice on purpose, because that is where models differ.

15%

Speed

Generation speed on log scale; 50 tokens/s or more earns full marks. Waiting less matters, with diminishing returns.

15%

Consistency

Run-to-run spread across three runs: a 2-point swing costs half the credit. A model with only one run gets 60% until it proves it can repeat.

Why not average public leaderboards?

They test full-precision models on data-centre hardware, rescale between versions, and rarely include the quantized builds people actually run at home. A number you can’t reproduce on your own machine isn’t much use for choosing one.

What it doesn’t measure — yet

Everyday assistant work, memory, safety and multi-agent delegation. Coding-agent results live on the lab page until enough models have them. The weights are version 1.0; any change will bump the version.

Measured on

The test bench.

Every number above came from this one machine — the AMD Strix Halo platform, where the GPU can use most of 128 GB of fast unified memory.

GMKtec EVO-X2 AI · AMD Ryzen AI Max+ 395 (Strix Halo)

128 GB
unified LPDDR5X-8000
256 GB/s
peak memory bandwidth
96 GB
addressable by the GPU
2 TB
PCIe 4.0 NVMe

Processor

Chip
AMD Ryzen AI Max+ 395 (Strix Halo)
CPU
16 Zen 5 cores · 32 threads
Clocks
3.0 GHz base · up to 5.1 GHz boost
Cache
16 MB L2 · 64 MB L3
Process
TSMC 4 nm
Power
Up to 140 W (GMKtec rating)AMD lists a 45–120 W configurable range; GMKtec rates this chassis at up to 140 W.

Graphics and AI

GPU
Radeon 8060S · 40 RDNA 3.5 compute units
GPU clock
Up to 2.9 GHz
Architecture ID
gfx1151
NPU
XDNA 2 · up to 50 TOPS
Total AI
Up to 126 TOPS (CPU + GPU + NPU)The NPU is not used by any result on this site; inference runs on the GPU.

Memory

Capacity
128 GB unified, soldered
Type
LPDDR5X-8000
Bus
256-bit
Bandwidth
256 GB/s theoretical peak8000 MT/s × 256 bits ÷ 8. Token generation on large models is mostly limited by this number.
GPU share
2 GB fixed carve-out + up to 96 GB shared (GTT)96 GB is the Linux default of three quarters of system memory.

Storage and I/O

Storage
2 TB M.2 PCIe 4.0 NVMe
USB
2 × USB4 (40 Gb/s) · 3 × USB 3.2 Gen 2 · 2 × USB 2.0
Display
HDMI 2.1 · DisplayPort 1.4
Network
2.5 GbE · Wi-Fi 7 · Bluetooth 5.4
Power supply
230 W

Software (September 2026 runs)

OS
Ubuntu 24.04.4 LTS, OEM kernel 7.0 · headless
Graphics stack
In-kernel amdgpu · Mesa 25.2.8 (RADV Vulkan)
Inference
llama.cpp b11238 (Vulkan) · ROCm for GLM-5.3-Flash
Also installed
Ollama 0.32
Agent harnesses
Pi 0.87.1 · OpenCode 1.18.33 · dsh 0.2.0-rc.1 · Hermes 0.21.5
Firmware
BIOS 1.12

On the radar

Not scored yet.

New open models are scouted every week from Hugging Face, r/LocalLLaMA, YouTube and public leaderboards. A model gets a score only after it runs here.

Downloaded · testing next · 4

  • Integrity-checked download complete. Compatibility smoke test next.

  • Integrity-checked download complete. Compatibility smoke test next.

  • A small distill of the current top open-weight family. Smoke test next.

  • Integrity-checked download complete. Compatibility smoke test next.

Shortlisted · 4

  • Ornith-1.5-35B-A3B36B total · 3B active

    Agentic-coding MoE with publisher-reported SWE-bench Verified of 79%. Small active size should be fast here.

  • Xing4.0-29B-A4B29B total · 4B active

    Agent-oriented MoE, Apache 2.0, 256K context; publisher reports 75 on SWE-bench Verified.

  • Successor to V4-Flash, which scored 18/20 here at IQ3. Needs a quant that fits 96 GB of GPU memory.

  • MiMo-V2.6-Flash309B total · 15B active

    Strong agent scores, MIT. Only fits at roughly 2–2.5 bits per weight — a stress test of aggressive quantization.

Blocked on this hardware · 1

  • Ternary Bonsai 2 27B27B at ~1.75 bits · 6 GB

    A ternary Qwen3.8-27B that claims 98% of full quality. Needs a custom llama.cpp fork with CUDA, Metal or CPU kernels — no AMD GPU path yet.

Too big for 128 GB · 3

  • MiMo-V2.6-Pro1.02T total · 42B active

    Top of the public open-weight rankings, but about four times too big for 128 GB at any usable quant.

  • Kimi K32.8T total

    Far beyond a single 128 GB machine.

  • GLM-5.3~750B total (reported)

    The full model; its Flash sibling already scored 20/20 here.

Sizes and benchmark claims on the radar are the publishers’ own, from the linked model cards. They are context, not part of the index.