Local AI Index · v1.0 · October 2, 2026
What’s worth running at home?
One machine. One set of graded tests. One score out of 100 for every open model I can run — measured, never copied from someone else’s leaderboard.
No. 1 right nowgpt-oss-120bMXFP4 · 48.5 tokens/s98.4The ranking
17 setups, scored.
Each bar is built from four parts. Longer is better; the colours show where the points came from.
- 01gpt-oss-120bMXFP4 · 3 runs98.4
- 02Gemma 4 31B QATUD-Q4_K_XL · 3 runs94.4
- 03Qwen3-Coder-NextUD-Q4_K_XL · 3 runs91.1
- 04GLM-5.3-FlashUD-IQ3_XXS · 1 run89.2
- 05KAT-Coder v2.5Q8_0 · 1 run83.0
- 06Qwen3-Coder-NextQ6 · 1 run83.0
- 07Muse Glimmer 30BQ8_0 · 1 run82.7
- 08Tiel-Coder 35BQ8_0 · 1 run81.4
- 09Nemotron 3 Super 120BUD-Q4_K_XL · 3 runs81.0
- 10DeepSeek V4-FlashUD-IQ3_XXS · 1 run79.7
- 11Qwen3.8-Flash-NextIQ4_XS-PLE · fixed sampling · 1 run78.4
- 12Qwen3.8-Flash-NextUD-IQ4_XS · fixed sampling · 1 run77.7
- 13Gemma 4 26B-A4B QATUD-Q4_K_XL · 1 run71.0
- 14Devstral 2 24BQ8_0 · 1 run70.0
- 15Qwen3.8-27BQ8_0 · 1 run · settings under audit66.0
- 16Gemma 4 12B QATUD-Q4_K_XL · 1 run62.0
- 17Mistral Small 4 119BUD-Q4_K_XL · 1 run61.7
Score breakdown table
| Model | Capability | Hard problems | Speed | Consistency | Index |
|---|---|---|---|---|---|
| gpt-oss-120b MXFP4 | 49.7 | 19.8 | 14.9 | 14.1 | 98.4 |
| Gemma 4 31B QAT UD-Q4_K_XL | 50.0 | 20.0 | 9.4 | 15.0 | 94.4 |
| Qwen3-Coder-Next UD-Q4_K_XL | 46.6 | 17.6 | 14.8 | 12.2 | 91.1 |
| GLM-5.3-Flash UD-IQ3_XXS | 50.0 | 20.0 | 10.2 | 9.0 | 89.2 |
| KAT-Coder v2.5 Q8_0 | 43.9 | 15.1 | 15.0 | 9.0 | 83.0 |
| Qwen3-Coder-Next Q6 | 43.4 | 16.3 | 14.3 | 9.0 | 83.0 |
| Muse Glimmer 30B Q8_0 | 47.5 | 18.0 | 8.2 | 9.0 | 82.7 |
| Tiel-Coder 35B Q8_0 | 42.3 | 15.6 | 14.5 | 9.0 | 81.4 |
| Nemotron 3 Super 120B UD-Q4_K_XL | 45.8 | 16.7 | 11.0 | 7.5 | 81.0 |
| DeepSeek V4-Flash UD-IQ3_XXS | 45.0 | 16.0 | 9.7 | 9.0 | 79.7 |
| Qwen3.8-Flash-Next IQ4_XS-PLE · fixed sampling | 42.5 | 14.0 | 12.9 | 9.0 | 78.4 |
| Qwen3.8-Flash-Next UD-IQ4_XS · fixed sampling | 42.5 | 14.0 | 12.2 | 9.0 | 77.7 |
| Gemma 4 26B-A4B QAT UD-Q4_K_XL | 35.0 | 12.0 | 15.0 | 9.0 | 71.0 |
| Devstral 2 24B Q8_0 | 36.9 | 15.5 | 8.6 | 9.0 | 70.0 |
| Qwen3.8-27B Q8_0 | 36.9 | 12.0 | 8.2 | 9.0 | 66.0 |
| Gemma 4 12B QAT UD-Q4_K_XL | 32.5 | 8.0 | 12.5 | 9.0 | 62.0 |
| Mistral Small 4 119B UD-Q4_K_XL | 28.5 | 10.3 | 13.8 | 9.0 | 61.7 |
How the score works
Four ingredients. Nothing hidden.
The weights favour getting the answer right, then getting it right on hard problems, then not making you wait, then doing it again.
50%
Capability
Mean score on all twenty graded scenarios.
20%
Hard problems
Mean score on the ten hard scenarios — counted twice on purpose, because that is where models differ.
15%
Speed
Generation speed on log scale; 50 tokens/s or more earns full marks. Waiting less matters, with diminishing returns.
15%
Consistency
Run-to-run spread across three runs: a 2-point swing costs half the credit. A model with only one run gets 60% until it proves it can repeat.
Why not average public leaderboards?
They test full-precision models on data-centre hardware, rescale between versions, and rarely include the quantized builds people actually run at home. A number you can’t reproduce on your own machine isn’t much use for choosing one.
What it doesn’t measure — yet
Everyday assistant work, memory, safety and multi-agent delegation. Coding-agent results live on the lab page until enough models have them. The weights are version 1.0; any change will bump the version.
Measured on
The test bench.
Every number above came from this one machine — the AMD Strix Halo platform, where the GPU can use most of 128 GB of fast unified memory.
GMKtec EVO-X2 AI · AMD Ryzen AI Max+ 395 (Strix Halo)
- 128 GB
- unified LPDDR5X-8000
- 256 GB/s
- peak memory bandwidth
- 96 GB
- addressable by the GPU
- 2 TB
- PCIe 4.0 NVMe
Processor
- Chip
- AMD Ryzen AI Max+ 395 (Strix Halo)
- CPU
- 16 Zen 5 cores · 32 threads
- Clocks
- 3.0 GHz base · up to 5.1 GHz boost
- Cache
- 16 MB L2 · 64 MB L3
- Process
- TSMC 4 nm
- Power
- Up to 140 W (GMKtec rating)AMD lists a 45–120 W configurable range; GMKtec rates this chassis at up to 140 W.
Graphics and AI
- GPU
- Radeon 8060S · 40 RDNA 3.5 compute units
- GPU clock
- Up to 2.9 GHz
- Architecture ID
- gfx1151
- NPU
- XDNA 2 · up to 50 TOPS
- Total AI
- Up to 126 TOPS (CPU + GPU + NPU)The NPU is not used by any result on this site; inference runs on the GPU.
Memory
- Capacity
- 128 GB unified, soldered
- Type
- LPDDR5X-8000
- Bus
- 256-bit
- Bandwidth
- 256 GB/s theoretical peak8000 MT/s × 256 bits ÷ 8. Token generation on large models is mostly limited by this number.
- GPU share
- 2 GB fixed carve-out + up to 96 GB shared (GTT)96 GB is the Linux default of three quarters of system memory.
Storage and I/O
- Storage
- 2 TB M.2 PCIe 4.0 NVMe
- USB
- 2 × USB4 (40 Gb/s) · 3 × USB 3.2 Gen 2 · 2 × USB 2.0
- Display
- HDMI 2.1 · DisplayPort 1.4
- Network
- 2.5 GbE · Wi-Fi 7 · Bluetooth 5.4
- Power supply
- 230 W
Software (September 2026 runs)
- OS
- Ubuntu 24.04.4 LTS, OEM kernel 7.0 · headless
- Graphics stack
- In-kernel amdgpu · Mesa 25.2.8 (RADV Vulkan)
- Inference
- llama.cpp b11238 (Vulkan) · ROCm for GLM-5.3-Flash
- Also installed
- Ollama 0.32
- Agent harnesses
- Pi 0.87.1 · OpenCode 1.18.33 · dsh 0.2.0-rc.1 · Hermes 0.21.5
- Firmware
- BIOS 1.12
On the radar
Not scored yet.
New open models are scouted every week from Hugging Face, r/LocalLLaMA, YouTube and public leaderboards. A model gets a score only after it runs here.
Downloaded · testing next · 4
Integrity-checked download complete. Compatibility smoke test next.
Integrity-checked download complete. Compatibility smoke test next.
- MiMo-V2.6-Distill-Qwen-9B9B (per name)
A small distill of the current top open-weight family. Smoke test next.
Integrity-checked download complete. Compatibility smoke test next.
Shortlisted · 4
- Ornith-1.5-35B-A3B36B total · 3B active
Agentic-coding MoE with publisher-reported SWE-bench Verified of 79%. Small active size should be fast here.
- Xing4.0-29B-A4B29B total · 4B active
Agent-oriented MoE, Apache 2.0, 256K context; publisher reports 75 on SWE-bench Verified.
- DeepSeek-V4.1-FlashCheck fit
Successor to V4-Flash, which scored 18/20 here at IQ3. Needs a quant that fits 96 GB of GPU memory.
- MiMo-V2.6-Flash309B total · 15B active
Strong agent scores, MIT. Only fits at roughly 2–2.5 bits per weight — a stress test of aggressive quantization.
Blocked on this hardware · 1
- Ternary Bonsai 2 27B27B at ~1.75 bits · 6 GB
A ternary Qwen3.8-27B that claims 98% of full quality. Needs a custom llama.cpp fork with CUDA, Metal or CPU kernels — no AMD GPU path yet.
Too big for 128 GB · 3
- MiMo-V2.6-Pro1.02T total · 42B active
Top of the public open-weight rankings, but about four times too big for 128 GB at any usable quant.
- Kimi K32.8T total
Far beyond a single 128 GB machine.
- GLM-5.3~750B total (reported)
The full model; its Flash sibling already scored 20/20 here.
Sizes and benchmark claims on the radar are the publishers’ own, from the linked model cards. They are context, not part of the index.