Strix Halo Bench · v0.1 · October 3, 2026
Every open model that fits this machine, on one table.
A benchmark-dashboard view of what a 128 GB home machine can actually run: index scores, category splits, agent harnesses and the Pareto frontier of quality against speed.
- Setups scored
- 17
- Graded scenarios
- 20
- Agent harnesses
- 4
- In the pipeline
- 12
Current answer
gpt-oss-120b (MXFP4) is the setup to run today — 19.88/20 at 48.5 tok/s, index 98.4. The perfect scorer Gemma 4 31B takes the planner seat when minutes don't matter. Four downloaded candidates are next in the queue; they get a row here only after they run here.
The leaderboard
Sorted, filtered, evidenced.
One row per configuration. The index bar shows where its points come from; the evidence column says how many independent runs back the number. Click a row for category strengths and its best agent run.
| # | Model | Index | /20 | Hard | tok/s | Load | Evidence | |
|---|---|---|---|---|---|---|---|---|
| 01 | gpt-oss-120bMXFP4 | Current team | 98.4 | 19.88 | 9.88/10 | 48.5 | 24s | Verified |
| 02 | Gemma 4 31B QATUD-Q4_K_XL | Current team | 94.4 | 20 | 10/10 | 10.6 | 9s | Verified |
| 03 | Qwen3-Coder-NextUD-Q4_K_XL | Contender | 91.1 | 18.63 | 8.8/10 | 46.8 | 18s | Verified |
| 04 | GLM-5.3-FlashUD-IQ3_XXS | Contender | 89.2 | 20 | 10/10 | 13.4 | 45s | Provisional |
| 05 | KAT-Coder v2.5Q8_0 | Retired | 83.0 | 17.55 | 7.55/10 | 51.4 | 15s | Provisional |
| 06 | Qwen3-Coder-NextQ6 | Retired | 83.0 | 17.36 | 8.16/10 | 41 | 27s | Provisional |
| 07 | Muse Glimmer 30BQ8_0 | Contender | 82.7 | 19 | 9/10 | 7.6 | 12s | Provisional |
| 08 | Tiel-Coder 35BQ8_0 | Retired | 81.4 | 16.92 | 7.8/10 | 43.6 | 15s | Provisional |
| 09 | Nemotron 3 Super 120BUD-Q4_K_XL | Contender | 81.0 | 18.33 | 8.33/10 | 16.8 | 30s | Verified |
| 10 | DeepSeek V4-FlashUD-IQ3_XXS | Contender | 79.7 | 18 | 8/10 | 11.8 | 36s | Provisional |
| 11 | Qwen3.8-Flash-NextIQ4_XS-PLE · fixed sampling | Retired | 78.4 | 17 | 7/10 | 28.2 | 33s | Provisional |
| 12 | Qwen3.8-Flash-NextUD-IQ4_XS · fixed sampling | Retired | 77.7 | 17 | 7/10 | 23.6 | 36s | Provisional |
| 13 | Gemma 4 26B-A4B QATUD-Q4_K_XL | Retired | 71.0 | 14 | 6/10 | 57.5 | 6s | Provisional |
| 14 | Devstral 2 24BQ8_0 | Retired | 70.0 | 14.76 | 7.76/10 | 8.5 | 12s | Provisional |
| 15 | Qwen3.8-27BQ8_0 | Under audit | 66.0 | 14.75 | 6/10 | 7.5 | 12s | Under audit |
| 16 | Gemma 4 12B QATUD-Q4_K_XL | Retired | 62.0 | 13 | 4/10 | 25.5 | 6s | Provisional |
| 17 | Mistral Small 4 119BUD-Q4_K_XL | Retired | 61.7 | 11.42 | 5.17/10 | 36.1 | 30s | Provisional |
Index = capability 50 · hard half 20 · speed 15 · consistency 15 (same weights as the AI Index). Evidence reflects independent full runs; provisional numbers move.
Quality against speed
The frontier nobody dominates.
Log-scale speed on the x-axis, index on the y. The dashed line is the Pareto frontier: every setup on it is the best there is at that speed.
Decision-ready
Four picks, each with its rule.
benchlm-style: if you won't read the table, read these. Every pick states how it was chosen, and carries the same evidence tier as its row.
Best overall
MXFP4
Highest index — accuracy on all twenty scenarios, hard half, speed and run-to-run consistency.
Verified — Two or more full runs
The perfect scorer
UD-Q4_K_XL
Full marks on every scenario, preferring the most independent runs. Slower than the champion — the planner seat.
Verified — Two or more full runs
Fastest tested
UD-Q4_K_XL
Highest generation speed of everything measured — for jobs where waiting is the cost.
Provisional — One full run so far
Speed-value pick
UD-Q4_K_XL
Fastest setup scoring 18+/20 once the overall champion is set aside — most capability per minute of waiting.
Verified — Two or more full runs
Where models break
Six categories, six bars each.
Mean pass rate per scenario category. The weak column is usually the same one: reasoning under the hard half.
| Model | Code | Math | Data | Tools | Context | Reasoning |
|---|---|---|---|---|---|---|
| gpt-oss-120bMXFP4 | 98% | 100% | 100% | 100% | 100% | 100% |
| Gemma 4 31B QATUD-Q4_K_XL | 100% | 100% | 100% | 100% | 100% | 100% |
| Qwen3-Coder-NextUD-Q4_K_XL | 93% | 100% | 100% | 100% | 100% | 68% |
| GLM-5.3-FlashUD-IQ3_XXS | 100% | 100% | 100% | 100% | 100% | 100% |
| KAT-Coder v2.5Q8_0 | 79% | 96% | 100% | 100% | 100% | 67% |
| Qwen3-Coder-NextQ6 | 79% | 100% | 73% | 100% | 100% | 80% |
| Muse Glimmer 30BQ8_0 | 83% | 100% | 100% | 100% | 100% | 100% |
| Tiel-Coder 35BQ8_0 | 69% | 100% | 67% | 100% | 100% | 93% |
| Nemotron 3 Super 120BUD-Q4_K_XL | 83% | 100% | 89% | 100% | 84% | 100% |
| DeepSeek V4-FlashUD-IQ3_XXS | 67% | 100% | 100% | 100% | 100% | 100% |
| Qwen3.8-Flash-NextIQ4_XS-PLE · fixed sampling | 67% | 75% | 100% | 100% | 100% | 100% |
| Qwen3.8-Flash-NextUD-IQ4_XS · fixed sampling | 83% | 75% | 100% | 100% | 100% | 67% |
| Gemma 4 26B-A4B QATUD-Q4_K_XL | 67% | 100% | 100% | 100% | 50% | 0% |
| Devstral 2 24BQ8_0 | 71% | 50% | 83% | 100% | 100% | 67% |
| Qwen3.8-27BQ8_0 | 83% | 50% | 67% | 100% | 100% | 58% |
| Gemma 4 12B QATUD-Q4_K_XL | 67% | 75% | 100% | 100% | 50% | 0% |
| Mistral Small 4 119BUD-Q4_K_XL | 67% | 63% | 50% | 100% | 25% | 30% |
Coding agents
Same models, four harnesses.
Six hidden-test repository tasks per harness. Score out of six; wall-clock minutes for the set. Harness defaults are part of what's measured.
| Model | Pi | OpenCode | dsh | Hermes |
|---|---|---|---|---|
| gpt-oss-120b | 6/6 · 4m | 6/6 · 6m | 6/6 · 6m | 6/6 · 11m |
| Gemma 4 31B QAT | 6/6 · 24m | 6/6 · 34m | 6/6 · 41m | 6/6 · 26m |
| Nemotron 3 Super 120B | 6/6 · 25m | — | 6/6 · 26m | 6/6 · 35m |
| Qwen3-Coder-Next Q4 | 5.83/6 · 4m | — | 5.83/6 · 6m | 5.83/6 · 8m |
| Gemma 4 12B QAT | 5.33/6 · 33m | 6/6 · 27m | 6/6 · 34m | 4.15/6 · 27m |
Tasks: Feature: tax rebate + CLI (Python) · Multi-file bug hunt (Python) · Implement JS from a written spec · Cross-file rename + deprecated alias · Messy CSV → JSON script · Debug a timezone / loop bug (Node). One harness bug changed history — OpenCode scored 1.04/6 until headless edit permissions were fixed, then 6/6. Harness results are single runs; treat close cells as ties.
The next round (evobench, 30 private agentic tasks, hidden-test scoring) widens this matrix to every verified harness and adds frontier references through the same adapters. When it lands, this table re-points at those results and the version becomes 1.0.
The pipeline
Scouted twice a week. Scored only here.
Wednesdays and Saturdays, a research pass over leaderboards, Hugging Face, llama.cpp releases and the Strix Halo community feeds this queue. A model gets a row on the leaderboard only after it runs on the machine.
Downloaded · testing next 4 of 12
- Nex-N2.5-mini
Integrity-checked download complete. Compatibility smoke test next.
- Occamy 1.0
Integrity-checked download complete. Compatibility smoke test next.
- MiMo-V2.6-Distill-Qwen-9B9B (per name)
A small distill of the current top open-weight family. Smoke test next.
- Ling-3.0-flash-Fin
Integrity-checked download complete. Compatibility smoke test next.
Shortlisted 4 of 12
- Ornith-1.5-35B-A3BOrnith AI · 36B total · 3B active
Agentic-coding MoE with publisher-reported SWE-bench Verified of 79%. Small active size should be fast here.
- Xing4.0-29B-A4BChina Telecom AI · 29B total · 4B active
Agent-oriented MoE, Apache 2.0, 256K context; publisher reports 75 on SWE-bench Verified.
- DeepSeek-V4.1-FlashDeepSeek · Check fit
Successor to V4-Flash, which scored 18/20 here at IQ3. Needs a quant that fits 96 GB of GPU memory.
- MiMo-V2.6-FlashXiaomi · 309B total · 15B active
Strong agent scores, MIT. Only fits at roughly 2–2.5 bits per weight — a stress test of aggressive quantization.
Blocked on this hardware 1 of 12
- Ternary Bonsai 2 27BPrism ML · 27B at ~1.75 bits · 6 GB
A ternary Qwen3.8-27B that claims 98% of full quality. Needs a custom llama.cpp fork with CUDA, Metal or CPU kernels — no AMD GPU path yet.
Too big for 128 GB 3 of 12
- MiMo-V2.6-ProXiaomi · 1.02T total · 42B active
Top of the public open-weight rankings, but about four times too big for 128 GB at any usable quant.
- Kimi K3Moonshot AI · 2.8T total
Far beyond a single 128 GB machine.
- GLM-5.3Z.ai · ~750B total (reported)
The full model; its Flash sibling already scored 20/20 here.
What changed
- 2026-10-03
Bench v0.1: table-first view of the September measurements. Scout cadence moves to twice a week — Wednesday and Saturday.
- 2026-10-02
Four candidates downloaded and integrity-checked; smoke tests next. Local AI Index v1.0 published.
- 2026-09-30
September window closed: 17 configurations, 25 full runs, four coding agents.
Method and evidence
Measured here, or not listed.
The policy is the same as the lab page's, stated benchlm-style: no interpolated numbers, no imported scores, an evidence tier on every row.
- · Every score comes from automatically graded scenarios or hidden tests — never a model judging another model.
- · Verified rows rest on two or more full runs; provisional rows on one. Under-audit rows are shown but flagged.
- · Publisher benchmark claims (SWE-bench and friends) appear only in the pipeline, as claims with links, never as scores.
- · The index weights are public and versioned: capability 50 · hard half 20 · speed 15 · consistency 15.
- 25 full runs across 17 configurations; 77 scenario-runs below full marks, of which 32 were unfinished (output budget) rather than wrong. Quants, backends and sampling differ per configuration — that is the point: this measures what you would actually run. Full limitations and protocol live on the methodology page.
Measured on
One machine, fully specified.
Every number above came from this hardware — the AMD Strix Halo platform with most of 128 GB of unified memory available to the GPU.
GMKtec EVO-X2 AI · AMD Ryzen AI Max+ 395 (Strix Halo)
- 128 GB
- unified LPDDR5X-8000
- 256 GB/s
- peak memory bandwidth
- 96 GB
- addressable by the GPU
- 2 TB
- PCIe 4.0 NVMe
Processor
- Chip
- AMD Ryzen AI Max+ 395 (Strix Halo)
- CPU
- 16 Zen 5 cores · 32 threads
- Clocks
- 3.0 GHz base · up to 5.1 GHz boost
- Cache
- 16 MB L2 · 64 MB L3
- Process
- TSMC 4 nm
- Power
- Up to 140 W (GMKtec rating)AMD lists a 45–120 W configurable range; GMKtec rates this chassis at up to 140 W.
Graphics and AI
- GPU
- Radeon 8060S · 40 RDNA 3.5 compute units
- GPU clock
- Up to 2.9 GHz
- Architecture ID
- gfx1151
- NPU
- XDNA 2 · up to 50 TOPS
- Total AI
- Up to 126 TOPS (CPU + GPU + NPU)The NPU is not used by any result on this site; inference runs on the GPU.
Memory
- Capacity
- 128 GB unified, soldered
- Type
- LPDDR5X-8000
- Bus
- 256-bit
- Bandwidth
- 256 GB/s theoretical peak8000 MT/s × 256 bits ÷ 8. Token generation on large models is mostly limited by this number.
- GPU share
- 2 GB fixed carve-out + up to 96 GB shared (GTT)96 GB is the Linux default of three quarters of system memory.
Storage and I/O
- Storage
- 2 TB M.2 PCIe 4.0 NVMe
- USB
- 2 × USB4 (40 Gb/s) · 3 × USB 3.2 Gen 2 · 2 × USB 2.0
- Display
- HDMI 2.1 · DisplayPort 1.4
- Network
- 2.5 GbE · Wi-Fi 7 · Bluetooth 5.4
- Power supply
- 230 W
Software (September 2026 runs)
- OS
- Ubuntu 24.04.4 LTS, OEM kernel 7.0 · headless
- Graphics stack
- In-kernel amdgpu · Mesa 25.2.8 (RADV Vulkan)
- Inference
- llama.cpp b11238 (Vulkan) · ROCm for GLM-5.3-Flash
- Also installed
- Ollama 0.32
- Agent harnesses
- Pi 0.87.1 · OpenCode 1.18.33 · dsh 0.2.0-rc.1 · Hermes 0.21.5
- Firmware
- BIOS 1.12