Strix Halo Bench · v0.1 · October 3, 2026

Every open model that fits this machine, on one table.

A benchmark-dashboard view of what a 128 GB home machine can actually run: index scores, category splits, agent harnesses and the Pareto frontier of quality against speed.

Setups scored
17
Graded scenarios
20
Agent harnesses
4
In the pipeline
12

Current answer

gpt-oss-120b (MXFP4) is the setup to run today — 19.88/20 at 48.5 tok/s, index 98.4. The perfect scorer Gemma 4 31B takes the planner seat when minutes don't matter. Four downloaded candidates are next in the queue; they get a row here only after they run here.

The leaderboard

Sorted, filtered, evidenced.

One row per configuration. The index bar shows where its points come from; the evidence column says how many independent runs back the number. Click a row for category strengths and its best agent run.

17 of 17 setups
#ModelIndex/20Hardtok/sLoadEvidence

Index = capability 50 · hard half 20 · speed 15 · consistency 15 (same weights as the AI Index). Evidence reflects independent full runs; provisional numbers move.

Quality against speed

The frontier nobody dominates.

Log-scale speed on the x-axis, index on the y. The dashed line is the Pareto frontier: every setup on it is the best there is at that speed.

5102040606080100generation speed (tokens/s, log scale)index /100best corner →gpt-oss-120b (MXFP4) — index 98.4, 48.5 tok/s, Current team, on the Pareto frontiergpt-oss-120bGemma 4 31B QAT (UD-Q4_K_XL) — index 94.4, 10.6 tok/s, Current teamQwen3-Coder-Next (UD-Q4_K_XL) — index 91.1, 46.8 tok/s, ContenderGLM-5.3-Flash (UD-IQ3_XXS) — index 89.2, 13.4 tok/s, ContenderKAT-Coder v2.5 (Q8_0) — index 83.0, 51.4 tok/s, Retired, on the Pareto frontierKAT-Coder v2.5Qwen3-Coder-Next (Q6) — index 83.0, 41 tok/s, RetiredMuse Glimmer 30B (Q8_0) — index 82.7, 7.6 tok/s, ContenderTiel-Coder 35B (Q8_0) — index 81.4, 43.6 tok/s, RetiredNemotron 3 Super 120B (UD-Q4_K_XL) — index 81.0, 16.8 tok/s, ContenderDeepSeek V4-Flash (UD-IQ3_XXS) — index 79.7, 11.8 tok/s, ContenderQwen3.8-Flash-Next (IQ4_XS-PLE · fixed sampling) — index 78.4, 28.2 tok/s, RetiredQwen3.8-Flash-Next (UD-IQ4_XS · fixed sampling) — index 77.7, 23.6 tok/s, RetiredGemma 4 26B-A4B QAT (UD-Q4_K_XL) — index 71.0, 57.5 tok/s, Retired, on the Pareto frontierGemma 4 26B-A4B Q…Devstral 2 24B (Q8_0) — index 70.0, 8.5 tok/s, RetiredQwen3.8-27B (Q8_0) — index 66.0, 7.5 tok/s, Under auditGemma 4 12B QAT (UD-Q4_K_XL) — index 62.0, 25.5 tok/s, RetiredMistral Small 4 119B (UD-Q4_K_XL) — index 61.7, 36.1 tok/s, Retired
Dashed line: the Pareto frontier — setups nobody beats on both quality and speed. Everything below-left of it trades one for the other.

Decision-ready

Four picks, each with its rule.

benchlm-style: if you won't read the table, read these. Every pick states how it was chosen, and carries the same evidence tier as its row.

Best overall

gpt-oss-120b98.4 / 100

MXFP4

Highest index — accuracy on all twenty scenarios, hard half, speed and run-to-run consistency.

Verified — Two or more full runs

The perfect scorer

Gemma 4 31B QAT20.00 / 20 · 3 runs

UD-Q4_K_XL

Full marks on every scenario, preferring the most independent runs. Slower than the champion — the planner seat.

Verified — Two or more full runs

Fastest tested

Gemma 4 26B-A4B QAT57.5 tok/s

UD-Q4_K_XL

Highest generation speed of everything measured — for jobs where waiting is the cost.

Provisional — One full run so far

Speed-value pick

Qwen3-Coder-Next46.8 tok/s · 18.63/20

UD-Q4_K_XL

Fastest setup scoring 18+/20 once the overall champion is set aside — most capability per minute of waiting.

Verified — Two or more full runs

Where models break

Six categories, six bars each.

Mean pass rate per scenario category. The weak column is usually the same one: reasoning under the hard half.

ModelCodeMathDataToolsContextReasoning
gpt-oss-120bMXFP498%100%100%100%100%100%
Gemma 4 31B QATUD-Q4_K_XL100%100%100%100%100%100%
Qwen3-Coder-NextUD-Q4_K_XL93%100%100%100%100%68%
GLM-5.3-FlashUD-IQ3_XXS100%100%100%100%100%100%
KAT-Coder v2.5Q8_079%96%100%100%100%67%
Qwen3-Coder-NextQ679%100%73%100%100%80%
Muse Glimmer 30BQ8_083%100%100%100%100%100%
Tiel-Coder 35BQ8_069%100%67%100%100%93%
Nemotron 3 Super 120BUD-Q4_K_XL83%100%89%100%84%100%
DeepSeek V4-FlashUD-IQ3_XXS67%100%100%100%100%100%
Qwen3.8-Flash-NextIQ4_XS-PLE · fixed sampling67%75%100%100%100%100%
Qwen3.8-Flash-NextUD-IQ4_XS · fixed sampling83%75%100%100%100%67%
Gemma 4 26B-A4B QATUD-Q4_K_XL67%100%100%100%50%0%
Devstral 2 24BQ8_071%50%83%100%100%67%
Qwen3.8-27BQ8_083%50%67%100%100%58%
Gemma 4 12B QATUD-Q4_K_XL67%75%100%100%50%0%
Mistral Small 4 119BUD-Q4_K_XL67%63%50%100%25%30%

Coding agents

Same models, four harnesses.

Six hidden-test repository tasks per harness. Score out of six; wall-clock minutes for the set. Harness defaults are part of what's measured.

ModelPiOpenCodedshHermes
gpt-oss-120b6/6 · 4m6/6 · 6m6/6 · 6m6/6 · 11m
Gemma 4 31B QAT6/6 · 24m6/6 · 34m6/6 · 41m6/6 · 26m
Nemotron 3 Super 120B6/6 · 25m—6/6 · 26m6/6 · 35m
Qwen3-Coder-Next Q45.83/6 · 4m—5.83/6 · 6m5.83/6 · 8m
Gemma 4 12B QAT5.33/6 · 33m6/6 · 27m6/6 · 34m4.15/6 · 27m

Tasks: Feature: tax rebate + CLI (Python) · Multi-file bug hunt (Python) · Implement JS from a written spec · Cross-file rename + deprecated alias · Messy CSV → JSON script · Debug a timezone / loop bug (Node). One harness bug changed history — OpenCode scored 1.04/6 until headless edit permissions were fixed, then 6/6. Harness results are single runs; treat close cells as ties.

The next round (evobench, 30 private agentic tasks, hidden-test scoring) widens this matrix to every verified harness and adds frontier references through the same adapters. When it lands, this table re-points at those results and the version becomes 1.0.

The pipeline

Scouted twice a week. Scored only here.

Wednesdays and Saturdays, a research pass over leaderboards, Hugging Face, llama.cpp releases and the Strix Halo community feeds this queue. A model gets a row on the leaderboard only after it runs on the machine.

Downloaded · testing next 4 of 12

  • Nex-N2.5-mini

    Integrity-checked download complete. Compatibility smoke test next.

  • Occamy 1.0

    Integrity-checked download complete. Compatibility smoke test next.

  • MiMo-V2.6-Distill-Qwen-9B9B (per name)

    A small distill of the current top open-weight family. Smoke test next.

  • Ling-3.0-flash-Fin

    Integrity-checked download complete. Compatibility smoke test next.

Shortlisted 4 of 12

  • Ornith-1.5-35B-A3BOrnith AI · 36B total · 3B active

    Agentic-coding MoE with publisher-reported SWE-bench Verified of 79%. Small active size should be fast here.

  • Xing4.0-29B-A4BChina Telecom AI · 29B total · 4B active

    Agent-oriented MoE, Apache 2.0, 256K context; publisher reports 75 on SWE-bench Verified.

  • DeepSeek-V4.1-FlashDeepSeek · Check fit

    Successor to V4-Flash, which scored 18/20 here at IQ3. Needs a quant that fits 96 GB of GPU memory.

  • MiMo-V2.6-FlashXiaomi · 309B total · 15B active

    Strong agent scores, MIT. Only fits at roughly 2–2.5 bits per weight — a stress test of aggressive quantization.

Blocked on this hardware 1 of 12

  • Ternary Bonsai 2 27BPrism ML · 27B at ~1.75 bits · 6 GB

    A ternary Qwen3.8-27B that claims 98% of full quality. Needs a custom llama.cpp fork with CUDA, Metal or CPU kernels — no AMD GPU path yet.

Too big for 128 GB 3 of 12

  • MiMo-V2.6-ProXiaomi · 1.02T total · 42B active

    Top of the public open-weight rankings, but about four times too big for 128 GB at any usable quant.

  • Kimi K3Moonshot AI · 2.8T total

    Far beyond a single 128 GB machine.

  • GLM-5.3Z.ai · ~750B total (reported)

    The full model; its Flash sibling already scored 20/20 here.

What changed

  1. 2026-10-03

    Bench v0.1: table-first view of the September measurements. Scout cadence moves to twice a week — Wednesday and Saturday.

  2. 2026-10-02

    Four candidates downloaded and integrity-checked; smoke tests next. Local AI Index v1.0 published.

  3. 2026-09-30

    September window closed: 17 configurations, 25 full runs, four coding agents.

Method and evidence

Measured here, or not listed.

The policy is the same as the lab page's, stated benchlm-style: no interpolated numbers, no imported scores, an evidence tier on every row.

  • · Every score comes from automatically graded scenarios or hidden tests — never a model judging another model.
  • · Verified rows rest on two or more full runs; provisional rows on one. Under-audit rows are shown but flagged.
  • · Publisher benchmark claims (SWE-bench and friends) appear only in the pipeline, as claims with links, never as scores.
  • · The index weights are public and versioned: capability 50 · hard half 20 · speed 15 · consistency 15.
    25 full runs across 17 configurations; 77 scenario-runs below full marks, of which 32 were unfinished (output budget) rather than wrong. Quants, backends and sampling differ per configuration — that is the point: this measures what you would actually run. Full limitations and protocol live on the methodology page.

Measured on

One machine, fully specified.

Every number above came from this hardware — the AMD Strix Halo platform with most of 128 GB of unified memory available to the GPU.

GMKtec EVO-X2 AI · AMD Ryzen AI Max+ 395 (Strix Halo)

128 GB
unified LPDDR5X-8000
256 GB/s
peak memory bandwidth
96 GB
addressable by the GPU
2 TB
PCIe 4.0 NVMe

Processor

Chip
AMD Ryzen AI Max+ 395 (Strix Halo)
CPU
16 Zen 5 cores · 32 threads
Clocks
3.0 GHz base · up to 5.1 GHz boost
Cache
16 MB L2 · 64 MB L3
Process
TSMC 4 nm
Power
Up to 140 W (GMKtec rating)AMD lists a 45–120 W configurable range; GMKtec rates this chassis at up to 140 W.

Graphics and AI

GPU
Radeon 8060S · 40 RDNA 3.5 compute units
GPU clock
Up to 2.9 GHz
Architecture ID
gfx1151
NPU
XDNA 2 · up to 50 TOPS
Total AI
Up to 126 TOPS (CPU + GPU + NPU)The NPU is not used by any result on this site; inference runs on the GPU.

Memory

Capacity
128 GB unified, soldered
Type
LPDDR5X-8000
Bus
256-bit
Bandwidth
256 GB/s theoretical peak8000 MT/s × 256 bits ÷ 8. Token generation on large models is mostly limited by this number.
GPU share
2 GB fixed carve-out + up to 96 GB shared (GTT)96 GB is the Linux default of three quarters of system memory.

Storage and I/O

Storage
2 TB M.2 PCIe 4.0 NVMe
USB
2 × USB4 (40 Gb/s) · 3 × USB 3.2 Gen 2 · 2 × USB 2.0
Display
HDMI 2.1 · DisplayPort 1.4
Network
2.5 GbE · Wi-Fi 7 · Bluetooth 5.4
Power supply
230 W

Software (September 2026 runs)

OS
Ubuntu 24.04.4 LTS, OEM kernel 7.0 · headless
Graphics stack
In-kernel amdgpu · Mesa 25.2.8 (RADV Vulkan)
Inference
llama.cpp b11238 (Vulkan) · ROCm for GLM-5.3-Flash
Also installed
Ollama 0.32
Agent harnesses
Pi 0.87.1 · OpenCode 1.18.33 · dsh 0.2.0-rc.1 · Hermes 0.21.5
Firmware
BIOS 1.12