I’m putting local models through the work I actually need done. The scoreboard is only the start: the real question is which model and agent setup can finish the task, check its work, and recover when things go wrong.
Measured locallySeptember 2026 results · updated October 2, 2026
16
model configurations in the recorded 20-task results
20
automatically graded scenarios per full run
4
coding agent harnesses compared so far
Current configuration
Two models, two jobs
The September round selected a practical local team. This is a coding and planning choice; the personal-assistant role is still open.
Plan / review
Gemma 4 31B QAT
20/20
Three full runs · 10.7 tokens/s
→
Build / execute
gpt-oss-120b
19.72–19.96/20
Three full runs · about 48–49 tokens/s
GLM-5.3-Flash reached 20/20 in one full run. It remains a heavyweight contender with a different serving requirement. No single model has been declared the universal winner.
Measured
Model results
September 2026 · scores out of 20
Each score shown is a recorded full run. “Hard” covers the final ten tasks. Speed is mean generation tokens per second, rounded where several runs exist. The rows are grouped by the research record, not by a single aggregate score.
Showing 16 of 16 configurations
Local model benchmark results
Model / quant
Full runs /20
Hard /10
Tok/s
Research status
Gemma 4 31B QATUD-Q4_K_XL
20 · 20 · 20
10
10.7
Reviewer / planner
GLM-5.3-FlashAJ-IQ3_XXS
20
10
13.4
Heavyweight contender
gpt-oss-120bMXFP4
19.96 · 19.72 · 19.96
9.72–9.96
48–49
Builder / daily driver
Nemotron 3 Super 120Blocal quant
19 · 19 · 17
7–9
16.8
Contender
Muse Glimmer 30BQ8
19
9
7.6
Contender
Qwen3-Coder-NextQ4
18.56 · 19.04 · 18.29
8.56–9.04
46–48
Contender
DeepSeek V4-FlashIQ3
18
8
11.8
Contender
Qwen3-Coder-NextQ6
17.36
8.16
41
Historical
KAT-Coder v2.5Q8
17.55
7.55
51.4
Historical
Qwen3.8-Flash-Nexttwo fair-test quants
17 · 17
7
23.6–28.2
Historical
Tiel-Coder 35BQ8
16.93
7.80
43.6
Historical
Devstral 2 24BQ8
14.76
7.76
8.5
Historical
Qwen3.8-27BQ8
14.75
6
7.5
Sampling audit pending
Gemma 4 26B-A4B QATlocal quant
14
6
57.5
Historical
Gemma 4 12B QATlocal quant
13
4
25.5
Historical
Mistral Small 4 119Blocal quant
11.42
5.17
36.1
Historical
Single scores mean one full run. Historical rows are not claims that weights are still installed. Qwen3.8-Flash-Next’s two 17/20 results used separate quants and corrected sampling; they are not repeated runs of one configuration. The Qwen3.8-27B sampling/template audit is still open, so its 14.75/20 needs caution. Merged partial runs are excluded.
Six repository tasks were run with hidden tests through Pi, OpenCode, DeepSeek Harness (dsh), and Hermes. These scores measure task completion with a model and harness together.
Coding harness scores out of six
Model
Pi
OpenCode
dsh
Hermes
gpt-oss-120b
6/6
6/6
6/6
6/6
Gemma 4 31B QAT
6/6
6/6
6/6
6/6
Gemma 4 12B QAT
5.33/6
6/6
6/6
4.15/6
The first OpenCode runs scored 1.04/6 for both leading models because headless permissions blocked edits. After correcting the setup, both scored 6/6. For gpt-oss, the six tasks took roughly 4 minutes in Pi, 6 in OpenCode, 6 in dsh, and 11 in Hermes; these are rounded run summaries, not controlled speed trials. Hermes’s coding time says nothing conclusive about its assistant quality.
What broke, and what it taught me
Thinking past the answer
An early Flash-Next setup exhausted its 16K output budget on six tasks before answering. Corrected sampling and a larger budget improved it to 17/20 on two quants.
A backend can look like a model failure
GLM-5.3-Flash froze the GPU twice on a Vulkan build. A ROCm run completed at 20/20. Serving details belong in every result.
Permissions matter
OpenCode’s initial low score came from its headless tool permissions. Repeating the benchmark after the fix reversed the conclusion.
Review caught the harness
A local planner–builder–reviewer self-test was rejected because generated cache files entered the patch. After that workflow bug was fixed, the self-test passed in one round.
From the workbench
How the local team took shape
August 2026 · Fit
A 128 GB unified-memory machine made much larger local models possible. Fit was only the first question; generation speed and tool behavior decided whether they were usable.
September · Fairness
The test grew from ten to twenty scenarios. Correct chat templates, tool parsing, output budgets, sampling, and headless permissions became part of the evaluation rather than afterthoughts.
September · Agent work
Six repository tasks ran through four coding harnesses. Hidden tests exposed real completions and also showed that the leading models had saturated this small suite.
October · Team
The planner–builder–reviewer workflow passed a local self-test after its own patch hygiene bug was fixed. The next phase moves beyond coding into daily assistant work.
Method and limits
How to read these numbers
One EVO-X2 with 128 GB shared memory, one local benchmark, and twenty graded questions: coding, math, SQL, tool use, logic, and long context.
Full-run scores came from the original September benchmark summaries. Six-task agent scores came from hidden-test summaries. Runs used different quants and sometimes different backends.
Several models have one run. The six coding tasks are now saturated for leading models, so 6/6 does not separate future contenders.
The current tables do not measure everyday assistant reliability, durable memory, safety, or delegation. No OpenMuse evaluation has run yet.
Selection favors correctness and completed work. Speed, memory use, tool reliability, and reproducibility still matter for a usable local system.
The plan is to screen viable distinct models, then run harder held-out tasks for finalists. I’ll keep separate leaders for personal assistance, coding, planning/review, and specialist work rather than crown one model from a single score.
Assistant workflows
Use the same synthetic scenarios in Hermes and a verified locally compatible OpenMuse implementation: research with citations, document work, budgeting, memory and recovery, tool failures, approval boundaries, and prompt injection.
Broader harness matrix
Retest coding in Pi, OpenCode, dsh, Hermes and Claude Code where a local endpoint works. Compare direct inference through llama.cpp and Ollama separately: they are serving engines, not equivalent coding agents. Add other relevant harnesses only after a compatible, reproducible adapter is verified.
New candidates
Nex-N2.5-mini, Occamy 1.0, MiMo-V2.6-Distill-Qwen-9B, and Ling-3.0-flash-Fin have completed integrity-checked downloads as of October 2. They are ready for compatibility smoke tests, not ranked results. None has a measured local score yet.
Delegation
Compare one agent with bounded workers under equal total token and tool budgets. Track coordinator and worker costs, verifier accuracy, and failures; spawning more agents alone is not evidence of greater intelligence.
The methodology is provisional and will be refined before the next scored round. New reports will date the cohort, model version, quant, backend, prompts, budgets, run counts, and failures.