Local intelligence lab / Research snapshot 01

How intelligent is local AI?

I’m putting local models through the work I actually need done. The scoreboard is only the start: the real question is which model and agent setup can finish the task, check its work, and recover when things go wrong.

Measured locallySeptember 2026 results · updated October 2, 2026

16

model configurations in the recorded 20-task results

20

automatically graded scenarios per full run

4

coding agent harnesses compared so far

Current configuration

Two models, two jobs

The September round selected a practical local team. This is a coding and planning choice; the personal-assistant role is still open.

Plan / review

Gemma 4 31B QAT

20/20

Three full runs · 10.7 tokens/s

Build / execute

gpt-oss-120b

19.72–19.96/20

Three full runs · about 48–49 tokens/s

GLM-5.3-Flash reached 20/20 in one full run. It remains a heavyweight contender with a different serving requirement. No single model has been declared the universal winner.

Measured

Model results

September 2026 · scores out of 20

Each score shown is a recorded full run. “Hard” covers the final ten tasks. Speed is mean generation tokens per second, rounded where several runs exist. The rows are grouped by the research record, not by a single aggregate score.

Showing 16 of 16 configurations

Local model benchmark results
Model / quantFull runs /20Hard /10Tok/sResearch status
Gemma 4 31B QATUD-Q4_K_XL20 · 20 · 201010.7Reviewer / planner
GLM-5.3-FlashAJ-IQ3_XXS201013.4Heavyweight contender
gpt-oss-120bMXFP419.96 · 19.72 · 19.969.72–9.9648–49Builder / daily driver
Nemotron 3 Super 120Blocal quant19 · 19 · 177–916.8Contender
Muse Glimmer 30BQ81997.6Contender
Qwen3-Coder-NextQ418.56 · 19.04 · 18.298.56–9.0446–48Contender
DeepSeek V4-FlashIQ318811.8Contender
Qwen3-Coder-NextQ617.368.1641Historical
KAT-Coder v2.5Q817.557.5551.4Historical
Qwen3.8-Flash-Nexttwo fair-test quants17 · 17723.6–28.2Historical
Tiel-Coder 35BQ816.937.8043.6Historical
Devstral 2 24BQ814.767.768.5Historical
Qwen3.8-27BQ814.7567.5Sampling audit pending
Gemma 4 26B-A4B QATlocal quant14657.5Historical
Gemma 4 12B QATlocal quant13425.5Historical
Mistral Small 4 119Blocal quant11.425.1736.1Historical

Single scores mean one full run. Historical rows are not claims that weights are still installed. Qwen3.8-Flash-Next’s two 17/20 results used separate quants and corrected sampling; they are not repeated runs of one configuration. The Qwen3.8-27B sampling/template audit is still open, so its 14.75/20 needs caution. Merged partial runs are excluded.

Download the sanitized result snapshot (JSON)

Measured

Coding agents change the outcome

Six repository tasks were run with hidden tests through Pi, OpenCode, DeepSeek Harness (dsh), and Hermes. These scores measure task completion with a model and harness together.

Coding harness scores out of six
ModelPiOpenCodedshHermes
gpt-oss-120b6/66/66/66/6
Gemma 4 31B QAT6/66/66/66/6
Gemma 4 12B QAT5.33/66/66/64.15/6

The first OpenCode runs scored 1.04/6 for both leading models because headless permissions blocked edits. After correcting the setup, both scored 6/6. For gpt-oss, the six tasks took roughly 4 minutes in Pi, 6 in OpenCode, 6 in dsh, and 11 in Hermes; these are rounded run summaries, not controlled speed trials. Hermes’s coding time says nothing conclusive about its assistant quality.

What broke, and what it taught me

Thinking past the answer

An early Flash-Next setup exhausted its 16K output budget on six tasks before answering. Corrected sampling and a larger budget improved it to 17/20 on two quants.

A backend can look like a model failure

GLM-5.3-Flash froze the GPU twice on a Vulkan build. A ROCm run completed at 20/20. Serving details belong in every result.

Permissions matter

OpenCode’s initial low score came from its headless tool permissions. Repeating the benchmark after the fix reversed the conclusion.

Review caught the harness

A local planner–builder–reviewer self-test was rejected because generated cache files entered the patch. After that workflow bug was fixed, the self-test passed in one round.

From the workbench

How the local team took shape

  1. August 2026 · Fit

    A 128 GB unified-memory machine made much larger local models possible. Fit was only the first question; generation speed and tool behavior decided whether they were usable.

  2. September · Fairness

    The test grew from ten to twenty scenarios. Correct chat templates, tool parsing, output budgets, sampling, and headless permissions became part of the evaluation rather than afterthoughts.

  3. September · Agent work

    Six repository tasks ran through four coding harnesses. Hidden tests exposed real completions and also showed that the leading models had saturated this small suite.

  4. October · Team

    The planner–builder–reviewer workflow passed a local self-test after its own patch hygiene bug was fixed. The next phase moves beyond coding into daily assistant work.

Method and limits

How to read these numbers

Read the testing protocol →

Planned · no scores yet

The next testing program

The plan is to screen viable distinct models, then run harder held-out tasks for finalists. I’ll keep separate leaders for personal assistance, coding, planning/review, and specialist work rather than crown one model from a single score.

Assistant workflows

Use the same synthetic scenarios in Hermes and a verified locally compatible OpenMuse implementation: research with citations, document work, budgeting, memory and recovery, tool failures, approval boundaries, and prompt injection.

Broader harness matrix

Retest coding in Pi, OpenCode, dsh, Hermes and Claude Code where a local endpoint works. Compare direct inference through llama.cpp and Ollama separately: they are serving engines, not equivalent coding agents. Add other relevant harnesses only after a compatible, reproducible adapter is verified.

New candidates

Nex-N2.5-mini, Occamy 1.0, MiMo-V2.6-Distill-Qwen-9B, and Ling-3.0-flash-Fin have completed integrity-checked downloads as of October 2. They are ready for compatibility smoke tests, not ranked results. None has a measured local score yet.

Delegation

Compare one agent with bounded workers under equal total token and tool budgets. Track coordinator and worker costs, verifier accuracy, and failures; spawning more agents alone is not evidence of greater intelligence.

The methodology is provisional and will be refined before the next scored round. New reports will date the cohort, model version, quant, backend, prompts, budgets, run counts, and failures.

Read the AI news digest