How I test local AI
A test result needs the model, the software around it, the task, and the evidence. This page separates what the September 2026 round measured from the protocol I want to refine for the next round.
Completed · September 2026
The existing evidence
Twenty-question model benchmark
Ten standard and ten hard scenarios cover programming, math, SQL, tool calls, structured output, long context, and logic. Answers were automatically graded. The raw score is out of twenty; the hard-half score is out of ten.
Six-task coding benchmark
Pi, OpenCode, DeepSeek Harness, and Hermes drove models through repository tasks. Hidden tests judged the resulting code. The harness and model form one result; changing permissions can change the outcome.
The public table transcribes original per-run summaries. It omits merged partial runs and private workspaces. Some configurations have only one run. Speed is local generation throughput, not task completion time.
Proposed · not yet scored
The next evaluation has three tracks
1. Model capability
Repeat the raw benchmark with verified templates, sampling, output budget, reasoning mode, tool parsing, and backend. Record truncation and load failures separately from wrong reasoning. Use at least three independent finalist runs.
2. Real-world assistance
Run the same synthetic scenarios through Hermes and a locally compatible OpenMuse implementation when feasible. Judge checkable artifacts, citations, numbers, tool use, memory, interruptions, recovery, and approval boundaries. Record unsupported features as coverage gaps.
3. Coding and review
Use held-out multi-file tasks with hidden tests and a separate review suite. Include the current four coding harnesses and other compatible agents after adapter checks. The original six tasks remain regression anchors because leading models already reach 6/6.
How the next results should be compared
Fairness controls
Freeze tasks and grading before seeing outcomes. Hold total tool and token budgets constant for controlled comparisons. Keep a separate practical time-to-finish view.
Role-specific leaders
Publish separate assistant, builder, planner/reviewer, and justified specialist results. Do not choose a universal winner from one aggregate.
Safety gates
An unauthorized action, cross-user data leak, forged execution claim, approval bypass, or hidden-test tampering blocks promotion even if the score is high.
Delegation controls
Compare one agent, independent attempts, and bounded worker teams. Count coordinator, worker, summarizer, retry, and verifier work in the same budget.
What every dated update will include
- Model identity, quantization, runtime and version, backend, template, sampling, context, and output limit.
- Task set, grader version, run count, per-run results, artifacts, time, tokens, and failure causes.
- Harness or application identity and permissions, helper models and external services, if any.
- Capability coverage and separate assistant, coding, and review conclusions.
- What remains untested, with the date and reason.
The details of the next assistant suite and weighting remain provisional. They will be settled before the next scored cohort, then recorded here with the dated results.