← Local AI lab

How I test local AI

A test result needs the model, the software around it, the task, and the evidence. This page separates what the September 2026 round measured from the protocol I want to refine for the next round.

Completed · September 2026

The existing evidence

Twenty-question model benchmark

Ten standard and ten hard scenarios cover programming, math, SQL, tool calls, structured output, long context, and logic. Answers were automatically graded. The raw score is out of twenty; the hard-half score is out of ten.

Six-task coding benchmark

Pi, OpenCode, DeepSeek Harness, and Hermes drove models through repository tasks. Hidden tests judged the resulting code. The harness and model form one result; changing permissions can change the outcome.

The public table transcribes original per-run summaries. It omits merged partial runs and private workspaces. Some configurations have only one run. Speed is local generation throughput, not task completion time.

Proposed · not yet scored

The next evaluation has three tracks

  1. 1. Model capability

    Repeat the raw benchmark with verified templates, sampling, output budget, reasoning mode, tool parsing, and backend. Record truncation and load failures separately from wrong reasoning. Use at least three independent finalist runs.

  2. 2. Real-world assistance

    Run the same synthetic scenarios through Hermes and a locally compatible OpenMuse implementation when feasible. Judge checkable artifacts, citations, numbers, tool use, memory, interruptions, recovery, and approval boundaries. Record unsupported features as coverage gaps.

  3. 3. Coding and review

    Use held-out multi-file tasks with hidden tests and a separate review suite. Include the current four coding harnesses and other compatible agents after adapter checks. The original six tasks remain regression anchors because leading models already reach 6/6.

How the next results should be compared

Fairness controls

Freeze tasks and grading before seeing outcomes. Hold total tool and token budgets constant for controlled comparisons. Keep a separate practical time-to-finish view.

Role-specific leaders

Publish separate assistant, builder, planner/reviewer, and justified specialist results. Do not choose a universal winner from one aggregate.

Safety gates

An unauthorized action, cross-user data leak, forged execution claim, approval bypass, or hidden-test tampering blocks promotion even if the score is high.

Delegation controls

Compare one agent, independent attempts, and bounded worker teams. Count coordinator, worker, summarizer, retry, and verifier work in the same budget.

What every dated update will include

The details of the next assistant suite and weighting remain provisional. They will be settled before the next scored cohort, then recorded here with the dated results.