Contents · 1 / 7
01Abstract 02The question 03Method 04The grid 05Efficiency 06Corrections 07Cite this
JackBench · research note

Seven-Model Pi–OpenCode Baseline: Harness Effects Are Real and Model-Specific

Abstract

Agent harnesses are usually chosen on reputation, and reputations assume the harness works the same for every model. This study asked whether that assumption holds. Seven hosted models ran the same 72 tasks in Pi and OpenCode: 1,008 planned attempts, 1,003 effective, a scored sample of 153 USD across 7,255 provider requests, deterministic grading throughout. One harness effect survived Holm correction: GPT-5.6 Sol gained 19.4 points in Pi (58/72 against 44/72, adjusted p = 0.009). The other six models showed no corrected effect in either direction. So the harness effect is real, and it belongs to the pairing rather than the harness. The takeaway: a harness that transforms one model can do nothing for yours, so measure your own model-harness pairing before trusting anyone else's ranking. This is a standalone study inside the JackBench programme; the local-model counterpart is the Local Qwen 3.6 Harness Baseline, and the two are never pooled.

The question

The registered research question: when model, task and feasible operating conditions are fixed, how much does harness choice change agent outcomes? A secondary question about the frontier–near-frontier gap shares this study's evidence; its frontier layer is reported separately and never as a model-only claim.

Method

Each model ran the same 72 tasks in each harness. Grading is deterministic, every attempt writes a sealed receipt, comparisons were registered in advance, and every reported p-value carries a Holm correction across the registered family. Of 1,008 planned attempts, 5 were lost to provider failures; the effective sample is 1,003.

Receipts

1,034 physical receipts including replacements; 32 provider-invalid attempts (1.08 USD) are excluded, and reported by identity and reason.

The grid

PiOpenCode
GPT-5.6 Sol confirmatory
80.6%
61.1%
Opus 5
72.2%
73.6%
GLM 5.2
65.3%
52.8%
Qwen 3.8 Max
61.1%
63.9%
K3
59.7%
61.1%
Qwen 3.7 Max
54.2%
45.8%
Fable 5withheld: 3 of 72 pairs lost to provider failures; 40/40 on the 69 complete pairs
Figure 1 as a table
GPT-5.6 SolPi 80.6%OpenCode 61.1%Holm p = 0.009
Opus 5Pi 72.2%OpenCode 73.6%p = 1.0
GLM 5.2Pi 65.3%OpenCode 52.8%p = 0.294
Qwen 3.8 MaxPi 61.1%OpenCode 63.9%p = 1.0
K3Pi 59.7%OpenCode 61.1%p = 1.0
Qwen 3.7 MaxPi 54.2%OpenCode 45.8%p = 0.898
Fable 5withheld, provider missingness
Figure 1 · Solve rate by model and harness, 72 tasks per cell. Sol's gap is the only effect that survived Holm correction. Data from the sealed reconciliation record.

Best cell in the grid: Sol on Pi at 80.6%. The model-specific spread is the finding: the same harness change helped Sol by 19 points and did nothing measurable for the other six.

Download the results (CSV)

Efficiency

K3 succeeded at near-identical rates in both harnesses whilst spending 5.62× more and taking 3.22× the median duration in OpenCode. Judged on solve rate alone the two set-ups look interchangeable for K3, which is exactly why the study records cost and duration as well.

Two more efficiency observations from the sealed record: Qwen 3.8 Max and Opus 5 gained no confirmatory quality from their more expensive OpenCode cells, and Sol's Pi cell cost slightly more than its OpenCode cell whilst producing 14 additional successes and the only adjusted-significant advantage.

Corrections and limitations

No corrections recorded. If something in here is wrong, tell me. Fixes get logged in this section with a date.

Known limits: two harnesses, one task set, and provider missingness that withheld Fable 5's comparison. Hosted results don't license claims about harnesses this study never ran; those are being measured separately and are listed on the research page with their own statuses.

Cite this

Tyler, J. (2026). Seven-Model Pi–OpenCode Baseline. Cold Anvil Studios. https://coldanvil.com/research/seven-model-pi-opencode-baseline/

@misc{tyler2026sevenmodel,
  author       = {Tyler, Jack},
  title        = {Seven-Model Pi--OpenCode Baseline},
  year         = {2026},
  publisher    = {Cold Anvil Studios},
  howpublished = {\url{https://coldanvil.com/research/seven-model-pi-opencode-baseline/}}
}