Research
The studies to come out of JackBench, the studio's evaluation system.
Completed studies
-
First Pi–OpenCode BaselineStudy CompleteAskedDo agent harnesses change outcomes for near-frontier models at all? TestedThree near-frontier models under Pi and OpenCode, preregistered, deterministically graded. LearnedThe pooled quality result was a null. The money told the real story: Kimi K3 tied on quality whilst spending 4× more under OpenCode. Judge a harness on cost per success, because pass rate alone hides it.4.06×cost, same quality
-
Seven-Model Pi–OpenCode BaselineStudy CompleteAskedWith the model held fixed, does the harness around it change how much work gets done? TestedSeven hosted models, the same 72 tasks in Pi and OpenCode, 1,008 attempts, deterministic grading. LearnedThe harness effect is real, and it belongs to the pairing. GPT-5.6 Sol gained 19.4 points in Pi; the other six didn't move after correction. Measure your own model-and-harness pairing before trusting a leaderboard.+19.4points, one model
-
Local Qwen 3.6 Harness BaselineStudy CompleteAskedDoes harness choice matter for a small local model as much as it does for hosted ones? TestedQwen 3.6 27B on the same 144 tasks in Pi and OpenCode, 720 attempts, run entirely on the studio's hardware. LearnedSame model, same tasks, 19.4 points apart. For local setups the harness matters as much as the model, and testing yours costs electricity rather than provider bills.+19.4points, harness alone
-
Blinded Judge-Panel QualificationStudy CompleteAskedCan a language model be trusted to judge another model's work, and how would you know? TestedBlind exams against locked known answers: 112 predictions per candidate, both candidate orders, repeated. LearnedFive models from five labs earned a seat. The exam also refuses: the studio's own local Qwen 3.8 scored 96.4% and was declined for changing its answers between repeats. Don't trust a judge that hasn't passed a blind exam first.5labs qualified
-
Hosted Claude Code, Codex and Hermes BridgeBeing writtenAskedHow do three more harnesses handle the task set the controlled baselines already use? Tested504 attempts across Claude Code, Codex and Hermes, sealed and closed out, roughly 170 USD.504attempts run
-
Local Qwen 3.8 Five-Harness StudyBeing writtenAskedHow does the newer local Qwen behave across five different harnesses? Tested715 attempts across Pi, OpenCode, Claude Code, Hermes and Codex on 143 tasks, all sealed, zero provider cost.715attempts run
-
Warp Agent CLI Ecological StudyBeing writtenAskedWhat does a managed agent product deliver out of the box, before anyone tunes anything? Tested400 attempts across five managed model selections, 28 USD as Warp reported it.400attempts run
Upcoming studies
-
Qwen 3.8 Regular versus Uncensored Security StudyBacklogAskedDoes an uncensored build of the same local model do better or worse at security work? TestedWill test: 216 red-team and blue-team sessions against synthetic security vulnerabilities in a disposable sandbox. The design is published; results come when the run is sealed.216sessions planned
-
Local Qwen 3.8 versus 3.6 Matched Pi StudyRunningAskedIs the newer Qwen generation actually better, like for like, when everything else is held fixed? TestedTesting now: a 480-attempt matched comparison at identical 4-bit quantisation, on the studio's hardware. Results stay sealed until the run completes.480attempts running