Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

| Source: arXiv AI

Tags: agentic-coding, claude-opus-4, GPT-5.5, benchmark, harness, deepagents

Empirical study on 800 runs across claude-opus-4-8 and gpt-5.5 finds no average advantage for vendor-native harnesses over neutral alternatives, though Opus 4.8 trails by 9 pp on 61 repository tasks while leading by 23.7 pp on 19 contest tasks — a stratification the authors flag as post-hoc and in need of a designed replication.

Details

Agentic coding systems pair a language model with a harness — the tools, prompts, and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume native pairings solve more tasks. This paper measures that assumption directly with 800 planned runs on a private, contamination-controlled suite. The main contrasts are claude-opus-4-8 under claude-agent-sdk vs. deepagents, and gpt-5.5 under openai-codex SDK vs. deepagents. Neither resolves a statistically significant average advantage: -1.25 pp for Opus 4.8 (48.8% vs 50.0%, CI: [-10.0, +7.5]) and +1.25 pp for GPT-5.5 (55.6% vs 54.4%, CI: [-4.4, +6.9]). Averages hide a sharp stratification in the Opus contrast: the native harness trails by 9.0 pp on the 61 repository tasks and leads by 23.7 pp on the 19 contest tasks (permutation p = 0.003). This partition was chosen after seeing the data — the authors explicitly flag this and call for a pre-registered replication. Cost figures are murky: 58 Anthropic runs left no usage record, placing the Opus cost ratio anywhere from 0.7 to 2.3 depending on allocation. This version corrects a telemetry defect in the August 2026 manuscript. Code, grading oracle, and derived aggregates are released; tasks remain private.