Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

| Source: MarkTechPost

Tags: HarnessDev, ByteDance Seed, agent harness, SWE-bench, Terminal-Bench, LLM evaluation, agentic AI, Opus 4.8

ByteDance Seed's HarnessDev benchmark tests 6 frontier LLMs on writing their own agent harnesses across 2,207 tasks—Opus 4.8 leads at 67.8 avg score but only 34 of 64 Evolution code changes generalize to held-out evaluation, exposing a core brittleness in LLM-authored agentic infrastructure.

Details

Most AI benchmarks treat the agent harness—the execution loop, tools, context management, and recovery logic around a model—as fixed scaffolding. HarnessDev, from researchers at ByteDance Seed, Georgia Tech, SUTD, M-A-P, and TokenWave.AI, flips that assumption: the artifact under evaluation is the runnable harness the LLM writes, not the answer it produces.\n\nSix creator LLMs were tested—Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and Seed 2.0 Pro—across two stages. Creation starts each model from a deliberately weak seed (file, search, and process primitives only; no loop, planner, verifier, or stopping rule) and asks it to build a full harness from scratch. Evolution then lets each model revise its frozen Creation harness using execution feedback from 100 SWE-bench Pro and 89 Terminal-Bench 2.1 tasks. All final harness versions are scored on 630 held-out SWE-Pro instances the creator never saw during development.\n\nResults under Self-Eval: Opus 4.8 leads with a 67.8 average against an 86.2 human-engineered reference. Domain gaps vary considerably—Opus 4.8 reaches 69.3 on SWE-Pro (vs 80.0 reference) and 84.6 on EQ-Bench3 writing tasks (actually beating the 83.7 reference). Search tasks show the widest gap: the best BrowseComp result was just 52.6 against a 92.2 reference. Counterintuitively, code volume did not predict quality—Gemini added only 1,006 net lines yet led Terminal-Bench at 68.8.\n\nThe headline finding: only 34 of 64 harness changes made during Evolution transferred to held-out evaluation. This overfitting pattern suggests current LLMs can build functional agent infrastructure for visible tasks but struggle to engineer generalizable agentic systems—a meaningful distinction as organizations build production multi-agent pipelines.