Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

| Source: MarkTechPost

Tags: Claude Code, LangChain, Amazon Bedrock, Manus, context engineering, agentic AI, agent harness

A practical breakdown of 4 harness-level mechanisms — context budgeting, offloading, compaction, and todo-state — that let agents complete 200+ tool-call tasks without losing goals, with concrete thresholds from Claude Code, LangChain Deep Agents, Manus, and Amazon Bedrock.

Details

Long-horizon agentic tasks fail in two predictable ways: context overflow (the window fills up) and goal loss (the original instruction drifts to the middle of the window where recall degrades). This MarkTechPost analysis argues these are harness problems, not model problems — and walks through how production systems solve them with specific implementation thresholds. Chroma's Context Rot report evaluated 18 LLMs including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, finding reliability degrades on even simple retrieval tasks as context grows. The mechanism: attention creates n² pairwise relationships per token, depleting a finite attention budget. For agent loops, Manus reports typical tasks need ~50 tool calls at a 100:1 input-to-output token ratio — meaning context fills fast and the original goal gets buried. On context budgeting: LangChain Deep Agents offloads tool responses exceeding 20,000 tokens to disk (replaced with a file path + 10-line preview) and triggers truncation when session context hits 85% of the model window. Claude Code caps auto-memory at 200 lines or 25KB, defers MCP tool schemas until needed via tool search, and represents any re-read file over 5,000 tokens as a path reference after compaction. The article promises to cover three additional mechanisms — compaction strategy, todo-state for goal persistence, and cross-session memory — though the retrieved content was partially truncated. For practitioners building agents: treat context as a depleting resource with hard budget constraints, implement offloading at specific token thresholds, and use todo-state as a persistence layer for goal tracking across tool calls.