Paper: 2609.20804 Authors: Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang Categories: cs.AI, cs.CL, cs.LG, cs.SE

The Gap

Autonomous coding agents have transitioned from curiosity to production infrastructure. Benchmark scoreboards like SWE-Bench Verified and Terminal-Bench attribute performance gains almost exclusively to base model capability. Yet in practice, models do not touch repositories raw; they run inside an execution harness.

The problem is that existing harness evaluations treat the scaffold as a monolithic black box. When OpenHands, SWE-agent, or Claude Code achieves high resolution rates, it remains unclear whether the win comes from superior tool interfaces, elaborate planning loops, or aggressive context compression. Designers routinely stack features—recoverable context caches, hierarchical planning, specialized editing tools—without measuring individual component contributions.

   PROBLEM: MONOLITHIC HARNESS EVALUATION

   Base Model (e.g. Claude / Qwen / DeepSeek)
                       |
                       v
   +-------------------------------------------------+
   | Monolithic Coding Harness                       |
   | [Planning?] + [Bash vs Custom Tools?] + [Cache?]| -> SWE-Bench Score
   +-------------------------------------------------+
                       |
   Gaps: Which component actually drives accuracy?
         Which components add overhead without gain?
                       |
                       v
   METHOD: CONTROLLED 3-COMPONENT FACTORIZATION
   Fixed Execution Loop + 176 Matched Settings
   [Planning (Off/On)] x [Action Space] x [Context Management]
                       |
                       v
   EVIDENCE: 4 Models on SWE-Bench Verified & Terminal-Bench 2.1
                       |
                       v
   CONCLUSION:
   - Rule elision > LLM summarization
   - Recoverable elision adds machinery models never use
   - Planning scaffolds weak models, saves cost on strong models
   - Bash-only interface wins for bash-proficient models

The Increment

One sentence: By holding the agent execution loop fixed and systematically isolating planning, action space, and context management across 176 matched settings, this paper demonstrates that harness components have distinct, decoupled behavioral levers: context management extends trajectory length, planning governs stopping conditions, and action spaces set editing granularity.

Core Mechanism

The study introduces a modular, lightweight harness designed for controlled component ablation. Rather than comparing disparate open-source frameworks whose loops differ in subtle ways, the authors fix the core interaction cycle: environment observation →\to model response →\to tool execution →\to environment feedback.

The authors vary three isolated dimensions across four LLMs:

  1. Context Management: Five distinct policies tested across 4 context-window budgets (16K, 32K, 64K, 128K).
    • Sliding window: drops oldest turns.
    • Rule-based elision: prunes verbose command outputs (compiler dumps, long git diffs) via heuristics.
    • LLM summarization: condenses previous interaction turns into prose summaries.
    • Staged elision-then-summarization: rule elision first, followed by summarization only when budget remains tight.
    • Recoverable elision: saves pruned text into numbered files with tool access to retrieve them.
  2. Action Space: Contrasting predefined programmatic primitives (dedicated view_file, edit_lines, create_file APIs) against a raw, bash-only execution interface (cat, sed, python, git).
  3. Planning Scaffold: Explicit upfront and dynamic step-by-step milestone planning versus zero-planning direct action.
   HARNESS DECOMPOSITION ARCHITECTURE

   Agent Interaction Turn
              |
              v
   +----------------------+
   |  Context Management  | -> Extends execution horizon
   |  (Rule > Summary)    |    (Prevents context overflow)
   +----------------------+
              |
              v
   +----------------------+
   |  Planning Scaffold   | -> Controls trajectory termination
   |  (Scaffold vs Saver) |    (Prevents endless looping)
   +----------------------+
              |
              v
   +----------------------+
   |  Action Space        | -> Dictates editing granularity
   |  (Bash vs Custom)    |    (Lower token overhead for bash)
   +----------------------+
              |
              v
   Environment Execution (Shell / Repo / Test Suite)

The load-bearing structural metaphor is scaffolding around a construction site. The base model is the stonemason.

  • Context management is the waste chute: if trash piles up, the mason trips and falls off the platform (overflow failure). Throwing out bulky drywall offcuts (rule-based elision) immediately clears the floor; cataloging every scrap into labeled bins in case the mason needs it later (recoverable elision) takes extra space and the mason never goes digging through the scrap bin.
  • Planning is the structural blueprint: an apprentice mason needs the blueprint pinned to every beam to avoid building in circles (accuracy scaffold); a master mason already knows the floor plan, so the blueprint simply reminds them when the job is done and prevents them from over-polishing (stopping condition & token cost saver).
  • Action space is the toolbelt: novice masons benefit from pre-shaped molding trowels (specialized tools); experienced masons prefer a single versatile hammer and chisel (bash shell), doing the same work faster with fewer tool changes.

Key Concepts

  • Rule-based Elision vs. Summarization: Truncating predictable high-entropy outputs (e.g. trimming terminal outputs longer than 50 lines or stripping unchanged file regions) costs zero inference tokens and preserves syntax. LLM summarization consumes tokens, introduces hallucination risks, and washes out exact compiler line numbers.
  • Recoverable Elision Overhead: Theoretical harness designs often give agents retrieval tools to fetch compressed context on demand. In practice, models almost never invoke these retrieval tools during multi-turn problem-solving, making the retrieval machinery pure deadweight.
  • Trajectory Lever Decomposition: Harness components affect execution orthogonally: context management prevents premature exit due to window exhaustion; planning dictates termination threshold; tool choice affects token payload per turn.

Framework Shift

Before (Monolithic Harness Assumption):
  More Scaffolding = Better Performance
  +--------------------------------------------------------------+
  | Add complex planning + recoverable context trees + 15 tools |
  +--------------------------------------------------------------+
  -> High token cost, unpredictable interaction effects, debugging nightmare

After (Decoupled Modular Harness):
  Component-Specific Matching by Model Capability & Context Budget
  - Context tight? -> Staged rule-based elision (zero-cost, syntax intact)
  - Capable model? -> Bash-only action space (lower token footprint)
  - Frontier model? -> Lightweight planning solely to bound trajectory cost

From “build the most elaborate agent scaffold possible,” the core shift is that harness features must be selected based on base-model capability and context budget, not accumulated as a virtue.

Expert Assessment

Problem choice: Outstanding. As frontier models converge on benchmark capabilities, engineering leverage has decisively shifted to runtime scaffolding. Dissecting the harness with rigorous controls addresses a real operational blind spot.

Method maturity: The experimental design is disciplined. Fixing the interaction loop while manipulating three orthogonal axes across 176 matched configurations eliminates the confounders that plague cross-framework comparisons. The observation that recoverable elision goes unused is an especially valuable negative result that spares engineers wasted implementation effort.

Experimental integrity: Tested across four models on both SWE-Bench Verified (code editing and test passing) and Terminal-Bench 2.1 (CLI tasks). The ablation on budget-constrained contexts (16K to 128K) is thorough and realistic. Baselines are consistent, and trajectory-level metrics (length, cost, failure modes) support the accuracy claims.

Writing quality: Direct, empirical, and free of hype. The breakdown between accuracy-oriented and cost-oriented benefits provides clear decision boundaries for system builders.

Verdict: strong accept — A definitive empirical benchmark on agent scaffolding that separates cargo-cult design from genuine runtime leverage.

Takeaways

  • Ditch recoverable context retrieval in coding agents. Use deterministic rule-based truncation of command outputs before attempting LLM summarization.
  • For top-tier models (Claude 3.5 Sonnet class), prefer a raw bash execution environment over extensive custom tool APIs; it reduces token overhead and improves velocity.
  • Use planning prompts not to make capable models smarter, but to force them to stop early and verify, cutting run costs.

论文: 2609.20804 作者: Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang 分类: cs.AI, cs.CL, cs.LG, cs.SE

缺口

自主编码智能体(Coding Agents)已从概念验证走向生产力基建。 但在各类评测榜单(如 SWE-Bench Verified、Terminal-Bench)上,大家通常将解题率的提升一股脑归功于底层大模型能力的进化。 在真实系统里,模型从不直接裸跑在代码仓库上,而是运行在一个被称为 Harness(脚手架/运行时容器)的系统之中。

现存的核心缺口在于:现有工作几乎都把 Harness 当作一个不可分割的单体黑盒来评测。 当某个智能体系统刷新榜单纪录时,我们根本无法分辨收益究竟来自精心设计的规划器、专门定制的文本编辑工具,还是精巧的上下文压缩策略。 业界工程师往往凭直觉堆叠特性——可恢复上下文缓存、多层规划、专属文件操作接口,却从不度量每个组件究竟带来了多少增益或冗余。

   问题:单体 Harness 的评估困境

   基座模型 (Claude / Qwen / DeepSeek 等)
                    |
                    v
   +-------------------------------------------------+
   | 单体黑盒脚手架 (Monolithic Coding Harness)        |
   | [规划模块?] + [Bash 还是自定义工具?] + [上下文策略?]| -> 最终得分
   +-------------------------------------------------+
                    |
   核心疑问:究竟是哪个组件决定了成功率?
            哪些组件徒增复杂度与成本却毫无正向收益?
                    |
                    v
   解法:受控的三组件正交因果解耦
   固定核心执行循环,构建 176 组严格对照实验
   [规划 (开启/关闭)] x [动作空间设计] x [上下文管理策略]
                    |
                    v
   证据:4 款基座模型在 SWE-Bench Verified 与 Terminal-Bench 2.1 上的实测
                    |
                    v
   结论:
   - 规则裁剪明显优于大模型摘要
   - 可恢复上下文机制(Recoverable Elision)模型几乎不用,纯属累赘
   - 规划对弱模型是提升准确率的拐杖,对强模型则是控制成本的刹车片
   - 具备 Bash 能力的模型,纯 Bash 接口比定制专用工具更省钱且更强

增量

一句话: 通过固定核心执行循环并在 176 组对照设定下正交剥离规划、动作空间与上下文管理,本文证明了 Harness 各组件具有清晰解耦的底层行为杠杆:上下文管理负责拓展执行寿命,规划负责控制终止时机,动作空间则决定了代码修改的操作粒度。

核心机制

为了彻底消除不同开源框架因交互循环实现细节不同导致的干扰,研究团队设计了一套极简受控 Harness。 该系统固定核心循环逻辑:环境观察 →\to 模型输出 →\to 工具调用 →\to 环境反馈。

实验在四款主流大模型上,系统化控制三类正交变量:

  1. 上下文管理策略:测试了 5 种不同策略,覆盖 16K、32K、64K、128K 四种上下文预算。
    • 滑动窗口(Sliding window):直接抛弃最早的历史轮次。
    • 基于规则的裁剪(Rule-based elision):通过正则或确定性规则修剪编译器冗长输出与大段未修改 diff。
    • LLM 摘要(LLM summarization):调用模型将历史多轮交互压制为自然语言概要。
    • 分段裁剪后摘要(Staged elision-then-summarization):先规则修剪,若预算仍告急再调用摘要。
    • 可恢复裁剪(Recoverable elision):将修剪掉的内容存入带编号的临时文件,提供专用工具供模型在需要时读回。
  2. 动作空间(Action Space):对比专门的结构化工具集(如独立提供的 view_file、edit_lines 等 API)与纯粹的原生 Bash 命令行交互(cat、sed、python、git)。
  3. 规划脚手架(Planning Scaffold):对比显式的前置里程碑与动态逐步规划,与完全不引入规划机制的直接执行。
   HARNESS 组件与智能体轨迹的底层作用机制

   智能体交互轮次
         |
         v
   +----------------------+
   |    上下文管理组件     | -> 决定轨迹能走多远(防爆窗口)
   |  (规则修剪 > 模型摘要) |
   +----------------------+
         |
         v
   +----------------------+
   |     规划脚手架       | -> 决定轨迹何时停下(控死循环与成本)
   | (弱模型支架/强模型刹车)|
   +----------------------+
         |
         v
   +----------------------+
   |      动作空间        | -> 决定修改代码的物理粒度与 Token 消耗
   |   (纯 Bash 胜出)     |
   +----------------------+
         |
         v
   环境反馈(执行结果 / 编译与测试状态)

这里的核喻是建筑工地的施工脚手架与泥瓦匠。 基座大模型就是泥瓦匠。

  • 上下文管理是建筑垃圾溜槽:如果砖块废料越堆越高,瓦匠就会被垃圾绊倒甚至摔下脚手架(上下文超限崩溃)。 把大块的废弃石膏板直接扔下溜槽(基于规则的确定性修剪),施工面立刻干净利落;而如果搞一套极其复杂的分类回收箱,把每块碎料编号打包,并配一个呼叫器让瓦匠随时调阅(可恢复裁剪),不仅占用脚手架面积,瓦匠在干活时也根本不会回头去翻垃圾箱。
  • 规划是施工图纸:刚入行的学徒瓦匠必须随时把图纸贴在每一根立柱上对照,否则就会反复拆砌、原地打转(规划作为弱模型的准确率支架);而经验丰富的老师傅早已对户型烂熟于心,图纸对他的唯一作用是明确验收标准,提醒他按时停工交房,避免无谓的精雕细刻(规划作为强模型的终止条件与 Token 节流阀)。
  • 动作空间是随身工具箱:生手需要一套分门别类的预制模板刀具(专用工具函数);而成熟的老师傅只用一把随身的多功能锤子和凿子(原汁原味的 Bash 命令行),动作行云流水,免去频繁翻箱换工具的开销。

关键概念

  • 规则修剪与模型摘要的优劣差异:对终端超长输出或无用 diff 进行确定性截断不消耗任何推理 Token,且绝不会破坏原始代码的语法精确度;而调用 LLM 做摘要不仅要付出额外的调用延迟与成本,还会抹杀具体的报错行号,甚至诱发幻觉。
  • 可恢复机制的实证失效(Unused Machinery):许多复杂脚手架理论上设计了「模型可主动检索被裁剪历史」的能力,但实测中各类大模型几乎从未调用该检索工具,复杂设计在现实中沦为纯粹的无效系统负担。
  • 轨迹杠杆的正交解耦(Trajectory Lever Decomposition):上下文管理管「生存时长」(避免提前撑爆退场),规划管「何处收手」(避免漫无目的闲逛),工具形态管「单步吞吐与开销」。

框架转变

之前(传统单体脚手架的堆砌思维):
  脚手架越复杂越好
  +--------------------------------------------------------------+
  | 无脑叠加多层规划 + 树状上下文回溯 + 15 个定制微操作工具       |
  +--------------------------------------------------------------+
  -> Token 消耗爆炸、组件间产生不可预测的相互负面干扰、极难调试

之后(基于因果解耦的轻量模块化设计):
  依模型能力与上下文预算量体裁衣
  - 上下文预算吃紧?-> 采用低成本分段规则修剪(零额外 Token,保留语法)
  - 模型本身够聪明?-> 直接提供纯 Bash 动作接口(通信负担最小化)
  - 顶尖基座模型?  -> 仅保留最轻量的规划作为终止判断,主动削减无谓开销

从「不惜代价搭建极尽复杂的 Agent 脚手架」,核心转变在于:脚手架特性的选择必须服务于基座模型的能力短板与预算边界,过度设计非但不能提升上限,反而会平添成本与故障点。

专家评审

选题眼光: 极佳。 在各大模型厂商基座能力逼近饱和的当下,软件工程落地的工程杠杆正加速向外部脚手架倾斜。 对 Harness 进行正规的组件级因果拆解,直击广大智能体开发者的认知盲区。

方法成熟度: 实验设计克制且严谨。 固定主交互循环、单变量控制三项正交属性并在 176 种配置下交叉验证,排除了不同框架间混乱的实现噪声。 发现「可恢复裁剪在实操中完全未被利用」这一负面结论极具工程指导价值,直接为后续架构设计避了坑。

实验诚意: 覆盖 SWE-Bench Verified 与 Terminal-Bench 2.1 两大权威基准,测试了 4 款不同代际与体量的模型。 在 16K 到 128K 四种严苛上下文窗口下的消融实验极具说服力,轨迹级指标(步数、花费、失败归因)扎实撑起了核心论点。

写作功力: 行文直接、数据详实,克制且无炒作词汇。 清晰区分了「准确率收益」与「成本节约收益」,为工程决策提供了可靠依据。

判决: 强接收 (strong accept) — 对智能体工程脚手架设计的一次彻底去魅与实证定调,兼具学术价值与直接的工程指导意义。

要点总结

  • 摒弃智能体脚手架里那些花哨的可恢复上下文检索机制;优先使用确定性的规则裁剪来压缩终端与文件冗余输出,慎用 LLM 摘要。
  • 面对 Claude 3.5 Sonnet 及同等级别的 Bash 熟练模型,尽量提供原生的 Bash Shell 环境,而非包一层厚重的专用文件编辑 API,这样能大幅降低 Token 消耗并提升执行流畅度。
  • 对强模型部署轻量规划的核心目的不是让它「更聪明」,而是给它设立明确的停机标准,防止陷入死循环或过度微调。