Paper: 2609.13082 Authors: Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Danyang Li, Zheng Yang, Jianwei Hu, Qiang Ma Categories: cs.AI
The Gap
Agentic systems are a promising way to automate embodied benchmark construction, and existing approaches fall short in two distinct ways. The first is coverage: they typically cover isolated stages or remain specialized to predefined environments and task families.
The second is stated as the more important one, and it is a structural property of pipelines rather than a limitation of any stage: multi-step construction produces dependent intermediate artifacts that are often passed downstream without artifact-specific verification, allowing local defects to propagate into the final benchmark.
That distinction matters because the two failures need different remedies. A coverage gap is closed by adding stages or broadening the environment set. Propagation is not closed by adding anything — it requires verification between stages, and specifically verification that is artifact-specific, since a defect in a scene specification and a defect in a task instruction are not detectable by the same check.
AUTOMATING EMBODIED BENCHMARK CONSTRUCTION
agentic systems are promising for this
existing approaches fall short in TWO DISTINCT WAYS
[1] COVERAGE
they cover ISOLATED STAGES
or remain SPECIALIZED TO PREDEFINED ENVIRONMENTS AND TASK
FAMILIES
[2] PROPAGATION (stated as the MORE IMPORTANT one)
MULTI-STEP CONSTRUCTION PRODUCES DEPENDENT INTERMEDIATE
ARTIFACTS THAT ARE OFTEN PASSED DOWNSTREAM WITHOUT
ARTIFACT-SPECIFIC VERIFICATION
-> ALLOWING LOCAL DEFECTS TO PROPAGATE INTO THE FINAL
BENCHMARK
<- a STRUCTURAL PROPERTY OF PIPELINES, not a limitation of any
single STAGE
[THE DISTINCTION MATTERS BECAUSE THE TWO FAILURES NEED DIFFERENT
REMEDIES]
a COVERAGE gap is closed by ADDING STAGES or BROADENING THE
ENVIRONMENT SET
PROPAGATION is NOT closed by adding anything
-> requires VERIFICATION **BETWEEN** STAGES
-> and specifically verification that is ARTIFACT-SPECIFIC, since
a defect in a SCENE SPECIFICATION and a defect in a TASK
INSTRUCTION ARE NOT DETECTABLE BY THE SAME CHECK
The Increment
One sentence: Before this paper, multi-step benchmark construction passed intermediates downstream unverified; after it, closed-loop synthesis verifies each artifact against its own contract and rolls back by provenance, producing seven benchmarks that discriminate model capability.
Core Mechanism
The framework transforms user-specified evaluation intents into complete embodied benchmark artifacts, and its organising idea is stated as a formulation: construction as Closed-Loop Benchmark Synthesis, integrating forward artifact synthesis with backward verification and repair. The loop is what distinguishes it — synthesis and verification are two directions through the same pipeline, not a pipeline followed by a check at the end.
The forward direction: Skill-Orchestrated Artifact Synthesis, which composes typed and reusable skills into executable workflows, while an artifact dependency graph records intermediate outputs and their dependencies. Two properties are designed in. Typed skills mean a skill’s interfaces are declared, so composition can be checked rather than assumed. And the dependency graph is the structural prerequisite for the backward direction — you cannot roll back to the right place without knowing what depended on what.
The backward direction: Requirement-Guided Verification and Repair, which applies artifact-specific contracts throughout construction and uses provenance to trigger local re-execution or upstream rollback when verification fails. The two mechanisms answer different needs. Artifact-specific contracts applied throughout means defects are caught at the artifact where they occur, rather than at the end where all that remains is a bad benchmark. And provenance-driven rollback is the part that makes repair efficient: on failure, the system recomputes locally where it can, and escalates upstream only when it must — which matters because a pipeline whose only remedy is to start over cannot be run at scale.
The results cover both track types. Six benchmarks are constructed in the Offline EQA Track, covering diverse embodied scenarios, plus one interactive benchmark containing 220 executable tasks in the Interactive Embodied Track. Having both matters for the claim: offline evaluation tests observation-based understanding while interactive execution tests closed-loop behaviour, so constructing both shows the framework is not specialised to one stage or one environment family.
And the evaluations are ordered from the artifact to the process. Representative MLLMs and embodied agents are evaluated, showing the benchmarks distinguish model capabilities in both observation-based understanding and closed-loop execution — so the artifacts are usable measuring instruments, not just constructed outputs. Then quality assessment and ablations validate benchmark quality and the effectiveness of verification and repair, which tests the paper’s own mechanism rather than only its products. And repair and skill-reuse analyses demonstrate efficient localized recovery and cross-benchmark reusability — the two properties the design was built to deliver: local rather than global repair, and skills reused across benchmarks rather than rebuilt.
THE ORGANISING IDEA, STATED AS A FORMULATION
CONSTRUCTION AS CLOSED-LOOP BENCHMARK SYNTHESIS,
INTEGRATING FORWARD ARTIFACT SYNTHESIS WITH BACKWARD VERIFICATION
AND REPAIR
<- the LOOP is what distinguishes it: SYNTHESIS and VERIFICATION
are TWO DIRECTIONS THROUGH THE SAME PIPELINE, not a PIPELINE
FOLLOWED BY A CHECK AT THE END
[THE FORWARD DIRECTION: SKILL-ORCHESTRATED ARTIFACT SYNTHESIS]
COMPOSES TYPED AND REUSABLE SKILLS INTO EXECUTABLE WORKFLOWS
an ARTIFACT DEPENDENCY GRAPH RECORDS INTERMEDIATE OUTPUTS AND THEIR
DEPENDENCIES
<- TYPED skills: a skill's INTERFACES ARE DECLARED, so COMPOSITION
CAN BE CHECKED rather than ASSUMED
<- the DEPENDENCY GRAPH is the STRUCTURAL PREREQUISITE for the
BACKWARD DIRECTION: you CANNOT ROLL BACK TO THE RIGHT PLACE
without knowing WHAT DEPENDED ON WHAT
[THE BACKWARD DIRECTION: REQUIREMENT-GUIDED VERIFICATION AND REPAIR]
APPLIES ARTIFACT-SPECIFIC CONTRACTS THROUGHOUT CONSTRUCTION
USES PROVENANCE TO TRIGGER LOCAL RE-EXECUTION OR UPSTREAM ROLLBACK
WHEN VERIFICATION FAILS
<- ARTIFACT-SPECIFIC CONTRACTS THROUGHOUT means defects are caught
AT THE ARTIFACT WHERE THEY OCCUR, rather than AT THE END where
ALL THAT REMAINS IS A BAD BENCHMARK
<- PROVENANCE-DRIVEN ROLLBACK is what makes REPAIR EFFICIENT:
on failure, RECOMPUTE LOCALLY where it can, ESCALATE UPSTREAM
only when it must
<- matters because a pipeline whose ONLY REMEDY IS TO START OVER
CANNOT BE RUN AT SCALE
[THE RESULTS COVER BOTH TRACK TYPES]
SIX benchmarks in the OFFLINE EQA TRACK, covering DIVERSE EMBODIED
SCENARIOS
ONE INTERACTIVE BENCHMARK with 220 EXECUTABLE TASKS in the
INTERACTIVE EMBODIED TRACK
<- HAVING BOTH matters for the claim: OFFLINE evaluation tests
OBSERVATION-BASED UNDERSTANDING while INTERACTIVE EXECUTION tests
CLOSED-LOOP BEHAVIOUR
-> constructing both shows the framework is NOT SPECIALISED to one
STAGE or one ENVIRONMENT FAMILY
[AND THE EVALUATIONS ARE ORDERED FROM THE ARTIFACT TO THE PROCESS]
representative MLLMs AND EMBODIED AGENTS evaluated
-> the benchmarks DISTINGUISH MODEL CAPABILITIES in BOTH
observation-based understanding AND closed-loop execution
<- the artifacts are USABLE MEASURING INSTRUMENTS, not just
constructed outputs
QUALITY ASSESSMENT AND ABLATIONS validate benchmark quality AND THE
EFFECTIVENESS OF VERIFICATION AND REPAIR
<- tests the paper's OWN MECHANISM, not only its PRODUCTS
REPAIR AND SKILL-REUSE ANALYSES demonstrate EFFICIENT LOCALIZED
RECOVERY and CROSS-BENCHMARK REUSABILITY
<- the two properties the design was BUILT to deliver: LOCAL
rather than GLOBAL repair, and SKILLS REUSED ACROSS BENCHMARKS
rather than REBUILT
Think of it as an assembly line with inspection stations rather than a final inspection at the loading dock. A final check catches only what is still visible at the end, and by then a mis-drilled hole is a property of the finished product rather than a fixable station error. Two details of the paper’s version map onto the machinery. The stations have checklists specific to what they handle — a scene specification and a task instruction are inspected differently — which is what “artifact-specific contracts” means. And each part carries a record of where it came from, so a failure sends the work back to the station that produced it rather than restarting the whole line. The result is that a defect stops at its station instead of becoming a property of the finished benchmark.
Key Concepts
- Propagation as the structural failure: unverified intermediates passing downstream. It is a property of the pipeline, not of any stage, which is why adding stages does not fix it.
- Artifact-specific contracts: different checks for different artifact types, applied throughout. A scene specification and a task instruction fail differently, so a single end-of-pipeline check cannot catch both.
- The dependency graph as the prerequisite for rollback: knowing what depended on what. Without it, repair has no target other than the whole build.
- Provenance-driven local repair: recompute locally where possible, escalate upstream only when necessary. It is what makes the loop affordable rather than a restart on every failure.
- Typed and reusable skills: declared interfaces for composition checks, and reuse across benchmarks. It makes both composition and repair checkable at the seams.
- Evaluating the mechanism, not only the products: ablations on verification and repair, and analyses of localised recovery and skill reuse.
Framework Shift
Before (stages automated, artifacts unverified):
agentic systems automate isolated stages, or predefined families
-> intermediates pass downstream without artifact-specific checks
-> a local defect becomes a property of the final benchmark
-> repair has no target short of rebuilding
After (closed-loop synthesis):
forward: typed, reusable skills composed into workflows, with an
artifact dependency graph
backward: artifact-specific contracts throughout, provenance-driven
local re-execution or upstream rollback
-> six offline benchmarks plus one interactive benchmark of 220 tasks
-> benchmarks discriminate capability; ablations validate the mechanism
-> repair is localised and skills are reused across benchmarks
From automating the stages of benchmark construction, to closing the loop between synthesis and verification, the core shift is that a defect should be caught at the artifact that produced it rather than discovered as a property of the finished benchmark.
Expert Assessment
Problem choice: Excellent, and the diagnosis identifies a failure mode that is easy to create and hard to notice. Automated pipelines naturally pass intermediates downstream, so naming propagation — rather than stage coverage — as the more important obstacle redirects the design question from “what stages do we need” to “where do we verify”.
Method maturity: The design is coherent because the two directions depend on each other: the artifact dependency graph in the forward pass is what makes provenance-driven rollback possible in the backward one, so the architecture is not two features bolted together. Making skills typed so composition is checkable rather than assumed is the right kind of precondition for automated assembly. And the evaluation order is well conceived: first showing the artifacts discriminate capability, then ablating the verification and repair mechanism, then demonstrating that repair is localised and skills reusable — which tests the design’s own claims rather than only its outputs.
Experimental integrity: The ablation of verification and repair is the most valuable element, because it tests the paper’s mechanism rather than its products, and the localised-recovery and skill-reuse analyses are each tied to a specific design promise. Constructing both offline and interactive benchmarks is a fair test of the generalisation claim rather than a single family. The limitation is that the evaluation is largely over the framework’s own outputs, so an independent party constructing a benchmark with the tool would be the stronger test of the reuse claims.
Writing quality: The formulation is precise and each component follows from it, which makes an architecture with several named parts legible. Because the practical claim is about repair cost, a short worked example of one verification failure and its resulting scope of recomputation would make the efficiency argument concrete.
Verdict: strong accept — it names the propagation failure that multi-step construction creates, designs a loop in which verification is artifact-specific and repair is provenance-scoped, and validates the mechanism through ablation rather than only through its outputs.
Takeaways
- Verify each artifact, not just the finished product. A defect caught at the end is a property of the benchmark rather than a fixable stage error.
- Make checks artifact-specific. Different artifact types fail differently, and one end-of-pipeline check cannot cover them.
- Record provenance if you want local repair. Rolling back to the right place requires knowing what depended on what.
- Ablate the mechanism, not only the product. Showing that verification and repair earn their cost is a different claim from showing the benchmarks are good.
论文: 2609.13082 作者: Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Danyang Li, Zheng Yang, Jianwei Hu, Qiang Ma 分类: cs.AI
缺口
智能体系统是自动化「具身基准构建」的一条有希望的路,而既有做法在两个彼此不同的方面不到位。第一是覆盖:它们通常只覆盖孤立的阶段,或者仍局限于预先定义的环境与任务族。
第二被陈述为更重要的那一个,而它是流水线的结构性属性、而不是任一阶段的局限:多步构建会产生彼此依赖的中间产物,而这些产物常常「未经产物特定的验证」就流向下游,使得局部缺陷传播进「成品基准」。
这个区分之所以要紧,是因为两种失效需要不同的补救。覆盖缺口可以通过增加阶段或拓宽环境集来补。而传播不能靠”增加任何东西”来补——它需要阶段之间的验证,而且具体需要产物特定的验证,因为”场景规格里的缺陷”与”任务指令里的缺陷”并不能被同一种检查发现。
自动化「具身基准构建」
智能体系统在这件事上有希望
而既有做法在「两个彼此不同」的方面不到位
[1] 「覆盖」
它们只覆盖「孤立的阶段」
或者仍局限于「预先定义的环境与任务族」
[2] 「传播」(被陈述为「更重要」的那一个)
「多步构建会产生彼此依赖的中间产物,
而这些产物常常「未经产物特定的验证」就流向下游」
-> 使得「局部缺陷传播进「成品基准」」
<- 这是「流水线的结构性属性」,
而不是任一「阶段」的局限
[「两种失效需要不同的补救」]
覆盖缺口可通过「增加阶段」或「拓宽环境集」来补
而传播「不能靠"增加任何东西"来补」
-> 需要「阶段之间」的验证
-> 而且具体需要「产物特定」的验证:因为
"场景规格里的缺陷"与"任务指令里的缺陷"
「并不能被同一种检查发现」
增量
一句话: 在这篇论文之前,多步基准构建把中间产物未经验证地传向下游;在这篇论文之后,闭环合成对每个产物按其自身契约验证、并依据来源回滚,产出七个能区分模型能力的基准。
核心机制
框架把用户指定的评测意图转化为完整的具身基准产物,而它的组织性想法被表述为一个形式化命题:把构建表述为「闭环基准合成」,把「前向的产物合成」与「后向的验证与修复」结合起来。 “闭环”才是让它不同的地方——合成与验证是穿过同一条流水线的两个方向,而不是”一条流水线之后再在末尾加一次检查”。
前向是「技能编排的产物合成」:它把「有类型的、可复用的技能」组合成可执行工作流,同时一张产物依赖图记录中间输出及其依赖关系。 有两条性质是设计进去的。“有类型”意味着技能的接口是被声明的,因此组合可以被检查、而不是被假定。而依赖图是后向方向的结构性前提:不知道”什么依赖于什么”,你就无法回滚到正确的位置。
后向是「需求引导的验证与修复」:它在整个构建过程中施加「产物特定的契约」,并在验证失败时利用「来源」触发局部重执行或上游回滚。 这两个机制回应的是不同的需要。“整个过程都施加产物特定的契约”意味着缺陷在它发生的那个产物处就被抓到,而不是在末尾——那里剩下的只是一个坏基准。而来源驱动的回滚才是让修复高效的东西:失败时,能在局部重算就在局部重算,只有在必须时才向上游升级——这一点要紧,因为一条”唯一补救就是从头再来”的流水线无法在规模上运行。
结果覆盖了两种赛道。 在 Offline EQA 赛道构建了六个基准,覆盖多样的具身场景;另有一个互动基准,包含 Interactive Embodied 赛道中 220 个可执行任务。两者兼具对主张很重要:离线评测检验基于观察的理解,而互动执行检验闭环行为——因此把两者都构建出来,说明这个框架并不专门化于某一阶段或某一环境族。
而评测是从「产物」排到「过程」的。 对代表性的 MLLM 与具身智能体做了评测,显示这些基准在”基于观察的理解”与”闭环执行”两方面都能区分模型能力——也就是说这些产物是可用的测量仪器,而不只是被构建出来的输出。接着质量评估与消融验证了基准质量、以及验证与修复的有效性——这检验的是论文自己的机制,而不只是它的产物。而修复与技能复用分析展示了高效的局部恢复与跨基准复用——正是这个设计被造出来要交付的两条性质:局部而非全局的修复,以及技能被跨基准复用而不是重建。
组织性想法,被表述为一个形式化命题
「把构建表述为「闭环基准合成」,
把「前向的产物合成」与「后向的验证与修复」结合起来」
<- "闭环"才是让它不同的地方:「合成与验证」是
「穿过同一条流水线的两个方向」,
而不是"一条流水线之后再在末尾加一次检查"
[前向:「技能编排的产物合成」]
把「有类型的、可复用的技能」组合成可执行工作流
一张「产物依赖图」记录中间输出及其依赖关系
<- "有类型":技能的「接口是被声明的」,
因此组合「可以被检查」、而不是被假定
<- 「依赖图」是后向方向的「结构性前提」:
不知道"什么依赖于什么",你「无法回滚到正确的位置」
[后向:「需求引导的验证与修复」]
在整个构建过程中施加「产物特定的契约」
并在验证失败时利用「来源」触发局部重执行或上游回滚
<- "整个过程都施加产物特定的契约":缺陷在「它发生的那个
产物处」就被抓到,而不是在末尾——那里剩下的只是一个坏基准
<- 「来源驱动的回滚」才是让修复「高效」的东西:
失败时能在局部重算就在局部重算,
「只有在必须时才向上游升级」
<- 要紧因为:一条"唯一补救就是从头再来"的流水线
「无法在规模上运行」
[结果覆盖了两种赛道]
OFFLINE EQA 赛道「六个」基准,覆盖「多样的具身场景」
另有「一个互动基准」,包含 INTERACTIVE EMBODIED 赛道中
「220 个可执行任务」
<- 「两者兼具」对主张很重要:离线评测检验「基于观察的理解」,
而互动执行检验「闭环行为」
-> 把两者都构建出来,说明这个框架「并不专门化于某一阶段
或某一环境族」
[评测从「产物」排到「过程」]
对代表性的 MLLM 与具身智能体做了评测
-> 这些基准在"基于观察的理解"与"闭环执行"两方面
「都能区分模型能力」
<- 这些产物是「可用的测量仪器」,而不只是被构建出来的输出
质量评估与消融验证了基准质量、「以及验证与修复的有效性」
<- 检验的是论文「自己的机制」,而不只是它的「产物」
修复与技能复用分析展示了「高效的局部恢复」与「跨基准复用」
<- 正是这个设计「被造出来」要交付的两条性质:
「局部」而非全局的修复,以及技能被「跨基准复用」
而不是「重建」
可以用**“一条带「工序检」的装配线,而不是只在装卸口做终检”来理解这件事: 终检只能抓到到最后还看得见的问题;而到那时,一个钻错的孔已经是成品的产品属性**,而不是一个可修的工位错误。 论文版本里有两个细节对得上这套机械。 各工位有各自的检查清单——场景规格与任务指令按不同方式受检——这就是”产物特定的契约”的意思。 而每个零件都带着”它从哪来”的记录,因此一次失败会把活儿送回生产它的那个工位,而不是整条线重开。 结果是:缺陷停在它的工位上,而不会变成成品基准的属性。
关键概念
- 以「传播」作为结构性失效: 未经验证的中间产物流向下游。它是流水线的属性、而不是任一阶段的属性——这正是”增加阶段”修不好它的原因。
- 产物特定的契约: 对不同产物类型用不同检查,且贯穿整个过程。场景规格与任务指令的失败方式不同,所以在流水线末尾做一次检查覆盖不了两者。
- 以依赖图作为回滚的前提: 知道”什么依赖于什么”。没有它,修复除了”整体重建”之外没有目标。
- 来源驱动的局部修复: 能局部重算就局部重算,只在必须时向上游升级。正是它让这个闭环负担得起,而不是每次失败都重开。
- 有类型、可复用的技能: 为组合检查提供已声明的接口,并跨基准复用。它让组合与修复在接缝处都可被检查。
- 检验机制,而不只是产物: 对验证与修复做消融,并分析局部恢复与技能复用。
框架转变
之前(阶段被自动化,产物未验证):
智能体系统自动化孤立阶段,或预定义的任务族
-> 中间产物未经产物特定的检查就流向下游
-> 一个局部缺陷成为成品基准的属性
-> 修复除了"重建"之外没有目标
之后(闭环合成):
前向:有类型、可复用的技能被组合成工作流,
并带一张产物依赖图
后向:贯穿始终的产物特定契约,
来源驱动的局部重执行或上游回滚
-> 六个离线基准 + 一个含 220 个任务的互动基准
-> 基准能区分能力;消融验证了机制
-> 修复是局部的,技能跨基准复用
从”把基准构建的各个阶段自动化”,转变为”在合成与验证之间闭合回路”,核心转变在于:缺陷应当在「产生它的那个产物」处被抓到,而不是作为「成品基准的属性」被发现。
专家评审
选题眼光: 极好,而这个诊断识别出的失效模式容易产生、却难以察觉。 自动化流水线天然会把中间产物传向下游;因此把传播——而不是阶段覆盖——点名为更重要的障碍,把设计问题从”我们需要哪些阶段”转向”我们在哪里做验证”。
方法成熟度: 设计之所以自洽,是因为两个方向彼此依赖:前向的那张产物依赖图,正是后向来源驱动回滚得以可能的原因——所以这个架构不是两个硬拼在一起的功能。把技能做成有类型的、使组合可被检查而非被假定,是自动装配所需的那种前提。而评测的顺序构思得好:先表明产物能区分能力,再对验证与修复机制做消融,最后展示”修复是局部的、技能可复用”——这检验的是设计自己的主张,而不只是它的输出。
实验诚意: 对验证与修复的消融是最有价值的一环,因为它检验的是论文的机制、而不是它的产物;而局部恢复与技能复用分析各自对应一条具体的设计承诺。同时构建离线与互动基准,是对”泛化”主张的公平检验、而不是只测单一族。 主要局限是:评测在很大程度上是对框架自身输出进行的,因此”由独立一方用这个工具去构建一个基准”会是检验那些复用主张的更强测试。
写作功力: 那个形式化命题很精确,且每个组件都由它推出——这让一个含有多个命名部件的架构变得可读。 由于实际主张关乎修复成本,若能给一个简短实例——一次验证失败、以及由此产生的重算范围——会让效率论证变得具体。
判决: 强接收(Strong Accept) — 它点出了多步构建所造成的「传播」失效,设计了一个”验证产物特定、修复由来源界定范围”的闭环,并通过消融而不只是通过输出验证了这个机制。
要点总结
- 逐个产物验证,而不只是验证成品。在末尾才抓到的缺陷,是基准的属性,而不是可修的工位错误。
- 让检查产物特定。不同产物类型的失败方式不同;流水线末尾的一次检查覆盖不了它们。
- 想要局部修复,就要记录来源。回滚到正确的位置,需要知道”什么依赖于什么”。
- 对机制做消融,而不只是对产物。证明”验证与修复值回它的成本”与”基准是好的”是两个不同的主张。