Paper: 2609.05401 Authors: Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No Categories: cs.RO, cs.CL
The Gap
Vision-language models (VLMs) are increasingly being promoted from passive perceptual modules to automated evaluators and reward functions for robotic manipulation. In reinforcement learning or trajectory ranking, a VLM reward model takes a video of a robot’s execution and a natural language goal, outputting a scalar progress score or success indicator. For this paradigm to function reliably, it demands a fundamental invariant: paraphrase invariance. If a robot is asked to “Pick radish, then place in pink bowl,” its trajectory should receive the exact same evaluation if the prompt is reworded to “After picking radish, place in pink bowl.”
Prior evaluations of VLM reward models tested accuracy against fixed task strings or simple binary success classifications. They ignored linguistic perturbations. When an agent policy optimizes against a reward model that is sensitive to superficial syntax, the policy does not learn robust motor control; it learns to exploit phrasing artifacts.
This paper establishes the extent of that failure mode. It shows that current state-of-the-art VLMs regularly contradict themselves on identical trajectories under human-verified paraphrases, turning verifiable task completions into failures and vice versa.
[SETTING] Robot Trajectory + Instruction -> VLM Reward Model -> Progress Score
|
v
[CURRENT ASSUMPTION] If semantic intent is preserved, reward score is stable
|
v
[TEST OF INVARIANCE] Evaluate identical video across verified paraphrases
|
+-------------------------+-------------------------+
v v v
[Lexical Synonym] [Syntactic Reorder] [Action-Goal Inversion]
"pink bowl" -> "magenta" "Pick X, then Y" -> "Place X in Y" ->
"After picking X, do Y" "Move X until in Y"
| | |
+-------------------------+-------------------------+
v
[OBSERVED REALITY] Contradiction in up to 43% of cases;
identical trajectory scores flip between 1 (failed) and 5 (completed);
RL policies trained on fragile rewards collapse.
The Increment
One sentence: Before this paper, VLM reward models were treated as plug-and-play evaluators with known visual blind spots; after it, they are proven to suffer from severe linguistic fragility, where minor syntactic rewrites invert reward signals on identical physical executions.
Core Mechanism
The authors formalize the vulnerability through RoboRMBench, a benchmark built on 2,390 real-robot trajectories across diverse manipulation tasks, paired with 21,673 human-verified paraphrases. These paraphrases span three operational tiers:
- Lexical variations: substituting object names, color descriptors, or container attributes with synonyms.
- Syntactic variations: altering clause structure, such as changing imperative sequences (“Do A then B”) into temporal subordinate clauses (“After doing A, perform B”).
- Action-goal variations: rewriting procedural commands into goal-state verification queries (“Ensure X is inside Y”).
To diagnose reward models, the paper introduces two primary fragility metrics alongside standard Mean Error (ME):
- Score Contradiction Rate (SCR): The proportion of paraphrase pairs where the reward model assigns divergent scores to the same trajectory that violate basic monotonicity.
- Flip Rate (FR): The percentage of cases where rewording an instruction causes the evaluated outcome to flip between failure (
score <= 2) and completion (score >= 4).
ROBORMBENCH DIAGNOSTIC ARCHITECTURE
[Real-Robot Video Trajectory: tau]
|
+--------+--------+
| |
v v
[Prompt g_1] [Prompt g_2] (Paraphrase)
"Pick radish..." "After picking radish..."
| |
+--------+--------+
|
v
[VLM Reward Model R_theta]
|
+--------+--------+
| |
v v
Score r_1 = 5 Score r_2 = 1 <-- CONTRADICTION!
| |
+--------+--------+
|
v
[Metric Suite: SCR, FR, and Preference Regret]
To explain why this happens, consider a structural metaphor of a courtroom stenographer acting as a judge. The judge is asked to grade a defendant’s compliance with an order. Instead of looking directly at whether the act occurred in the courtroom video footage, the judge parses the grammar of the indictment. If the clerk writes “The subject opened the vault and took the ledger,” the judge awards full marks. If the clerk writes “Prior to taking the ledger, the subject opened the vault,” the judge gets confused by the inverted clause timing, glances at the exact same surveillance tape, and rules that the subject failed to follow the order.
Key Concepts
- Paraphrase Fragility: The vulnerability of multi-modal reward models where superficial syntactic or lexical rewrites of a command induce massive swings in predicted progress scores on identical video input.
- Score Contradiction Rate (SCR): A pairwise metric quantifying how often a model rates trajectory A higher than trajectory B under prompt 1, but rates trajectory B higher than trajectory A under prompt 2.
- Semantic Equivalence Filter: An automated verification pipeline combined with human cross-auditing to guarantee that paraphrased prompts share 100% semantic identity in the robotics domain before benchmark inclusion.
Framework Shift
Before (Standard VLM Reward Evaluation):
[Trajectory] + [Single Fixed Prompt] ---> [VLM Judge] ---> [Reward Scalar]
(Blind to linguistic variance; conflates visual error with prompt sensitivity)
After (RoboRMBench Invariance Protocol):
[Trajectory] + [Paraphrase Set G] ---> [VLM Judge] ---> [Reward Vector]
|
+------------------+------------------+
v v
[Accuracy (ME)] [Invariance (SCR/FR)]
From evaluating a single prompt-response point to evaluating the reward distribution across the semantic equivalence class, the core shift is exposing how reward models hallucinate task failure purely from linguistic variation.
Expert Assessment
Problem choice: Real and urgent. As robotics shifts from hardcoded reward functions to foundation model scoring, uninspected linguistic fragility directly pollutes policy gradients and offline model selection.
Method maturity: The construction of RoboRMBench is rigorous. Collecting 2,390 physical robot executions and verifying 21,673 paraphrases with human cross-checks prevents the benchmark itself from being polluted by semantic drift.
Experimental integrity: The authors test both leading proprietary models (GPT-4o, Gemini 1.5 Pro) and open-source models (Qwen2-VL, InternVL2, LLaVA-OneVision). The findings are consistent across architectures: even frontier models exhibit flip rates exceeding 25% on syntactic reorderings.
Writing quality: The distinction between lexical, syntactic, and action-goal rewrites is cleanly maintained throughout all ablation tables.
Verdict: strong accept — A benchmark that catches a blind spot in robotic foundation models before community bad habits harden into standard practice.
Takeaways
- Never evaluate a robotic policy or reward model against a single phrasing of a goal. Always test over a semantic equivalence cluster.
- When fine-tuning VLMs for reward modeling, syntactic augmentation is just as critical as visual data augmentation.
- In policy deployment, ensemble-averaging reward outputs over 3-5 deterministic syntactic variations acts as an immediate test-time defense against reward hacking.
论文: 2609.05401 作者: Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No 分类: cs.RO, cs.CL
缺口
在具身智能与机器人学习中,视觉语言模型(VLM)正逐渐从单纯的视觉感知模块,升级为策略训练中的自动化奖励函数(Reward Model)。 在强化学习探索或离线轨迹排序中,VLM 负责观察机器人的动作视频并比对指令文本,输出一个进度标量或成功概率。 这种机制能够成立,必须依赖一个不可妥协的前提:同义改写不变性(Paraphrase Invariance)。 如果机器人完成了“先捡起萝卜,再放入粉色碗中”,那么将指令换成“捡起萝卜之后,放到粉色碗里”,模型给出的奖励必须完全一致。
以往对 VLM 奖励模型的评测,几乎全部局限在单一固定指令模板下的准确率或二元成功判断,完全忽略了语言层面的扰动。 如果底层奖励函数对表层句式极度敏感,下游策略学到的就根本不是稳健的物理操作,而是利用提示词漏洞进行的“奖励黑客(Reward Hacking)”。
这篇论文正是为了刺破这一盲区而生。 研究团队证实:即便是最先进的专有与开源 VLM,在面对人工核验的同义改写时,也会在完全相同的真实机器人视频上给出截然相反的评分,将完成判定为彻底失败。
[场景] 机器人执行视频 + 目标指令 -> VLM 奖励模型 -> 进度评分
|
v
[既有假设] 只要语义一致,奖励评分就保持稳定
|
v
[不变性压力测试] 对完全相同的执行视频输入不同同义句
|
+-------------------------+-------------------------+
v v v
[词汇同义替换] [句式顺序倒置] [动作-目标态重构]
"粉色碗" -> "洋红色碗" "先做A再做B" -> "把A放进B" ->
"做完A后执行B" "移动A直到其位于B内"
| | |
+-------------------------+-------------------------+
v
[实测事实] 评分矛盾率高达 43%;
同一条轨迹在 1 分(失败)与 5 分(成功)之间剧烈震荡;
基于此类奖励训练的强化学习策略迅速崩溃。
增量
一句话: 在这篇论文之前,VLM 奖励模型被当作仅有视觉盲区的即插即用评估器;在这篇论文之后,学界证实其存在严重的语言脆弱性——仅仅改变从句先后顺序,就能在完全相同的物理轨迹上将奖励信号彻底颠倒。
核心机制
研究团队构建了涵盖 2,390 条真实机器人操控轨迹与 21,673 条人工校验同义句的基准数据集 RoboRMBench。 这些同义改写被系统化解构为三个层级:
- 词汇层改写:替换物体名、颜色属性或容器名称为常见同义词。
- 句法层改写:调整从句结构,例如将祈使连词(“执行 A 然后 B”)变为时间状语从句(“在执行 A 之后,进行 B”)。
- 动作-目标态改写:将动作过程描述重写为终态约束验证(如“确认 A 已经进入 B”)。
为了量化这种语义脆弱性,论文提出了两个核心诊断指标:
- 评分矛盾率(SCR, Score Contradiction Rate):在同义改写提示词对下,模型对同一轨迹输出破坏单调性评分的比例。
- 翻转率(FR, Flip Rate):仅仅因为指令句式重写,导致同一轨迹在“失败”(<=2分)与“成功”(>=4分)之间发生定性逆转的概率。
ROBORMBENCH 评测拓扑
[真实机器人轨迹视频: tau]
|
+-------+-------+
| |
v v
[指令 g_1] [同义改写 g_2]
"先抓萝卜再入碗" "抓完萝卜后放入碗中"
| |
+-------+-------+
|
v
[VLM 奖励模型 R_theta]
|
+-------+-------+
| |
v v
评分 r_1 = 5 评分 r_2 = 1 <-- 产生恶性矛盾!
| |
+-------+-------+
|
v
[指标体系: SCR 矛盾率, FR 颠倒率, 策略后悔值]
用一个法庭书记员兼法官的核喻来理解这个机制: 法官负责审查监控录像,判断嫌疑人是否执行了特定指令。 但他不是把注意力集中在录像里的物理动作上,而是在死抠起诉书的措辞。 当控方写“嫌疑人打开了保险箱,取走了账本”,法官打出满分。 当辩方同义改写为“在取走账本之前,嫌疑人打开了保险箱”,法官被颠倒的句式语法彻底绕晕,扫了一眼同一段监控,断定被告根本没做这件事。
关键概念
- 同义词脆弱性(Paraphrase Fragility):多模态奖励模型对输入文本的表层句式极为敏感,在视觉信号完全恒定的情况下产生剧烈评分波动。
- 评分矛盾率(SCR):用于衡量模型在不同同义提示词下,对成对轨迹优劣关系判断发生逆转的频率。
- 语义等价过滤(Semantic Equivalence Filter):结合大模型初步生成与双盲人工交叉复审的清洗机制,确保基准中的所有改写在物理世界中具备绝对同等含义。
框架转变
之前(传统 VLM 奖励模型评测):
[执行轨迹] + [单一固定指令] ---> [VLM 评判器] ---> [标量奖励值]
(忽略语言变化维度,将语言敏感性与视觉识别错误混为一谈)
之后(RoboRMBench 不变性评测框架):
[执行轨迹] + [同义改写集合 G] ---> [VLM 评判器] ---> [奖励输出分布]
|
+-----------------------+-----------------------+
v v
[基准准确度 (ME)] [不变性鲁棒度 (SCR/FR)]
从孤立测试单点准确率,转变为评估模型在语义等价类空间内的奖励分布收敛性,核心转变在于将“表层语言过拟合”从视觉能力中彻底剥离出来进行审计。
专家评审
选题眼光: 极为毒辣且切中要害。 具身智能领域正在大规模采用 VLM 进行无监督强化学习与轨迹重排,如果底层奖励信号本身充斥着未被察觉的句式噪声,训练出的策略必然存在巨大隐患。
方法成熟度: RoboRMBench 的构建流程十分扎实。 收集 2,390 条实体机器人操作轨迹,并结合两道人工校验确保 21,673 条改写的严格等价,彻底规避了以往自动化数据合成中的语义漂移缺陷。
实验诚意: 评测覆盖了 GPT-4o、Gemini 1.5 Pro 等商业巨头模型以及 Qwen2-VL、InternVL2、LLaVA-OneVision 等开源前沿。 实验数据令人警醒:即便顶尖商业模型,在复杂句式调整下的评分翻转率也普遍超过 25%。
写作功力: 词汇、句法与目标态三种改写梯度的划分极为清晰,实验图表直接有力。
判决: 强接收(Strong Accept) — 为机器人基座模型评测敲响了警钟,提供了急需的标准工具。
要点总结
- 永远不要用单一措辞的 prompt 评估机器人策略的表现或训练强化学习奖励模型,必须在同义等价集上测试分布。
- 在对 VLM 进行奖励微调时,文本端的句式对抗增强与视觉端的数据增强同等关键。
- 在实际工程推理阶段,对 3 到 5 个同义变体输入进行奖励打分并取中位数或平均值,是防范奖励黑客最低成本的防御方案。