Paper: 2604.19638 Authors: Josue Torres-Fonseca, Naihao Deng, et al. (University of Michigan, Boise State) Categories: cs.AI, cs.CV
Problem
Multimodal LLMs are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient.
Existing safety benchmarks (ASIMOV, MM-SafetyBench) focus on hazard recognition through question-answering on static images/videos - a critical gap remains in evaluating action-oriented safety.
SafetyALFRED Benchmark
Built on ALFRED embodied agent benchmark, augmented with six categories of real-world kitchen hazards:
- Sharp object handling
- Hot surface awareness
- Liquid spillage prevention
- Chemical safety
- Fall prevention
- Electrical hazards
Key Finding: Alignment Gap
Evaluated 11 state-of-the-art models (Qwen, Gemma, Gemini families):
- Models can accurately recognize hazards in QA settings
- Average mitigation success rates are low in embodied planning contexts
Static QA evaluations are insufficient for physical safety because they don’t test whether models can take corrective actions.
Takeaways
- Need paradigm shift toward benchmarks that prioritize corrective actions in embodied contexts
- Hazard recognition ≠ Hazard mitigation
- Future safety evaluation must test action, not just perception
论文: 2604.19638 作者: Josue Torres-Fonseca, Naihao Deng等(密歇根大学、博伊西州立大学) 分类: cs.AI, cs.CV
问题
多模态LLM越来越多地被用作交互环境中的自主代理,但它们主动应对安全危险的能力仍然不足。
现有的安全基准(ASIMOV、MM-SafetyBench)专注于通过问答形式在静态图像/视频上识别危险——在评估行动导向安全方面存在关键差距。
SafetyALFRED基准
建立在ALFRED具身代理基准之上,增加了六类现实厨房危险:
- 锋利物体处理
- 高温表面意识
- 液体溢出预防
- 化学品安全
- 跌倒预防
- 电气危险
关键发现:对齐差距
评估了11个SOTA模型(Qwen、Gemma、Gemini系列):
- 模型能够在QA设置中准确识别危险
- 在具身规划环境中缓解成功率平均较低
静态QA评估对于物理安全是不够的,因为它们不测试模型是否能采取纠正行动。
要点总结
- 需要向优先考虑具身环境中纠正行动的基准进行范式转变
- 危险识别 ≠ 危险缓解
- 未来安全评估必须测试行动,而不仅仅是感知