Paper: 2604.19638 Authors: Josue Torres-Fonseca, Naihao Deng, et al. (University of Michigan, Boise State) Categories: cs.AI, cs.CV

Problem

Multimodal LLMs are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient.

Existing safety benchmarks (ASIMOV, MM-SafetyBench) focus on hazard recognition through question-answering on static images/videos - a critical gap remains in evaluating action-oriented safety.

SafetyALFRED Benchmark

Built on ALFRED embodied agent benchmark, augmented with six categories of real-world kitchen hazards:

  1. Sharp object handling
  2. Hot surface awareness
  3. Liquid spillage prevention
  4. Chemical safety
  5. Fall prevention
  6. Electrical hazards

Key Finding: Alignment Gap

Evaluated 11 state-of-the-art models (Qwen, Gemma, Gemini families):

  • Models can accurately recognize hazards in QA settings
  • Average mitigation success rates are low in embodied planning contexts

Static QA evaluations are insufficient for physical safety because they don’t test whether models can take corrective actions.

Takeaways

  • Need paradigm shift toward benchmarks that prioritize corrective actions in embodied contexts
  • Hazard recognition ≠ Hazard mitigation
  • Future safety evaluation must test action, not just perception

论文: 2604.19638 作者: Josue Torres-Fonseca, Naihao Deng等(密歇根大学、博伊西州立大学) 分类: cs.AI, cs.CV

问题

多模态LLM越来越多地被用作交互环境中的自主代理,但它们主动应对安全危险的能力仍然不足

现有的安全基准(ASIMOV、MM-SafetyBench)专注于通过问答形式在静态图像/视频上识别危险——在评估行动导向安全方面存在关键差距。

SafetyALFRED基准

建立在ALFRED具身代理基准之上,增加了六类现实厨房危险

  1. 锋利物体处理
  2. 高温表面意识
  3. 液体溢出预防
  4. 化学品安全
  5. 跌倒预防
  6. 电气危险

关键发现:对齐差距

评估了11个SOTA模型(Qwen、Gemma、Gemini系列):

  • 模型能够在QA设置中准确识别危险
  • 在具身规划环境中缓解成功率平均较低

静态QA评估对于物理安全是不够的,因为它们不测试模型是否能采取纠正行动。

要点总结

  • 需要向优先考虑具身环境中纠正行动的基准进行范式转变
  • 危险识别 ≠ 危险缓解
  • 未来安全评估必须测试行动,而不仅仅是感知