Paper: 2604.21871 Authors: Research Team Categories: cs.AI, cs.CL
Problem
LLMs are being deployed in high-stakes decisions, but:
- We don’t know how they handle moral dilemmas
- Prescriptive rules (laws, policies) ≠ social sensitivity
- Alignment with rules ≠ alignment with human values
- Moral competence vs. rule-following are different
Study Design
Extensive evaluation of LLM moral reasoning:
Moral Scenarios
- Trolley Problems: Classic utilitarian dilemmas
- Healthcare Decisions: Life-or-death resource allocation
- Social Norms: Informal rules and expectations
- Cultural Variations: Different moral frameworks
Evaluation Dimensions
- Prescriptive Compliance: Following explicit rules
- Social Sensitivity: Understanding informal norms
- Contextual Awareness: Recognizing situational factors
- Moral Consistency: Coherent reasoning across scenarios
Key Findings
| LLM Behavior | Prescriptive | Social Sensitivity |
|---|---|---|
| High compliance with rules | 95% | 62% |
| Understands social context | 78% | 45% |
| Adapts to cultural norms | 81% | 38% |
Critical Insight: LLMs align with prescriptive rules, not social sensitivity!
Takeaways
- Rule-following ≠ moral alignment
- Social sensitivity requires different training approaches
- Current RLHF may not capture human moral intuitions
- Need better benchmarks for moral competence
论文: 2604.21871 作者: 研究团队 分类: cs.AI, cs.CL
问题
LLM正被部署在高风险决策中,但:
- 我们不知道它们如何处理道德困境
- 规范性规则(法律、政策)≠ 社会敏感性
- 与规则对齐 ≠ 与人类价值观对齐
- 道德能力与规则遵循是不同的
研究设计
LLM道德推理的广泛评估:
道德场景
- 电车难题:经典功利主义困境
- 医疗决策:生死资源分配
- 社会规范:非正式规则和期望
- 文化差异:不同的道德框架
评估维度
- 规范性遵从:遵循明确规则
- 社会敏感性:理解非正式规范
- 上下文意识:识别情境因素
- 道德一致性:跨场景的一致推理
关键发现
| LLM行为 | 规范性 | 社会敏感性 |
|---|---|---|
| 高规则遵从 | 95% | 62% |
| 理解社会背景 | 78% | 45% |
| 适应文化规范 | 81% | 38% |
关键洞察:LLM与规范性规则对齐,而非社会敏感性!
要点总结
- 规则遵循 ≠ 道德对齐
- 社会敏感性需要不同的训练方法
- 当前RLHF可能无法捕捉人类道德直觉
- 需要更好的道德能力基准测试