Paper: 2604.20652 Authors: Nattavudh Powdthavee Categories: cs.AI, cs.CL
Abstract
Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We tested this in a preregistered experiment across seven leading LLMs and twelve investment scenarios covering legitimate, high-risk, and objectively fraudulent opportunities, combining 3,360 AI advisory conversations with a 1,201-participant human benchmark.
Key Findings
Contrary to predictions:
- Motivated investor framing did NOT suppress AI fraud warnings (if anything, it marginally increased them)
- Endorsement reversal occurred in fewer than 3 in 1,000 observations
- Human advisors endorsed fraudulent investments at 13-14% baseline rate
- All LLMs: 0% endorsement of fraud
- Humans suppressed warnings under pressure at 2-4x the AI rate
Results
| Group | Fraud Endorsement Rate |
|---|---|
| Human Advisors | 13-14% |
| All LLMs | 0% |
AI systems currently provide more consistent fraud warnings than lay humans in an identical advisory role.
Takeaways
- LLMs resist motivated reasoning better than humans in fraud detection
- Even when investors are “persuaded,” AI maintains fraud warning consistency
- Model behavior under pressure is crucial for safety-critical applications
- Human-like preferences (RLHF) don’t undermine AI’s ability to detect fraud
论文: 2604.20652 作者: Nattavudh Powdthavee 分类: cs.AI, cs.CL
摘要
基于人类反馈训练的大语言模型可能会在投资者已经被说服存在欺诈机会时抑制欺诈警告。我们在一个预注册实验中测试了这一点,涉及7个领先LLM和12个投资场景(包括合法、高风险和客观欺诈机会),结合3360次AI咨询对话与1201名人类参与者作为基准。
关键发现
与预测相反:
- 有动机的投资者框架没有抑制AI欺诈警告(如果有任何影响,反而略有增加)
- 支持反转发生在少于千分之三的观察中
- 人类顾问欺诈支持率为13-14%的基线
- 所有LLM:0%欺诈支持率
- 人类在压力下抑制警告的比率是AI的2-4倍
实验结果
| 组别 | 欺诈支持率 |
|---|---|
| 人类顾问 | 13-14% |
| 所有LLM | 0% |
AI系统在相同的咨询角色中提供的欺诈警告比普通人类更加一致。
要点总结
- 在欺诈检测中,LLM比人类更能抵制动机推理
- 即使投资者已被”说服”,AI仍保持欺诈警告一致性
- 压力下的模型行为对安全关键应用至关重要
- 类人偏好(RLHF)不会削弱AI检测欺诈的能力