Paper: 2608.06151 Authors: Thomas H. Costello, Nathaniel Rabb, Michael Nicholas Stagnaro, Gordon Pennycook, David Rand Categories: cs.HC, cs.AI
The Gap
There is now solid evidence that conversational dialogue with an LLM can reduce belief in conspiracy theories. That finding was surprising when it landed, because it cut against the dominant account of conspiracy belief as identity-protective and therefore immune to evidence. It turned out that tailored, patient, factually specific counter-argument moves people — and that an LLM can supply it at scale in a way a human interlocutor cannot.
But every demonstration of this ran on established conspiracy theories: the moon landing, 9/11, vaccines. That setting has a property doing quiet work in the result. An established conspiracy theory has an established rebuttal. Decades of journalism, investigation, and debunking sit in the training data, so the model is not reasoning its way to a counter-argument — it is retrieving a well-developed one and tailoring the delivery.
This leaves the case that actually matters untested. Conspiracy theories are most consequential and most contagious in the first days after a major public event, before facts are settled and while the interpretive frame is still being fixed. That window has none of the properties the earlier work relied on. There is no debunking literature. There may not yet be an authoritative account of what happened. The model’s own training data may predate the event entirely. And beliefs formed in that window are the ones that harden and propagate.
So the open question is not “can LLMs debunk conspiracies” — that is answered — but whether the effect survives the removal of the scaffolding that made it easy.
Established finding: LLM dialogue reduces belief
in well-known conspiracy theories
|
v
Unexamined enabling condition:
those theories have decades of debunking in training data
|
v
[This paper] Remove the scaffolding.
Test conspiracies forming in the days after a crisis
|
+-------------------------+
| |
v v
July 2024 Sept 2025
Trump assassination Kirk assassination
attempt (N = 472) (N = 1035)
| |
+-----------+-------------+
|
v
LLM dialogue vs irrelevant-topic LLM chat
vs static fact sheet
|
v
Result: significant belief reduction in BOTH
|
v
Follow-up 1-2 months later: reduced belief in
DIFFERENT conspiracies after LATER crisis events
The Increment
One sentence: Before this paper, we knew LLMs could talk people out of conspiracy theories that already had a debunking record; after it, we know the effect works on conspiracies that are still forming, and that it generalizes to other conspiracies months later.
Core Mechanism
The design is a field experiment run against the clock, and the timing is the methodological contribution. Both studies launched in the days immediately following a real crisis event — the July 2024 assassination attempt on Donald Trump, and the September 2025 assassination of Charlie Kirk. U.S. adults who held conspiratorial views about the event were recruited and assigned to a multi-turn conversation with an LLM prompted to reduce their conspiracy belief.
The control conditions are chosen to isolate two different alternative explanations, and this is where the design earns its keep. One control has participants discuss an irrelevant topic with an LLM. This absorbs everything about the experience that is not the argument: the novelty of talking to an AI, the attention and engagement, the experience of being listened to, any general reflective effect of a long conversation. The other control shows a static fact sheet — the same corrective information, delivered without dialogue. This isolates the interactive component specifically: if a fact sheet worked as well, the whole LLM apparatus would be an expensive way to deliver text. The treatment beat both, in both experiments.
The most interesting result is the one that was not the primary hypothesis. Participants were followed up one to two months later, in the wake of subsequent crisis events, and showed reduced belief in different conspiracy theories. This is a different kind of claim from the main effect. Belief in the target conspiracy going down is consistent with a narrow reading — the model supplied specific counter-evidence about a specific event, and it landed. Belief in unrelated later conspiracies going down is not explicable that way. Something transferred that was not event-specific content: some disposition, some raised threshold, some habit of interrogating a claim before adopting it.
Timeline of Experiment (both replications)
crisis event occurs
|
| days
v
recruit adults holding conspiratorial views
|
+----+--------------------+--------------------+
| | |
v v v
TREATMENT CONTROL A CONTROL B
multi-turn LLM LLM chat about static
dialogue, prompted irrelevant topic fact sheet
to reduce belief | |
| | |
| isolates: | |
| the ARGUMENT | |
| isolates: isolates:
| AI novelty + information
| engagement without dialogue
| | |
+----+--------------------+--------------------+
|
v
immediate measure: belief in target conspiracy
--> treatment significantly lower (both experiments)
|
| 1-2 months, NEW crisis events occur
v
follow-up measure: belief in DIFFERENT conspiracies
--> still reduced
Here is the structural metaphor. Think of conspiracy belief as an immune response to an ambiguous event. When something shocking and unexplained happens, the mind needs an account fast, and in the absence of a settled one it will manufacture something that fits the available fragments and the person’s priors. The account, once adopted, resists revision — that is the identity-protective story the field long believed.
Earlier LLM-debunking work was, in this metaphor, treating a chronic infection with a well-characterized drug. The conspiracy was old, the counter-arguments were thoroughly developed, and the model’s job was to administer a known therapy with unusual patience and personalization.
This paper does something different: it intervenes during the acute phase, before the belief has consolidated, using no pre-existing therapy because none exists yet. And the surprising finding is not just that intervening early works — it is that the effect looks less like a drug and more like a vaccine. Months later, facing a new pathogen — a different crisis, a different conspiracy — the treated participants respond less strongly. Whatever the conversation did, it did not merely neutralize one belief; it left something behind that engaged with the next one. That is why the downstream result is the paper’s most important sentence, and also its most fragile one.
Key Concepts
-
Unfolding versus established conspiracies: The distinction the whole paper rests on. An established conspiracy has three things an unfolding one lacks: a documented factual record, a body of existing rebuttals in the model’s training data, and a belief that has already hardened in the holder. Removing all three at once is a much harder test, and it is also the test that corresponds to real deployment — nobody needs an intervention against moon-landing denial with the urgency they need one in the 72 hours after a national shock. A concrete illustration of the difficulty: for the September 2025 event, a model may have no training data about the event whatsoever, so it cannot supply facts about what happened. Whatever it is doing must be more procedural than factual.
-
The two-control design: Why a single control would have been uninformative. Comparing LLM dialogue against nothing would confound the argument with the experience. People who spend twenty minutes in an engaged conversation, with a novel and attentive interlocutor, about a topic they care about, may moderate their views for reasons having nothing to do with what was argued — social desirability, reflection, or simple fatigue with the position. The irrelevant-topic control absorbs that. Separately, comparing against nothing would leave open that the model’s content was all that mattered, in which case a PDF would do. The fact-sheet control absorbs that. Beating both is what licenses the claim that interactive, tailored counter-argument is the active ingredient.
-
Downstream transfer: Reduced belief in different conspiracies, following different crisis events, one to two months later. The reason this is the most theoretically loaded result: it is inconsistent with the intervention working purely through content delivery. If the mechanism were “the model told me true things about event X,” there is no route by which that lowers belief about unrelated event Y months later. The result therefore points toward something dispositional — a shift in how readily a person adopts a conspiratorial frame at all. It is also the result most exposed to alternative explanations, since repeat participation in a study about conspiracy beliefs is itself an intervention.
Framework Shift
Before (prior LLM-debunking work): After (this paper):
established conspiracy conspiracy forming NOW
(decades old) (days old)
| |
v v
rich debunking record no debunking record;
in training data event may postdate
| training data
v |
model retrieves + tailors v
a known counter-argument model must work
| procedurally, not
v from stored rebuttals
belief in THAT theory drops |
| v
v belief drops -- AND
(transfer untested) transfers to OTHER
conspiracies months later
|
v
looks less like a cure,
more like a vaccine
From showing that LLMs can deliver mature counter-arguments to old conspiracy theories, to showing they can move beliefs in the hours when beliefs are actually being formed, the core shift is from debunking as retrieval to debunking as a real-time, and apparently durable, cognitive intervention.
Expert Assessment
Problem choice: Excellent, and the field-timing is the reason. The gap between “works on the moon landing” and “works on the thing everyone is arguing about right now” is exactly the gap that determines whether the earlier finding has any practical consequence, and it is a gap most researchers would leave open because closing it requires being ready to launch a study within days of an unpredictable event. Running it twice, on two separate real crises fourteen months apart, is not just replication — it is evidence that the team built standing capacity to do this, which is itself a contribution to how this research area can be done. The paper sits at the point where the LLM-persuasion literature stops being an interesting curiosity and starts being a policy question.
Method maturity: The two-control design is the right design and is executed with unusual discipline. Choosing an irrelevant-topic LLM conversation as one arm shows the authors anticipated the strongest deflationary reading of their own result, which is the mark of a serious experimental group. What the design cannot do — and this is inherent rather than a criticism of execution — is tell us which feature of the dialogue matters. The treatment is a bundle: personalization, patience, interactivity, specific evidence, non-judgmental tone. Beating a fact sheet establishes that dialogue beats text; it does not establish which property of dialogue is load-bearing, and that is what anyone trying to deploy or defend against this would most need to know.
Experimental integrity: Sample sizes are respectable and the near-tripling from Experiment 1 to Experiment 2 suggests the second was appropriately powered on the basis of the first. Self-reported belief measures remain the field’s standard and the field’s weakness: demand characteristics are a live concern when a participant has just spent twenty minutes with an interlocutor who was visibly trying to change their mind, and the natural response is to report movement. The irrelevant-topic control does not fully absorb this, since only the treatment arm makes the experimenter’s goal legible. The downstream result is where I would apply the most pressure. It is the paper’s most striking claim and it rests on a follow-up sample that has necessarily attrited, and where the participants have been primed by the original study to think about conspiracy beliefs. Attrition in this setting is unlikely to be random — people who found the conversation persuasive may be likelier to return. None of this makes the result wrong, but it is the claim that most needs a preregistered, adversarially designed replication before it is treated as established.
Writing quality: The abstract is clear and admirably restrained given how quotable the finding is; “shed light on the psychology of emerging conspiracies” is a more modest framing than the result would license, which is the right instinct. The corner cut is that the downstream transfer result is reported in a single sentence and carries a disproportionate share of the paper’s theoretical weight. A section that took the alternative explanations for that result seriously — attrition, priming, regression to the mean — would strengthen the paper more than any additional experiment.
Verdict: strong accept — The hard version of an important question, answered twice on real events with controls that anticipate the obvious objections; the durable-transfer claim needs more scrutiny than it gets, but the main result is solid and consequential.
Takeaways
Three things a practitioner can steal:
-
Two controls, one for the experience and one for the content. This design pattern transfers directly to evaluating any LLM intervention — tutoring, therapy-adjacent tools, customer support, behavior change. Run one arm where the model does something irrelevant (absorbing novelty, attention, and engagement effects) and one arm delivering the same information statically (absorbing the content). Only what beats both is attributable to interactive, tailored generation. Most product-side evaluations of LLM features run neither control and therefore cannot distinguish their model from a well-written FAQ.
-
Test the capability where the scaffolding is absent. The generalizable methodological move: identify the condition that made your earlier result easy — abundant training data, a settled ground truth, a well-documented answer — and construct the case where it is missing. This is the difference between a capability that benchmarks well and one that survives deployment, and it applies to retrieval systems, agents, and evaluation suites as much as to persuasion. Ask what your system is retrieving rather than reasoning, then remove the thing it retrieves.
-
Distinguish content effects from dispositional effects by looking for transfer. If an intervention only changes beliefs about the thing it argued about, it is a content effect and must be re-delivered for every new instance. If it changes beliefs about unrelated things later, something dispositional moved, and the economics are entirely different — one interaction with lasting spillover rather than perpetual per-case delivery. Measuring off-target outcomes at follow-up is cheap and is the only way to tell these apart.
论文: 2608.06151 作者: Thomas H. Costello, Nathaniel Rabb, Michael Nicholas Stagnaro, Gordon Pennycook, David Rand 分类: cs.HC, cs.AI
缺口
现在已有扎实证据表明,与大语言模型的对话能够降低人们对阴谋论的相信。 这个发现刚出来时令人意外,因为它反驳了当时主流的解释:阴谋论信念是身份保护性的,因而对证据免疫。 结果证明,量身定制的、有耐心的、事实具体的反驳确实能打动人——而 LLM 能以人类对话者做不到的规模提供这种反驳。
但所有这类论证跑的都是已成型的阴谋论:登月、9/11、疫苗。 这个设定里有一个性质在悄悄做工作。 一个已成型的阴谋论,配有一套已成型的反驳。 数十年的新闻调查、追查、辟谣都躺在训练数据里,所以模型并不是推理出一个反驳——它是检索出一个发育完整的反驳,然后调整递送方式。
这就让真正要紧的那种情形未被测试。 阴谋论最有后果、也最有传染性的时刻,是重大公共事件发生后的头几天,那时事实尚未落定,解释框架还在被固定。 那个窗口不具备此前工作所倚仗的任何性质。 没有辟谣文献。 甚至可能还没有一个关于「发生了什么」的权威说法。 模型的训练数据可能整个早于该事件。 而在那个窗口里形成的信念,恰恰是会硬化并传播开的那些。
所以真正悬着的问题不是「LLM 能不能破除阴谋论」——那已经有答案了——而是当那些让它变容易的脚手架被撤走后,效应还在不在。
已有结论:LLM 对话能降低对
知名阴谋论的相信
|
v
未被检视的使能条件:
那些阴谋论在训练数据里有几十年的辟谣积累
|
v
[本文] 撤走脚手架。
测试危机后数日内正在成形的阴谋论
|
+-------------------------+
| |
v v
2024 年 7 月 2025 年 9 月
特朗普遇刺未遂 Charlie Kirk 遇刺
(N = 472) (N = 1035)
| |
+-----------+-------------+
|
v
LLM 对话 vs 与 LLM 聊无关话题
vs 静态事实清单
|
v
结果:两次实验中信念均显著降低
|
v
1-2 个月后追访:在更晚发生的危机事件之后,
对**不同**阴谋论的相信也降低了
增量
一句话: 这篇论文之前,我们知道 LLM 能劝退那些已经有辟谣积累的阴谋论;这篇论文之后,我们知道这套效应在阴谋论尚在成形时也管用,而且数月之后能泛化到别的阴谋论上。
核心机制
这套设计是一场与时钟赛跑的实地实验,而时机本身就是方法论上的贡献。 两项研究都在真实危机事件发生后的数日内启动——2024 年 7 月对特朗普的暗杀未遂,以及 2025 年 9 月 Charlie Kirk 遇刺。 研究招募了对该事件持阴谋论看法的美国成年人,让他们与一个被提示去降低其阴谋论信念的 LLM 进行多轮对话。
对照条件的选择隔离了两种不同的替代解释,而这正是这套设计立身之处。 一组对照让参与者与 LLM 讨论一个无关话题。 这吸收掉了这段体验中除论证之外的一切:和 AI 说话的新鲜感、被投入的注意力、被倾听的体验、长对话本身带来的任何一般性反思效应。 另一组对照展示静态事实清单——同样的纠正信息,但不通过对话递送。 这专门隔离了互动成分:如果一张事实清单效果一样好,那整套 LLM 装置就只是一种昂贵的送文本方式。 在两次实验中,实验组都胜过了这两组。
最有意思的结果并不是主假设。 研究在 1 到 2 个月后对参与者做了追访,时间点落在后续危机事件发生之际,结果显示他们对不同的阴谋论的相信也降低了。 这是一个与主效应性质不同的论断。 对目标阴谋论的相信下降,符合一种狭义读法——模型针对一个特定事件提供了特定的反证,而它奏效了。 但对若干个月后不相关的新阴谋论的相信也下降,就无法这样解释。 迁移过去的东西不是事件特定的内容:而是某种倾向、某种被抬高的门槛、某种「在采纳一个说法之前先盘问它」的习惯。
实验时间线(两次重复均如此)
危机事件发生
|
| 数日
v
招募持阴谋论看法的成年人
|
+----+--------------------+--------------------+
| | |
v v v
实验组 对照组 A 对照组 B
多轮 LLM 对话, 与 LLM 聊 静态事实
被提示去降低信念 无关话题 清单
| | |
| 隔离出: | |
| 论证本身 | |
| 隔离出: 隔离出:
| AI 新鲜感 + 没有对话的
| 投入感 信息递送
| | |
+----+--------------------+--------------------+
|
v
即时测量:对目标阴谋论的相信
--> 实验组显著更低(两次实验均如此)
|
| 1-2 个月,新的危机事件发生
v
追访测量:对**不同**阴谋论的相信
--> 依然降低
下面是结构性比喻。 把阴谋论信念想成对一个含混事件的免疫反应。 当某件震撼且无解释的事发生时,心智需要一个说法,而且要快;在缺乏定论的情况下,它会制造一个能拼上现有碎片、也贴合此人先验的东西。 这个说法一旦被采纳,就抗拒修正——那就是领域长期相信的身份保护叙事。
在这个比喻里,此前的 LLM 辟谣工作是用一种成分明确的药去治一场慢性感染。 阴谋论是旧的,反驳论证发育得很充分,模型的活儿是以超常的耐心和个性化去施用一套已知疗法。
本文做的是另一件事:它在急性期介入,在信念尚未固化之前,而且不依赖任何既有疗法,因为疗法还不存在。 而令人意外的发现不只是「早介入有效」——而是这个效应看起来不像药,更像疫苗。 数月之后,面对一个新的病原体——另一场危机、另一个阴谋论——被干预过的参与者反应更弱。 不管那场对话做了什么,它都不只是中和了一个信念;它留下了某种会去应对下一个信念的东西。 这就是为什么那个下游结果是全文最重要的一句话,同时也是最脆弱的一句。
关键概念
-
正在成形 vs 已成型的阴谋论: 全文赖以立足的区分。 一个已成型的阴谋论拥有三样正在成形者所缺的东西:一份有据可查的事实记录、模型训练数据里的一整套现成反驳、以及一个在持有者心中已经硬化的信念。 一次性拿掉这三样,是一个难得多的测试,而且它恰好对应真实的部署场景——没有人需要以「国家级震荡后 72 小时内」那种紧迫程度,去针对登月否认论做干预。 这个难度有一个具体的说明:对 2025 年 9 月那起事件,模型可能完全没有关于该事件的训练数据,因此它无法提供关于发生了什么的事实。 那么它所做的事,必然更偏程序性而非事实性。
-
双对照设计: 为什么单一对照会毫无信息量。 把 LLM 对话与「什么都不做」相比,会把论证和体验混杂在一起。 一个人花二十分钟,与一个新奇而专注的对话者,就一个自己在意的话题投入交谈,其观点可能因为与所论证内容毫不相干的理由而缓和——社会赞许性、反思,或者单纯对自己立场的疲倦。 无关话题对照吸收掉了这一层。 另外,与「什么都不做」相比,也留下了「模型的内容才是全部」这种可能,那样一份 PDF 就够了。 事实清单对照吸收掉了这一层。 同时胜过两者,才使得「互动式、定制化的反驳论证是有效成分」这个论断得到许可。
-
下游迁移: 1 到 2 个月后、在不同危机事件之后,对不同阴谋论的相信也降低了。 它之所以是理论负载最重的结果:它与「干预纯粹通过内容递送起作用」不相容。 如果机制是「模型告诉了我关于事件 X 的真实情况」,那没有任何路径能让它在数月后降低关于不相关事件 Y 的信念。 因此这个结果指向某种倾向性的东西——一个人有多容易采纳阴谋论框架,这件事本身发生了移动。 它同时也是最暴露于替代解释的结果,因为「重复参与一项关于阴谋论信念的研究」本身就是一次干预。
框架转变
之前(此前的 LLM 辟谣工作): 之后(本文):
已成型的阴谋论 此刻正在成形的阴谋论
(几十年) (数日)
| |
v v
训练数据里有丰富的 无辟谣记录;事件甚至
辟谣积累 可能晚于训练数据
| |
v v
模型检索 + 定制 模型必须程序性地工作,
一个已知反驳 而非依赖存量反驳
| |
v v
对**那个**理论的相信下降 相信下降 —— 并且在数月后
| 迁移到**其他**阴谋论
v |
(迁移未被测试) v
看起来不像解药,
更像疫苗
从「证明 LLM 能对旧阴谋论递送成熟的反驳论证」,到「证明它能在信念真正被塑造的那几个小时里改变信念」,核心转变是把辟谣从一种检索行为,变成一种实时的、且看起来是持久的认知干预。
专家评审
选题眼光: 极好,而实地时机正是理由。 「在登月问题上有效」和「在此刻所有人都在争论的事情上有效」之间的那道缝,恰恰是决定此前发现有无实践后果的那道缝;而多数研究者会把它留着不补,因为补上它要求你有能力在一个不可预测的事件发生后数日内启动一项研究。 把它跑两次,在相隔十四个月的两起独立真实危机上,就不只是重复验证——它证明这个团队建起了做这件事的常备能力,而这本身就是对「这个研究方向该怎么做」的一项贡献。 这篇论文所在的位置,正是 LLM 说服文献从一件有趣的奇闻变成一个政策问题的那个点。
方法成熟度: 双对照是正确的设计,且执行得格外有纪律。 把「与 LLM 聊无关话题」选作其中一组,说明作者预判到了针对自己结果的最强通缩式读法,这是一个严肃实验团队的标志。 这套设计做不到的事——而这是内在局限,不是执行上的批评——是告诉我们对话的哪一个特征在起作用。 实验处理是一个捆绑包:个性化、耐心、互动性、具体证据、不带评判的语气。 胜过事实清单确立了「对话胜过文本」;它没有确立对话的哪一项性质在承重,而这恰恰是任何想部署它、或想防御它的人最需要知道的。
实验诚意: 样本量体面,而从实验一到实验二接近三倍的扩大,说明第二次是基于第一次做了恰当的功效计算。 自陈式信念测量仍是这个领域的标准,也是这个领域的软肋:当一个参与者刚刚花二十分钟,和一个明显在试图改变他想法的对话者待在一起时,需求特征是一个现实顾虑,而自然的反应就是报告出变化。 无关话题对照并不能完全吸收这一点,因为只有实验组会让实验者的目的变得可读。 我最想施压的是下游结果。 它是全文最惊人的论断,而它建立在一个必然有流失的追访样本上,且参与者已被原研究启动过「思考阴谋论信念」这件事。 这种设定下的流失不太可能是随机的——觉得那场对话有说服力的人,可能更愿意回来。 这些都不能说明结果是错的,但它是在被当作定论之前,最需要一次预注册的、对抗性设计的重复验证的论断。
写作功力: 摘要清晰,而且考虑到这个发现有多好被引用,其克制令人敬佩;「为新兴阴谋论的心理学提供了启示」是一个比结果所允许的更谦逊的说法,这是对的直觉。 偷懒之处在于,下游迁移结果只用一句话报告,却承载了全文理论分量中不成比例的一份。 如果有一节认真对待该结果的替代解释——流失、启动效应、均值回归——对论文的加固会强过任何一次追加实验。
判决: 强接收 —— 一个重要问题的困难版本,在真实事件上被回答了两次,对照设计预判了显而易见的反驳;持久迁移这一论断所受的审视少于它应得的,但主结果扎实且有后果。
要点总结
实践者可以从这篇论文「偷」走三样东西:
-
两个对照:一个管体验,一个管内容。 这个设计模式可以直接迁移到评估任何 LLM 干预——辅导、心理健康相邻工具、客服、行为改变。 跑一组让模型做无关的事(吸收新鲜感、注意力、投入感效应),再跑一组把同样的信息静态递送(吸收内容)。 只有同时胜过两者的部分,才能归因于互动式的定制化生成。 多数产品侧对 LLM 功能的评估这两组都没跑,因此无法把自己的模型和一份写得不错的 FAQ 区分开。
-
到脚手架缺席的地方去测能力。 可推广的方法论动作是:找出那个让你此前结果变容易的条件——充足的训练数据、已定论的真值、有据可查的答案——然后构造它缺席的情形。 这就是「在基准上好看的能力」和「能在部署中活下来的能力」之间的差别,而它对检索系统、智能体、评估套件的适用性,不亚于对说服。 问一句你的系统是在检索还是在推理,然后把它检索的那样东西撤掉。
-
通过找迁移来区分内容效应和倾向效应。 如果一次干预只改变了它所论证的那件事的相关信念,那是内容效应,每来一个新实例就得重新递送一次。 如果它在之后改变了不相关事物上的信念,那就是有倾向性的东西发生了移动,其经济学完全不同——一次互动带来持久外溢,而不是无休止地按例递送。 在追访时测量非靶标结果很便宜,而且这是区分二者的唯一办法。