Paper: 2608.23543 Authors: Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers, Padhraic Smyth Categories: cs.AI
The Gap
The promise of AI assistance is straightforward — better performance now. The worry is equally straightforward: that assistance substitutes for the very effort through which skill is built, so the performance gain is rented rather than owned. Both claims are widely asserted, and the second is stated far more often than it is measured.
Measuring it requires a design that most studies do not have. You need to observe the same people before assistance exists, while it is available, and after it is taken away. That third phase is the one that matters and the one usually missing, because it is the only phase in which you can see whether the skill stayed. And you need to separate three quantities that a naive before-and-after comparison conflates: how good someone was to begin with, how good they are afterward, and how much they individually changed.
Without that separation, the obvious analysis is also the wrong one. If you predict post-AI performance from AI-assisted performance, you will be reading a number that was produced partly by the AI.
THE THREE PHASES (and why most studies skip the third)
[1] BEFORE AI baseline ability
|
[2] DURING AI performance goes UP <- easy to measure
| <- the "promise"
[3] AFTER AI assistance REMOVED <- the phase that matters
| <- usually missing
v
[THE CONFOUND] predict post-AI ability from
AI-assisted performance?
-> that score was partly produced by the AI
-> overestimates what the person can do alone
|
v
[GAP] the skill-retention question needs all three
phases AND a decomposition of ability
The Increment
One sentence: Before this study, the claim that AI assistance erodes skill was plausible but largely unmeasured; after it, a three-phase controlled experiment with a Bayesian latent ability model shows that AI users perform worse after assistance is removed, that extrapolating from assisted scores overestimates them, and that independent reasoning effort predicts the size of the gain.
Core Mechanism
The design is a logic-puzzle experiment with on-demand AI assistance, and it has two features that make it work.
Three phases with the same participants. Tasks are completed before, during and after AI is available. The “after” phase is the whole point: it is the only measurement taken when the crutch is gone, and it is what separates a rented gain from an owned one.
Experimental variation in the cost of asking. The AI request cost is varied across participants, which turns “did they use it” from a self-selected choice into something closer to an assigned condition. The finding is unsurprising and important: lower-cost assistance induces more frequent AI use. Friction is the lever on usage, which means the amount of assistance is a design parameter rather than a fixed property of the tool.
Three results follow.
Participants who requested AI assistance during the access phase performed worse on the task after assistance was removed. This is the direct measurement of the concern, and it is a performance deficit, not merely an absence of gain.
Their subsequent unassisted performance is overestimated when predicted from earlier AI-assisted performance. This is the finding with the widest practical reach. The natural way to estimate a person’s ability is to look at what they recently achieved — but recent achievement was co-produced by the AI. Any downstream decision that extrapolates from assisted scores inherits a systematic optimism.
Greater independent problem-solving effort during the access phase is associated with larger gains in latent ability. This is the mechanism, and it is where the Bayesian latent ability model does its work. By separating initial ability, post-AI ability, and participant-specific skill change, the analysis can ask what predicts the change. The answer is independent effort — consistent with the interpretation that skill development is weaker precisely when assistance substitutes for independent reasoning.
ANALYSIS STRUCTURE (Bayesian latent ability model)
observed: puzzle outcomes across three phases
|
v
decompose into:
[ initial ability ] who they were at the start
[ post-AI ability ] who they are at the end
[ skill change ] the participant-specific delta
|
v
PREDICTOR OF THE DELTA:
independent problem-solving effort during the AI phase
-> more independent effort => larger latent gain
-> assistance that REPLACES reasoning => weaker gain
|
v
ALSO: predicting post-AI ability from AI-assisted
scores systematically OVERESTIMATES it
Think of it as learning to navigate versus following a navigator. Someone driving a new city with turn-by-turn directions arrives fine, on time, and learns almost nothing about the street layout — the navigation replaced the wayfinding rather than supporting it. Take the phone away and they are worse off than a driver who spent the same trips getting lost and correcting. The two telling details from the study map onto this cleanly: cheap directions mean you ask for directions constantly, and the passenger’s confidence after the trip is a bad predictor of whether they could drive it alone tomorrow.
Key Concepts
- Rented versus owned performance: a gain that persists only while the aid is present versus one that survives its removal. The three-phase design exists solely to distinguish these, and the “after” phase is the only place the distinction is observable.
- Overestimation from assisted scores: predicting unassisted ability from assisted performance inherits the AI’s contribution as if it were the person’s. It is a measurement error with policy consequences, because assessment decisions are typically made from recent achievement.
- Effort as the mediator: independent problem-solving effort predicts the size of the latent ability gain. This matters because it identifies a variable a person can actually control, and it gives a concrete reading of what “substituting for reasoning” means behaviourally.
- Friction as the usage lever: varying request cost changes how often assistance is used. If the goal is to preserve skill development, the cost of asking is a design parameter, not an incidental UI detail.
Framework Shift
Before (assistance evaluated by the gain it produces):
measure performance with and without the tool
-> assistance helps (true, and short-sighted)
-> ability estimated from recent assisted output
After (assistance evaluated by what survives removal):
measure before / during / after
-> during: gain (real)
-> after: DEFICIT vs those who did not use it
-> ability from assisted scores: OVERESTIMATED
-> predictor of retained gain: independent effort
assistance that REPLACES reasoning -> skill loss
assistance that SUPPORTS reasoning -> retained gain
From judging an aid by the performance it produces while present, to judging it by what the user can still do when it is gone, the core shift is that the interesting measurement is taken after the tool has been removed.
Expert Assessment
Problem choice: Excellent, and this is the right moment for it. The skill-erosion worry is now a standard objection in discussions of AI assistance, and it was being argued from intuition. The three-phase design is the minimum needed to test it, and it is notable how few studies implement the third phase.
Method maturity: The Bayesian latent ability model is doing real work rather than decorating the result. Without it, the natural analysis — predict unassisted performance from assisted performance — produces a biased answer, which the paper demonstrates rather than merely asserts. Varying request cost to make usage less self-selected is a sensible design choice, though it governs usage rather than randomising the decision to use AI, so the causal reading of the deficit still rests on the assumption that cost variation is as-good-as-random with respect to unmeasured ability.
Experimental integrity: The honesty of stating both a performance deficit and an overestimation effect — the two results most likely to be contested — is to the authors’ credit, and the direction of the deficit is consistent across the analysis. The scope caveat is the logic-puzzle setting: a task with short horizons, immediate feedback and no transfer requirement. Whether the same erosion appears in open-ended, long-horizon work where assistance may scaffold rather than replace is genuinely unknown, and the paper does not overclaim it.
Writing quality: The distinction between “skill development is weaker when assistance substitutes for independent reasoning” and the stronger claim “AI makes people worse” is maintained carefully, which is exactly the discipline this topic needs. The paper would be more useful with one concrete illustration of the overestimation — a plot of predicted-versus-actual unassisted performance against the diagonal — since that is the finding most likely to be actioned by anyone doing assessment.
Verdict: strong accept — it converts a widely repeated intuition into a measured effect with an identified mechanism, and it names a measurement error that has consequences well beyond this experiment.
Takeaways
- Evaluate any assistance tool on the “after” phase. A gain measured while the aid is present is not evidence of retained capability, and that phase is usually the one nobody runs.
- Do not extrapolate unassisted ability from assisted performance. The score was co-produced by the tool, and the resulting estimate is systematically optimistic.
- Treat the cost of asking as a design parameter. Friction reliably changes how often assistance is used, which makes it a lever on how much independent reasoning is displaced.
- Track independent effort, not just outcomes. Effort during the assisted phase predicted the latent ability gain here, which means it is both a diagnostic and the thing to protect if you want the skill to survive.
论文: 2608.23543 作者: Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers, Padhraic Smyth 分类: cs.AI
缺口
AI 辅助的承诺很直白——现在表现更好。 而担忧同样直白:辅助替代了技能赖以建立的那份努力,于是这份能力提升是”租来的”,而不是”自己的”。两种说法都被广泛主张,而后者被反复陈述的次数,远多于它被真正测量的次数。
要测量它,需要一个大多数研究并不具备的设计。 你需要观察同一批人在辅助尚未存在时、辅助可用时、以及辅助被撤除后的表现。第三阶段才是关键,也恰恰是通常缺失的那一阶段——因为只有在那个阶段,你才能看到技能是否留下了。 而且你还需要把三样被”朴素的前后对比”混在一起的东西分开:一个人原本有多强、他之后有多强、以及他本人改变了多少。
缺少这份拆分,最显而易见的分析方式也恰好是错误的那种。 如果你用”有 AI 辅助时的表现”去预测”撤除 AI 后的表现”,你读到的那个数字,有一部分本来就是 AI 产出的。
三个阶段(以及为什么多数研究跳过第三个)
[1] AI 之前 基线能力
|
[2] AI 期间 表现上升 <- 容易测
| <- 这就是"承诺"
[3] AI 之后 辅助被撤除 <- 真正要紧的阶段
| <- 通常缺失
v
[混淆] 用"有辅助时的表现"去预测"无辅助时的能力"?
-> 那个分数本来就部分由 AI 产出
-> 高估了这个人独自能做到的水平
|
v
[缺口] 技能保留这个问题,既需要三个阶段,
也需要对能力做拆分
增量
一句话: 在这项研究之前,“AI 辅助会侵蚀技能”这个说法看似合理却基本未被测量;在这项研究之后,一个三阶段对照实验加上贝叶斯潜在能力模型表明:使用 AI 的人在辅助撤除后表现更差,用其”有辅助时的成绩”做外推会高估其能力,而独立推理的投入量可以预测能力增长的大小。
核心机制
设计本身是一个按需提供 AI 辅助的逻辑谜题实验,而它之所以能成立,靠两个特点。
同一批人、三个阶段。 任务分别在 AI 可用之前、可用期间、以及可用之后完成。“之后”这个阶段才是全部重点:它是唯一一次在拐杖撤走之后进行的测量,也正是它把”租来的提升”与”自己的提升”区分开。
通过实验手段改变”求助的代价”。 AI 请求的代价在不同受试者之间被设定为不同水平,这就把”他们是否用了”从一个自选行为,变成了更接近被分配的条件。结论不出人意料但很重要:求助代价越低,使用 AI 的频率就越高。 摩擦是使用量的杠杆——这意味着”辅助的量”是一个设计参数,而不是工具的一个固定属性。
由此得出三条结果。
在辅助可用阶段请求过 AI 的受试者,在辅助被撤除后表现更差。 这是对那份担忧的直接测量,而且它是一个能力赤字,而不只是”没有获得增益”。
用其”有辅助时的表现”去预测其”撤除后的表现”,会系统性地高估后者。 这是实践影响面最广的一条。估计一个人能力的自然方式,就是看他最近做到了什么——但”最近做到的”是 AI 共同产出的。任何从”有辅助成绩”做外推的下游决策,都会继承这份系统性的乐观偏差。
在可用阶段投入更多独立解题努力,与潜在能力的更大增长相关联。 这就是机制所在,也是贝叶斯潜在能力模型发挥作用的地方。通过把初始能力、AI 后的能力、以及个体特有的技能变化拆开,分析得以追问:什么能预测这个变化量。答案是独立投入——这与”当辅助替代了独立推理时,技能发展就更弱”这一解读相一致。
分析结构(贝叶斯潜在能力模型)
观测:三个阶段中的谜题作答结果
|
v
拆分出:
[ 初始能力 ] 他们在起点上是谁
[ AI 后能力 ] 他们在终点上是谁
[ 技能变化 ] 个体特有的增量
|
v
对「增量」的预测因子:
在 AI 阶段投入的独立解题努力
-> 独立投入更多 => 潜在能力增长更大
-> 辅助「替代」推理 => 增长更弱
|
v
另外:用「有辅助时的成绩」预测「撤除后的能力」,
会系统性地高估它
可以用**“自己认路 vs 跟着导航开”来理解这件事: 一个人在新城市里靠逐向语音导航开车,能顺利、准时抵达,而对街道格局几乎一无所知——导航替代了认路,而不是支持**了认路。把手机拿走,他比一个把同样这些行程用来迷路、纠错的司机还要更差。 研究里那两个耐人寻味的细节,正好对上:指路太便宜,你就会不停地问路;而一趟行程结束后副驾驶的信心,并不能预测他明天独自能不能开。
关键概念
- 租来的表现 vs 自己的表现: 只在辅助在场时成立的提升,与在辅助撤除后仍然存活的提升。三阶段设计的存在,正是为了区分这两者;而”之后”阶段是唯一能观测到这种区别的地方。
- 由辅助成绩导致的高估: 用”有辅助时的表现”预测”无辅助时的能力”,会把 AI 的贡献当成这个人自己的贡献一并继承下来。这是一种带有决策后果的测量误差,因为评估决策通常就是依据”最近做到了什么”来做的。
- 以投入为中介变量: 独立解题努力可以预测潜在能力增长的大小。这一点重要,因为它指出了一个人确实能控制的变量,也让”替代推理”在行为层面有了具体的读法。
- 以摩擦为使用量杠杆: 改变求助代价,就能改变辅助被使用的频率。如果目标是保住技能发展,那么”求助的代价”是一个设计参数,而不是无关紧要的界面细节。
框架转变
之前(用「辅助带来的增益」评估辅助):
测量有工具/无工具时的表现
-> 辅助有帮助(是真的,但短视)
-> 用最近的有辅助输出估计能力
之后(用「撤除后剩下什么」评估辅助):
测量 之前 / 期间 / 之后
-> 期间:有增益 (真实)
-> 之后:相对未使用者出现 DEFICIT(赤字)
-> 由辅助成绩推能力:被高估
-> 保留增益的预测因子:独立投入
替代推理的辅助 -> 技能流失
支持推理的辅助 -> 增益被保留
从”用辅助在场时产生的表现来评判一个工具”,转变为”用工具被拿走之后使用者还能做什么来评判它”,核心转变在于:真正有意思的那次测量,是在工具撤除之后才做的。
专家评审
选题眼光: 极好,而且时机正对。 “技能侵蚀”这份担忧,如今已是讨论 AI 辅助时的标准反驳,而它此前主要靠直觉在争论。三阶段设计是检验它的最低配置,而值得注意的是:真正实现第三阶段的研究少之又少。
方法成熟度: 贝叶斯潜在能力模型是在真干活,而不只是给结论做装饰。 没有它,最常见的分析方式——用有辅助时的表现去预测无辅助时的表现——会给出有偏答案;而论文是证明了这一点,而不只是断言。改变求助代价、让使用行为不那么自选,是合理的设计选择;不过它调节的是使用量,而非随机化”是否使用 AI”这个决定,因此对那个赤字的因果解读,仍然依赖”代价的变动相对于未观测能力而言近似随机”这一假设。
实验诚意: 把”表现赤字”与”高估效应”这两个最可能引发争议的结果都直白写出来,值得肯定;赤字的方向在整个分析中是一致的。范围上的保留在于”逻辑谜题”这一设定:任务周期短、反馈即时、且不要求迁移。在开放式、长周期的工作中——辅助可能起到脚手架而非替代作用——是否出现同样的侵蚀,目前确实未知,而论文也没有过度主张。
写作功力: 论文谨慎地维持了”当辅助替代独立推理时技能发展更弱”与更强的断言”AI 让人变差”之间的区分,这正是这个话题所需要的纪律。若能补一张具体图示——把”预测的无辅助表现”与”实际的无辅助表现”对对角线的散点画出来——会更有用,因为那正是最可能被任何做评估的人付诸行动的那条结论。
判决: 强接收(Strong Accept) — 它把一个被反复重复的直觉,变成一个被测出的效应,并指出了机制;同时命名了一种影响远超本实验的测量误差。
要点总结
- 评估任何辅助工具时,请看”之后”那个阶段。在辅助在场时测到的增益,并不能证明能力被保留;而那个阶段恰恰是通常没人做的一环。
- 不要用”有辅助时的表现”去外推”无辅助时的能力”。那个分数是工具共同产出的,由此得到的估计会系统性地偏乐观。
- 把”求助的代价”当成设计参数。摩擦会稳定地改变辅助的使用频率,因此它是控制”多少独立推理被替代掉”的一根杠杆。
- 追踪独立投入,而不只是结果。在本研究中,辅助阶段的独立投入预测了潜在能力的增长——这意味着它既是一个诊断指标,也是”想让技能留下来就必须守住的东西”。