Paper: 2610.06851 Authors: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros Categories: cs.AI, cs.CL, cs.LG
The Gap
The prevailing narrative in frontier AI development is that reasoning is synthesized during post-training. Under this view, base language models merely predict web text distributions; it is reinforcement learning (RL) with verifiable rewards or self-play (e.g., DeepSeek-R1, OpenAI o1 style training) that teaches the network how to search, backtrack, and construct chain-of-thought derivations.
However, this assumption overlooks the structural composition of the pre-training corpus itself. Rigorous mathematical proofs, code reviews, and competitive programming walkthroughs on the web do not appear at random; they are preceded by characteristic conversational openers, formatting cues, and transitional markers.
Could it be that the base model has already learned these reasoning pathways during pre-training, and is merely waiting for the right trigger token to steer its generation trajectory into those specialized document subspaces?
PROBLEM: DOES RL TEACH REASONING, OR MERELY RETRIEVE IT?
Dominant View:
Base Model (Unstructured Prediction)
|
v Expensive RL / Verifiable Rewards (RLVR)
Synthesizes Search & Reasoning Abilities from Scratch
|
v MATH-500: 42% -> 80%
--------------------------------------------------------------
THIS PAPER'S HYPOTHESIS & EXPERIMENTAL INTERVENTION:
Base Model + Starting Token Cue (e.g., ".\n\nOkay" or "Alright,")
|
v Forced Prefix
MATH-500 Accuracy Jumps Immediately:
- Olmo-3-7B: 42% -> 78% (Recovers 90% of RL gain without training!)
- Qwen3-14B: 72% -> 87%
|
v
Causal Pretraining Interventions:
- Inject arbitrary token "chicken" as prefix to math derivations
- "chicken" becomes an equally potent reasoning trigger!
- "Think duck duck goose" matches "Think step by step"
|
v
CONCLUSION: RL primarily boosts probability of pre-existing prefix cues
The Increment
One sentence: By demonstrating that fixing minimal opening token cues allows un-fine-tuned base models to match the mathematical and coding accuracy of RL-trained checkpoints, this paper proves that complex reasoning policies reside latent in pre-trained weights and are routed through document-distribution associations formed during pre-training.
Core Mechanism
The authors conduct a series of controlled causal interventions on base model generations across standard mathematical reasoning benchmarks (MATH-500, GSM8K) and competitive coding:
- Prefix Cue Forcing: Rather than allowing the base model to sample its opening tokens, the authors force-feed candidate starting tokens at the beginning of the completion.
- Appending
".\n\nOkay"to the prompt before generation instantly elevates Olmo-3-7B’s MATH-500 pass@1 from 42% to 78%. - Appending
"Alright,"elevates Qwen3-14B from 72% to 87%, nearly matching its fully RL-tuned sibling.
- Appending
- Mechanism of RL: When analyzing the checkpoint before and after reinforcement learning, the authors observe that RL drastically sharpens the model’s likelihood of emitting these exact reasoning cues as its initial tokens. Once those initial tokens are emitted, the autoregressive rollout follows the reasoning manifold that pre-training had already established.
- Causal Pretraining Manipulation:
- To rule out innate linguistic semantics of English words like “Okay,” the researchers pre-trained synthetic models from scratch.
- By prepending arbitrary nonsense words—such as
"chicken"or"quack"—to mathematical and algorithmic walkthroughs in the pre-training data, those arbitrary tokens become fully functional reasoning triggers. - Similarly, instructing the model with
"Think duck duck goose"elicited identical chain-of-thought derivations as the canonical"Think step by step"if that phrase co-occurred with technical problem solutions.
- Internal Representation Routing: Probing intermediate hidden states reveals that forcing a reasoning cue immediately shifts the network’s internal representations toward the subspace occupied by technical lecture notes, textbooks, and forum solutions, away from casual social media chatter.
ACTIVATION ROUTING THROUGH PREFIX TOKENS
Query: "Find all real roots of P(x) = x^3 - 3x + 1..."
|
v
+-------------------------------------------------------------+
| Unconditioned Base Model: Samples generic conversational tokens |
| -> Routes to casual web chat manifold -> Fails / Hallucinates |
+-------------------------------------------------------------+
|
v Inject Prefix Cue: ".\n\nOkay"
+-------------------------------------------------------------+
| Conditioned Base Model: Hidden states snap into Math/Code |
| document manifold -> Emits rigorous derivation -> Solves! |
+-------------------------------------------------------------+
The load-bearing structural metaphor is a repertory theater actor backstage.
- The actor has spent years memorizing the complete works of Shakespeare, quantum physics lectures, and barroom stand-up comedy (pre-training).
- When pushed onto a pitch-black stage without props or direction, the actor starts bantering like a tavern patron because tavern talk is the most common default mode of human speech.
- You do not need to send the actor to graduate school for four years (reinforcement learning) to teach them quantum mechanics. You simply hand them a tweed jacket and chalk (the opening token cue
".\n\nOkay"). The moment they put on the jacket, the entire professor persona activates instantly, delivering a flawless physics derivation from memory. RL does not teach the physics; it simply teaches the actor to grab the tweed jacket by default when hearing the word “problem.”
Key Concepts
- Reasoning Cue: A small set of opening tokens (such as punctuation followed by conversational transitional phrases) that conditions the Transformer’s attention onto pre-training document distributions containing step-by-step technical problem solving.
- Causal Data Association: The associative link formed in next-token prediction between discourse markers that introduce problem solutions in training text and the structured logic that follows.
- Manifold Routing vs. Skill Acquisition: The distinction between a model learning how to logically deduce new theorems versus a model navigating to a pre-existing latent capability space.
Framework Shift
Before (The Post-Training Emergence Paradigm):
Pre-training: Ingests raw knowledge without reasoning ability.
RL Post-Training: Teaches the model search policies, self-correction, and reasoning logic.
After (The Cue-Driven Document Routing Paradigm):
Pre-training: Ingests both knowledge AND sophisticated reasoning patterns from high-quality text.
RL Post-Training: Primarily aligns the opening token probabilities to point into the reasoning manifold.
Proof: Forcing the cue on the base model recovers ~90% of RL's benchmark gains for free!
From “RL creates reasoning out of whole cloth,” the core shift is that base models already possess profound reasoning circuits formed during pre-training, which can be summoned by simply steering the model’s opening discourse markers.
Expert Assessment
Problem choice: Groundbreaking and provocative. The AI community has spent billions of compute hours fine-tuning and running RL on models under the assumption that base models cannot reason on their own. Exposing the latent reasoning capability of unaligned checkpoints reshapes our understanding of LLM intelligence.
Method maturity: The causal data intervention experiments are airtight. By substituting arbitrary non-words (“chicken”) into pretraining data and observing the emergence of the identical cue effect, the authors decisively eliminate confounding semantic arguments.
Experimental integrity: Tested across distinct model families (Olmo, Qwen) and rigorous benchmarks (MATH-500, coding suites). Safety evaluations further confirm that the same cue-based routing governs compliance versus refusal behaviors.
Writing quality: Exemplary clarity, lucid storytelling, and sharp empirical precision.
Verdict: strong accept — One of the most clarifying mechanistic interpretability and pre-training papers of the year, dismantling key myths surrounding the necessity of heavy post-training for basic reasoning tasks.
Takeaways
- When evaluating or deploying base models, do not let them sample opening tokens unconstrained; experiment with prefix-cue forcing (
".\n\nOkay","Alright,","Let's think carefully:") to unlock immediate reasoning jumps without fine-tuning. - During pre-training data curation, ensure that high-quality technical and mathematical reasoning content is paired with clean, consistent introductory markers.
- Re-evaluate post-training RL pipelines: measure whether your RL algorithm is truly discovering novel problem-solving topologies or merely optimizing the likelihood of starting token triggers.
论文: 2610.06851 作者: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros 分类: cs.AI, cs.CL, cs.LG
缺口
在当前大模型研发的流行叙事中,存在一个近乎共识的认知:复杂的推理能力是靠后训练阶段(Post-training)的强化学习(RL)「教出来」的。 传统观念认为,海量预训练只给基座模型灌入了静态知识与文本续写直觉;只有通过类似 DeepSeek-R1、OpenAI o1 那样带可验证奖励(RLVR)的大规模搜索与自我博弈强化学习,模型才学会了回溯、拆解步骤以及缜密的思维链(Chain-of-Thought)推导。
然而,这一认知忽略了预训练语料本身的宏观组织结构。 互联网上高质量的数学证明、竞赛代码题解以及科学论坛讨论,从来不会平白无故地突然冒出来——它们几乎总是紧随在某些特定的过渡词、段落标点或口语化引导词之后。
基座大模型真的缺乏推理能力吗? 抑或,它的大脑里早就刻满了完整的推理计算图,只是在默认情况下由于开头的采样概率被冲淡,未能精准定位到这些高智商文档分布中?
核心疑问:强化学习究竟是创造了推理,还是仅仅激活了先验记忆?
传统主流观点:
基座模型 (仅有零散知识与胡乱续写)
|
v 昂贵且漫长的强化学习 (RL / RLVR 搜索)
从零学会严密的逻辑推演与思维链展开
|
v MATH-500 得分:42% -> 80%
--------------------------------------------------------------
本文的核心猜想与因果实验干预:
未经任何后训练的纯净基座模型 + 固定起始前缀 Token (如 ".\n\nOkay" 或 "Alright,")
|
v 在生成开头强行注入极简提示前缀
MATH-500 准确率产生惊人跳跃:
- Olmo-3-7B:从 42% 暴涨至 78% (零训练直接收复 RL 超过 90% 的涨幅!)
- Qwen3-14B:从 72% 飙升至 87%
|
v
预训练因果干预实验验证:
- 故意将无意义词汇 "chicken" 绑定到预训练数学题解开头
- "chicken" 瞬间成为威力等同的顶级推理激活开关!
- "Think duck duck goose" 诱发的效果与 "Think step by step" 毫无二致
|
v
结论:RL 的核心作用仅在于拉高这些本就存在的前缀诱导词的输出概率
增量
一句话: 通过证实仅需在前缀强制填入极简起始词(如「.\n\nOkay」)即可让未经微调的基座大模型在数学与代码基准上打出媲美强化学习检查点的惊人表现,本文揭示了复杂推理能力本就潜伏在预训练权重之中,而所谓推理激活本质上是向特定训练语料文档流的表征路由。
核心机制
研究团队在多个权威数学与编程评测集(MATH-500、GSM8K 及算法竞赛题库)上,对基座模型展开了细致的因果干预分析:
- 前缀提示词强制注入(Prefix Cue Forcing):
- 阻止模型在开头自由采样口语废话,而是在模型回答的第 1 个 Token 强制写入特定的字符组合。
- 仅仅在提示后拼接
".\n\nOkay",Olmo-3-7B 在 MATH-500 上的 pass@1 准确率便从 42% 陡增至 78%。 - 拼接
"Alright,",Qwen3-14B 的正确率从 72% 拔高至 87%,直接逼近其重度 RL 训练后的衍生版本。
- 解密 RL 的真实物理效应:
- 对比强化学习前后的权重变化,作者发现 RL 的最大贡献是在遇到推理任务时,将模型第一步吐出这些特定推理引导词的概率拉升了数倍。
- 一旦第一步落子在这些引导词上,后续的自回归生成便自然而然地滑入预训练时期习得的高阶推理流形之中。
- 预训练层面的因果实证:
- 为排除 “Okay” 本身的自然语言语义干扰,研究者从头训练了微型验证模型。
- 人为在预训练语料中的算法解答前插入毫无关联的单词(如
"chicken"或"quack")。 - 结果证实:只要在语料中建立共现关联,输入
"chicken"就能完美诱发深度推演;而让模型"Think duck duck goose"(想丢手绢)与经典的"Think step by step"能够取得完全一致的思维链爆发效果。
- 内部表征的流形切换:
- 隐藏状态探针显示,注入引导词的瞬间,模型的注意力与表征向量立刻脱离社交闲聊空间的分布,瞬移到由教科书、学术论文与代码仓库主导的隐空间流形。
前缀 Token 对表征流形的路由机制
数学问题:"求解多项式 P(x) = x^3 - 3x + 1 的全部实数根..."
|
v
+-------------------------------------------------------------+
| 无干预的基座模型:开头采样网络口水词,流向普通的闲聊文档流形 |
| -> 导致逻辑松散、推导半途而废或产生严重计算幻觉 |
+-------------------------------------------------------------+
|
v 强制注入前缀词: ".\n\nOkay"
+-------------------------------------------------------------+
| 受控基座模型:隐藏状态瞬间被拽入高密度专业数学与代码题解流形 |
| -> 自动输出严密细致的逐步代数推导 -> 命中正确答案! |
+-------------------------------------------------------------+
这里的核喻是剧团后台储备丰富的资深老戏骨。
- 这位老戏骨博览群书,脑子里背熟了莎士比亚全集、理论物理推导、以及街头插科打诨的全部台词(预训练海量语料)。
- 如果灯光突然亮起,什么道具都不给,他很可能随口唠起家常,因为世俗闲聊是他见过最密集的日常对话。
- 你根本不需要送他去读四年理论物理博士(漫长的后训练强化学习)来教他怎么算量子力学;你只需要在开场铃响时,塞给他一副黑框眼镜和一件粗花呢西装(前缀引导词
".\n\nOkay")。 老戏骨一披上这件衣服,大学教授的人格瞬间上身,倒背如流地写下黑板推导。 强化学习并没有教会他物理学;强化学习只是在台后踢了他一脚,逼他看到题目时下意识去抓那件粗花呢西服。
关键概念
- 推理诱导词(Reasoning Cue):一类特定的起始 Token 组合(通常包含特定标点、换行与过渡转折连词),能够强制引导自回归注意力锚定在预训练时期记录的高质量推演文档分布上。
- 语料因果关联(Causal Data Association):在下一个 Token 预测训练中,特定语篇过渡标记与后续严密逻辑链之间建立起的强依赖条件概率。
- 流形路由与能力习得(Manifold Routing vs. Skill Acquisition):区别在于模型是真正通过算法从无到有学会了逻辑搜索,还是仅仅学会在已有庞大技能库中拨动正确的索引开关。
框架转变
之前 (后训练能力涌现假说):
预训练:只是粗放灌输世界知识,模型本质没有逻辑与推理能力。
RL 后训练:通过奖惩信号从无到有合成了思维链、自省和探索搜索策略。
之后 (语料提示词路由假说):
预训练:知识与高级推理策略早已完整封存在海量优质解题语料中。
RL 后训练:主要作用是调整初始 Token 概率,避免模型落入低智闲聊分布。
铁证:仅靠在基座模型输入端强制指定一个引导词,即可免费白嫖 90% 的 RL 涨幅!
从「迷信强化学习是推理能力的唯一造物主」,核心转变在于:基座大模型本就蕴含着极为强大的推理网络,我们所需要的,仅仅是用合适的前缀标记在起点处拨动正确的路由开关。
专家评审
选题眼光: 极具颠覆性与反思价值。 全行业在后训练与强化学习上烧掉了天文数字般的算力,默认基座模型是「不会思考的蛮荒状态」。 本文用极简的实验设计戳破了这一神话,直指预训练大模型真正的能力边界。
方法成熟度: 因果干预实验设计堪称教科书级别。 利用无意义代称(“chicken”)改写预训练语料并复现出完全一致的推理激发现象,干脆利落地排除了自然语言词义解释的混淆变量。
实验诚意: 横跨 Olmo 与 Qwen 两大顶级开源模型家族,并在 MATH-500、代码评测以及安全拒答分类等多个正交任务上交叉验证,数据信服力拉满。
写作功力: 逻辑清晰、步步紧扣,直击模型可解释性与后训练机制的核心。
判决: 强接收 (strong accept) — 本年度最具启发性的大模型机理研究之一,对重新理解预训练质量与后训练真实价值具有里程碑式的指导意义。
要点总结
- 在评测或直接调用未经微调的基座大模型时,切勿放任其自由生成开头;尝试强制注入
".\n\nOkay"或"Alright,"等前缀引导词,常常能零成本激活强劲的思维链推理。 - 预训练数据工程团队应格外重视优质解题、代码排错语料的篇章开篇格式,确保其具备清晰且规范的引导标记,便于模型建立稳固的条件概率通道。
- 重新审视现有的 RLVR 算法指标:严格排查模型在强化学习期间究竟是学到了更优的思维探索拓扑,还是仅仅记住了几个高收益的开篇引导 Token。