Paper: 2609.40360 Authors: Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng Categories: cs.AI, cs.CL, cs.LG

The Gap

Reinforcement learning with verifiable rewards (RLVR)—most notably Group Relative Policy Optimization (GRPO) used in DeepSeek-R1, QwQ, and OpenR1—has unlocked phenomenal leaps in mathematical and algorithmic reasoning. By generating a group of candidate reasoning rollouts and normalizing rewards across the group, models learn self-correction and deep exploration.

However, models trained under GRPO display an infuriating, fragile pathology: extreme sensitivity to semantically irrelevant prompt perturbations. Changing a character’s name from Alice to Bob, altering the order of introductory clauses, or reformatting whitespace routinely causes a previously solved AIME competition problem to fail completely.

The flaw lies in the mathematical heart of GRPO’s credit assignment: GRPO assigns the exact same scalar advantage A^\hat{A} to every single token in the generated response.

If a rollout arrives at the correct answer (Reward = 1), GRPO rewards all 2,000 tokens equally. This indiscriminately reinforces both genuine logical deductions and fragile, spurious tokens that happen to overfit to specific prompt quirks.

   THE UNIFORM CREDIT TRAP IN GRPO

   Prompt: "Farmer John has 15 cows and 7 sheep..."
               |
               v
   Generated Trajectory:
   [Token 1..50: "Let John's sheep be s..."]  <-- Spurious superficial phrasing!
   [Token 51..500: True mathematical deduction] <-- Genuine logical engine!
   [Token 501: "Answer is 22"]                  <-- Correct! (Reward = 1)
               |
               v
   GRPO Update:
     Advantage A = +1.0 assigned UNIFORMLY to all 501 tokens!
     -> Reinforces spurious surface tokens just as strongly as core math!
     -> Model becomes brittle: change "John" to "Mary" and reasoning collapses!
               |
               v
   SCAPO SOLUTION: Semifactual Prompt Interventions
     Perturb surface phrasing while keeping math facts identical ("Even if Mary...")
     Measure which tokens drift in probability -> Down-weight unstable tokens!

The Increment

One sentence: By applying semifactual prompt interventions to measure token-level causal stability and selectively damping policy gradient advantages on prompt-sensitive tokens during training, SCAPO eliminates spurious dependencies in GRPO, boosting AIME 2024–2026 accuracy by 5.63 percentage points on Qwen3-4B.

Core Mechanism

SCAPO (Semifactual Credit-Augmented Policy Optimization) introduces a causally grounded credit refinement layer on top of standard GRPO:

  1. Semifactual Prompt Invariance: A semifactual intervention alters the surface expression of a prompt x→x~x \to \tilde{x} (e.g. narrative context, variable symbols, syntactic phrasing) while guaranteeing that the underlying ground-truth mathematical premise and target answer remain strictly invariant: P(Answer∣x)≡P(Answer∣x~)P(\text{Answer} | x) \equiv P(\text{Answer} | \tilde{x})
  2. Token Probability Drift Measurement: For a generated successful response trajectory y=(y1,y2,…,yT)y = (y_1, y_2, \dots, y_T), SCAPO probes how the conditional token probabilities shift when evaluated under the perturbed prompt x~\tilde{x}: Δt=∣log⁡πθ(yt∣x,y<t)−log⁡πθ(yt∣x~,y<t)∣\Delta_t = | \log \pi_\theta(y_t | x, y_{<t}) - \log \pi_\theta(y_t | \tilde{x}, y_{<t}) | A high Δt\Delta_t reveals that token yty_t is heavily tethered to superficial prompt quirks rather than mathematical necessity.
  3. Credit Damping (Not Inflation): Crucially, SCAPO does not hand out bonus rewards for stability (which would incentivize models to output bland, generic tautologies). Instead, it damps the positive advantage on unstable tokens during early training: A^tSCAPO=A^GRPO⋅σ(Stabilityt)\hat{A}_t^{\text{SCAPO}} = \hat{A}^{\text{GRPO}} \cdot \sigma(\text{Stability}_t) Spurious tokens have their gradient updates throttled, forcing the model to allocate its capacity to robust, invariant reasoning steps.
   SCAPO TOKEN-LEVEL ADVANTAGE REFINEMENT

   Token Stream:   [ Let ]  [ John ]  [ have ]  [ x ]  ...  [ x = 15 + 7 ]
                     |        |         |        |             |
   Semifactual       |        |         |        |             |
   Drift (Delta):   Low     HIGH       Med      Low           ZERO
                     |        |         |        |             |
   Stability Head:  Keep    DAMP      Damp     Keep          FULL
   Advantage:       +1.0    +0.1      +0.4     +1.0          +1.0
                     ^        ^                                ^
                     |        |                                |
   Result: Spurious phrasing is pruned; mathematical logic receives full credit!

The structural metaphor is a marathon runner being coached across different weather conditions.

  • GRPO is a naive coach who only looks at the finish stopwatch: if the runner wins a race on a sunny day wearing yellow socks, the coach forces them to keep wearing the exact same socks, drink the exact same brand of sparkling water, and run with the exact same lucky hat (rewarding spurious tokens). When it rains in London, the runner is psychologically ruined.
  • SCAPO is a sports scientist who tests the runner in wind tunnels (semifactual interventions): they test the runner in rain, mist, and headwinds. They immediately see that cardiovascular stamina and stride cadence (invariant mathematical logic) determine the win in every climate, while the yellow socks (high-drift tokens) have zero causal bearing on speed. The coach tells the runner to stop obsessing over the lucky socks, focusing all praise and training strictly on lung capacity and stride cadence.

Key Concepts

  • Semifactual Intervention: In causal inference, an intervention that modifies antecedent conditions without changing the truth value of the factual consequence (“Even if X′X' had occurred, YY would still have occurred”).
  • Token-Level Credit Assignment: Differentiating the reinforcement signal across the tokens of a sequence rather than treating the entire generated response as a monolithic scalar unit.
  • Spurious Dependency in Reasoning: When an autoregressive model conditions its next reasoning step on stylistic, syntactic, or lexical idiosyncrasies of the prompt that possess zero logical necessity.

Framework Shift

Before (GRPO Uniform Token Credit Assignment):
  Reward = 1 -> Assign A = +1.0 to ALL 2,000 tokens equally
  -> Model overfits to superficial prompt formatting and variable names
  -> Fragile to out-of-distribution prompts, fails on rephrased AIME problems

After (SCAPO Causal Credit Refinement):
  Semifactual prompt probe -> Measure token probability drift
  -> Down-weight advantage on high-drift, fragile tokens
  -> Zero bonus credit for stability alone (avoids generic degeneration)
  -> +5.6% on AIME 2024-2026, uniform gains across all OOD math benchmarks

From “blindly rewarding all tokens in a successful trajectory,” the core shift is using semifactual causal invariance to filter out spurious tokens, ensuring that reinforcement learning only cements truly invariant reasoning.

Expert Assessment

Problem choice: Exceptional. GRPO has become the de facto default algorithm for reasoning post-training, but its uniform token credit assignment has always been an open theoretical compromise. Diagnosing and repairing its spurious prompt sensitivity tackles a fundamental flaw.

Method maturity: The causal formulation is disciplined. The authors’ design choice to use stability solely as a damping gate (rather than granting additive rewards for stability) is a vital insight that prevents the model from collapsing into conservative, repetitive gibberish.

Experimental integrity: Tested on frontier open-weights (Qwen3-4B-Base and Qwen3-1.7B-Base). Evaluating across three consecutive years of AIME competitions (2024, 2025, 2026) alongside exhaustive out-of-distribution math benchmarks confirms that the accuracy gains (+5.63% and +4.17%) reflect genuine mathematical generalization rather than test-set leakage.

Writing quality: Exemplary clarity. The causal diagrams and token-drift heatmaps illustrate the mechanism with undeniable transparency.

Verdict: strong accept — An essential upgrade to Group Relative Policy Optimization that should be adopted across all open-source reasoning model training pipelines.

Takeaways

  • In RLVR and GRPO, never assume every token in a correct reasoning trace deserves equal credit; uniform advantage breeds prompt fragility.
  • Use semifactual prompt variations (changing names, numbers, or phrasing while preserving truth) to measure which tokens are fragile artifacts.
  • Never reward stability directly; only use stability to down-weight advantages on fragile tokens during early optimization epochs.

论文: 2609.40360 作者: Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng 分类: cs.AI, cs.CL, cs.LG

缺口

基于可验证奖励的强化学习(RLVR)——尤其是以 DeepSeek-R1、QwQ 和 OpenR1 为代表所采用的组相对策略优化(GRPO)——推动大语言模型在数学与代码逻辑推理领域实现了前所未有的爆发。 通过在一个 Prompt 下采样一组候选解题轨迹并进行组内优势归一化,模型自发学会了试错、回溯与深度探索。

然而,经过 GRPO 训练的模型暴露出一个极其脆弱的痼疾:对提示词中完全无关紧要的表层表述具有神经质般的极度敏感。 仅仅把题目背景故事里的角色名字从「张三」改成「李四」,或者稍微调换一下已知条件的叙述次序,原本能够轻松做对的 AIME 级别高难度竞赛题就会当场推理崩溃。

这病根深植于 GRPO 信用分配(Credit Assignment)的数学底层: GRPO 将整条轨迹推演出来的单一标量优势值 A^\hat{A},毫无差别地均摊给了生成序列里的每一个 Token。

如果某条长达两千词的推演轨迹最终碰巧答对了(Reward = 1),GRPO 会给其中的全部两千个词都打上相同的满额正向奖励。 这导致那些纯属巧合、完全过拟合于特定提示词句式的「虚假关联 Token」,获得了与真正攻克数学关隘的「核心逻辑 Token」完全等同的强化力度。

   GRPO 粗放式均摊奖励的致命陷阱

   输入题目:"农夫约翰养了 15 头牛和 7 只羊..."
              |
              v
   模型生成轨迹:
   [前 50 词:"让约翰的羊群记为 s..."]  <-- 纯属偶然的表层修辞废话!
   [中间 500 词:严密的因数分解与方程推导] <-- 真正决定生死的数学内核!
   [最终词:"答案是 22"]                  <-- 最终答对!(获得奖励 1.0)
              |
              v
   传统 GRPO 的粗暴更新:
     计算出优势值 A = +1.0,整齐划一地涂抹在全部 501 个 Token 上!
     -> 虚假的表层套话被赋予了与数学逻辑完全相等的反向传播梯度!
     -> 导致模型极度脆弱:一旦把「约翰」改成「玛丽」,整个推演瞬间崩盘!
              |
              v
   SCAPO 的因果手术刀:半事实提示词微扰(Semifactual Intervention)
     改变题干的人名与措辞,但严格保全底层的数学事实与数值关系不变
     探测哪些 Token 的生成概率在微扰下剧烈漂移 -> 定点削减虚假 Token 的收益!

增量

一句话: 通过引入保持数学事实严格不变的「半事实提示词微扰」来度量 Token 级别的因果稳定性,并在训练中外科手术式地削减不稳定 Token 的优势权重,SCAPO 根除了 GRPO 中的虚假关联毒瘤,在 Qwen3-4B 上将 AIME 2024~2026 年竞赛题的平均准确率大幅提升了 5.63 个百分点。

核心机制

SCAPO(半事实信用增强策略优化)在标准 GRPO 之上叠加了一层极具因果解释力的信用精细化过滤器:

  1. 半事实提示词不变性(Semifactual Invariance): 所谓半事实微扰,是指对题目文本进行表层重写 x→x~x \to \tilde{x}(修改无关的故事叙述、置换代数未知数符号、调整语法倒装),但从逻辑底层严格确保题目的数学公理与标准答案绝对不变: P(答案∣x)≡P(答案∣x~)P(\text{答案} | x) \equiv P(\text{答案} | \tilde{x})
  2. Token 级概率漂移测绘: 针对一条已经跑通的正确解题轨迹 y=(y1,y2,…,yT)y = (y_1, y_2, \dots, y_T),算法考察当提示词被替换为半事实变体 x~\tilde{x} 时,各个位置上 Token 的条件概率漂移幅度: Δt=∣log⁡πθ(yt∣x,y<t)−log⁡πθ(yt∣x~,y<t)∣\Delta_t = | \log \pi_\theta(y_t | x, y_{<t}) - \log \pi_\theta(y_t | \tilde{x}, y_{<t}) | 若某个位置的 Δt\Delta_t 极高,铁证如山地说明该 Token 严重寄生于特定的提示词措辞表象,而非源于坚不可摧的数学公理推演。
  3. 优势阻尼抑制(而非盲目加分): 极为关键的是,SCAPO 绝不给稳定的 Token 额外发放奖励加成(那会导致模型学会吐出大量毫无信息量的废话同义反复)。 它的设计极为克制:仅在训练初期对那些高漂移、不稳定的脆弱 Token 施加优势阻尼下调: A^tSCAPO=A^GRPO⋅σ(稳定性t)\hat{A}_t^{\text{SCAPO}} = \hat{A}^{\text{GRPO}} \cdot \sigma(\text{稳定性}_t) 虚假套话的梯度更新被就地打压,强迫模型将有限的参数容量全额倾注在真正具备泛化稳定性的核心数学推导上。
   SCAPO 细粒度 Token 级优势校正流程

   生成的 Token 流: [ 设 ]   [ 约翰 ]   [ 拥有 ]   [ x ]  ...  [ x = 15 + 7 ]
                      |        |          |        |              |
   半事实微扰漂移量: 低       极高       中等     极低           绝对为零!
                      |        |          |        |              |
   稳定性阻尼门控:   保留     大幅削减   适度削减 保留           全额保留!
   最终实际获得优势: +1.0     +0.1       +0.4     +1.0           +1.0
                      ^        ^                                  ^
                      |        |                                  |
   更新结局:虚假句式依赖被迅速剥离剪枝,真正的不变逻辑获得最强强化!

这里的核喻是特种兵教官在恶劣多变的天气模拟仓里考核狙击手。

  • GRPO 就像一个只看靶心环数的糊涂裁判:如果狙击手在晴空万里下穿了一件花哨的红色披风打出了十环,裁判便把功劳归结于这件红色披风,并要求他在未来的实战考核里必须焊死这件披风(均摊奖励诱发虚假依赖)。结果到了雨雪泥泞的实战环境,狙击手因披风挂枝直接暴露阵亡。
  • SCAPO 则是引入了风洞与暴雨模拟的硬核教官(半事实因果微扰):教官在测试中随机模拟侧风、降雨和强光,但射击靶位和物理弹道规律保持不变。 教官一眼就看穿:不管外面刮风还是下雨,扣动扳机的均匀呼吸与肌肉记忆(不变的数学逻辑)才是十环的真正原因,而那件红色披风(高漂移的表层词)在侧风中只会乱晃添乱。 教官毫不留情地剥夺披风的加分项,将全部奖赏集中在呼吸节奏与击发动作上,锻造出在任何极端未知战场都能一击必杀的真正神枪手。

关键概念

  • 半事实因果微扰(Semifactual Intervention):因果推断领域中的高级反事实形态,即改变先验条件中的无关因,验证结论果是否依然保持稳固不变。
  • Token 级信用分配(Token-Level Credit Assignment):告别以整条句子或段落为单位的粗放宏观打分,将因果权责精准下沉到每个离散词的微观神经元激发上。
  • 推理中的虚假依赖(Spurious Dependency):自回归模型将输出条件错误绑定在输入提示词的语序、特定标点或语气词上,产生形式上看似连贯实则弱不禁风的虚假逻辑。

框架转变

之前(GRPO 粗放的均摊奖励更新):
  结局答对 -> 整条两千词的轨迹全员获得标量 A = +1.0
  -> 模型死记硬背提示词的特定叙述句式与无关变量名称
  -> 泛化能力极为脆弱,换个题干问法便当场抓瞎崩溃

之后(SCAPO 基于半事实的因果精准授信):
  半事实因果微扰探测 -> 精确测定各个 Token 的概率漂移敏感度
  -> 对高度过拟合表层句式的虚假词施加优势阻尼削弱
  -> 严禁给稳定性单独发糖(杜绝保守废话塌缩)
  -> AIME 竞赛题平均准确率暴涨 5.63%,全面横扫 OOD 外部数学评测集

从「只要结局正确就盲目嘉奖轨迹里的每一个字」,核心转变在于:利用半事实因果微扰照妖镜,剔除混迹在成功轨迹中的虚假寄生 Token,确保强化学习只夯实真正具有宇宙不变性的逻辑内核。

专家评审

选题眼光: 极为深刻且直击要害。 GRPO 已成为全行业训练推理模型的通用底座,但其「一刀切」式信用分配的粗糙妥协始终是学术界的隐痛。 直面该瓶颈并给出因果级解法,极富学术与工业穿透力。

方法成熟度: 理论逻辑极为克制。 最精妙之处在于没有把稳定性做成加分项(那会诱导模型说车轱辘废话),而是将其设计为针对虚假词的阻尼衰减门控,展现了深厚的强化学习优化功力。

实验诚意: 在 Qwen3-4B 与 Qwen3-1.7B 两代基座上完成了极高难度的 AIME 20242026 三年竞赛全真试题评测。 在数学推理与全新分布外(OOD)基准上录得 45 个百分点的坚实净增长,彻底排除了数据污染的嫌疑。

写作功力: 因果图与 Token 级漂移热力图极具视觉冲击力,说理透彻,论证严丝合缝。

判决: 强接收 (strong accept) — 强化学习推理对齐领域的范式升级力作,是每位大模型后训练算法工程师必读的奠基性改良成果。

要点总结

  • 在使用 GRPO 或 PPO 训练推理大模型时,切勿盲信「成功轨迹中的每个词都有功」;粗放均摊只会把提示词的表层偏见烙进模型深层。
  • 引入半事实微扰(保留题目数学本质不变,扰动叙述外壳),利用条件概率漂移精准识别虚假寄生 Token。
  • 严禁对稳定性给予额外加分奖励;只应使用稳定性指标来对虚假词的优势值施加定向削减与降噪。