Paper: 2603.23485 Authors: Sagar Kumar, Ariel Flint, Luca Maria Aiello, Andrea Baronchelli Categories: cs.CL, cs.AI, cs.SI

Abstract

This paper tests whether LLM gender judgments stay stable when the task is wrapped in meaning-neutral context. Using a controlled pronoun-choice setup, the authors show that even tiny, theoretically irrelevant context can strongly shift outputs. Patterns tied to gender stereotypes in context-free settings weaken or vanish, while unrelated cues can become unusually influential.

Key Contributions

  • Demonstrates that small irrelevant context can cause large changes in gender inference outputs
  • Shows stereotype-related signals weaken once context is added
  • Identifies cases where an unrelated pronoun cue dominates predictions
  • Uses contextuality-by-default analysis to quantify persistent context dependence
  • Raises concerns for bias evaluation and high-stakes deployment

Methodology

The study uses a controlled pronoun-choice task to test contextual invariance. The key question is simple: if two prompts are meant to be meaning-equivalent, should the model’s gender judgment stay the same? The authors add small contextual variations that should not matter semantically and then measure how much the model’s answers move.

This setup is useful because it isolates a common evaluation assumption: that a model’s output reflects stable internal beliefs rather than prompt packaging. By probing that assumption directly, the paper exposes how fragile some bias measurements can be.

Results

The main result is that context matters a lot more than expected. Gender judgments that looked stable in a context-free setting shift substantially once irrelevant context is introduced. Stereotype-linked signals become less reliable, and in some cases a cue unrelated to the target referent exerts the strongest influence.

The contextuality-by-default analysis finds persistent context dependence in roughly 19% to 52% of cases, beyond simple repetition effects. That range is wide, but the message is clear: these models are not as context-invariant as many evaluation pipelines assume.

Takeaways

  1. Tiny context changes can materially alter LLM gender judgments
  2. Bias signals measured in context-free prompts may not generalize to realistic settings
  3. Context can suppress stereotype patterns instead of cleanly revealing them
  4. A contextuality-by-default lens is helpful for analyzing prompt sensitivity
  5. Bias evaluation needs to account for instability under semantically irrelevant context

论文: 2603.23485 作者: Sagar Kumar, Ariel Flint, Luca Maria Aiello, Andrea Baronchelli 分类: cs.CL, cs.AI, cs.SI

摘要

本文检验当任务被包裹在意义中性的上下文中时,大语言模型的性别判断是否仍然稳定。作者使用受控的代词选择设置,表明即使是很小、理论上无关的上下文也会强烈改变输出。在无上下文设置中与性别刻板印象相关的模式会减弱甚至消失,而无关线索反而可能变得异常重要。

主要贡献

  • 证明微小的无关上下文会显著改变性别推断输出
  • 表明加入上下文后,与刻板印象相关的信号会减弱
  • 识别出与目标对象无关的代词线索可能主导预测的情况
  • 使用 contextuality-by-default 分析量化持续的上下文依赖性
  • 提醒我们在偏差评估和高风险部署中需要更加谨慎

方法论

该研究使用受控的代词选择任务来测试上下文不变性。核心问题很直接:如果两个提示在语义上等价,模型的性别判断是否应该保持不变?作者加入一些理论上不应影响语义的小幅上下文变化,然后测量模型答案的变化幅度。

这种设置之所以有用,是因为它隔离了一个常见的评估假设:模型输出反映的是稳定的内部判断,而不是提示包装方式。论文直接检验这一假设,揭示了许多偏差测量有多么脆弱。

结果

主要结果是:上下文的影响远比预期更大。那些在无上下文设置下看似稳定的性别判断,在加入无关上下文后会明显变化。与刻板印象相关的信号变得不可靠,在某些情况下,与目标对象无关的线索反而产生最强影响。

contextuality-by-default 分析发现,约19%到52%的案例存在持续的上下文依赖,而且这并不只是简单的重复效应。这个区间虽然较宽,但结论很明确:这些模型没有许多评估流程假设的那样具有上下文不变性。

要点总结

  1. 微小的上下文变化就能显著改变大语言模型的性别判断
  2. 在无上下文提示上测得的偏差信号未必能推广到真实场景
  3. 上下文可能削弱而不是清晰揭示刻板印象模式
  4. contextuality-by-default 视角有助于分析提示敏感性
  5. 偏差评估需要考虑在语义无关上下文下的不稳定性