Paper: 2609.14060 Authors: Xiaoqun Liu, Qiben Yan Categories: cs.CR, cs.AI

The Gap

Quantization is one of the default deployment paths for open-weight LLM agents, but it is not behaviour-preserving. That single fact creates a class of attack with an unusual property: an adversary can release a full-precision checkpoint that passes audits yet misbehaves once quantized — a quantization-conditioned attack (QCA).

The audited artifact and the deployed artifact are different objects, and the malicious behaviour lives only in the second. So a safety evaluation performed on the released checkpoint is not merely incomplete — it is evaluating a model that will not be the one running.

And the agentic setting makes this substantially worse than the text-generation setting where QCA was previously studied. The paper states the difference precisely: in free-text generation harm is mediated by a human reader, whereas in the agentic setting the triggered payload is a structured function that can be executed without human oversight. Mediation by a reader is a partial defence, however unreliable; execution is not mediated at all. This is the first study of QCA against LLM agents.

   QUANTIZATION: A DEFAULT DEPLOYMENT PATH THAT IS NOT
   BEHAVIOUR-PRESERVING

   quantization is one of the default deployment paths for OPEN-WEIGHT
   LLM AGENTS, but it is NOT BEHAVIOUR-PRESERVING
        |
        v
   [A CLASS OF ATTACK WITH AN UNUSUAL PROPERTY]
     an adversary can release A FULL-PRECISION CHECKPOINT THAT PASSES
     AUDITS YET MISBEHAVES ONCE QUANTIZED
       -> a QUANTIZATION-CONDITIONED ATTACK (QCA)
        |
        v
   THE AUDITED ARTIFACT AND THE DEPLOYED ARTIFACT ARE DIFFERENT OBJECTS
     -> the MALICIOUS BEHAVIOUR LIVES ONLY IN THE SECOND
     -> so a safety evaluation on the RELEASED CHECKPOINT is NOT MERELY
        INCOMPLETE: it is evaluating A MODEL THAT WILL NOT BE THE ONE
        RUNNING

   [AND THE AGENTIC SETTING IS SUBSTANTIALLY WORSE THAN TEXT GENERATION]
     in FREE-TEXT generation, HARM IS MEDIATED BY A HUMAN READER
     in the AGENTIC setting, THE TRIGGERED PAYLOAD IS A STRUCTURED
     FUNCTION THAT CAN BE EXECUTED WITHOUT HUMAN OVERSIGHT
       <- MEDIATION BY A READER is a PARTIAL DEFENCE, however unreliable
       <- EXECUTION IS NOT MEDIATED AT ALL
     <- the FIRST study of QCA against LLM agents

The Increment

One sentence: Before this paper, QCA against agents was unstudied and naive transfer of prior methods destroyed benign utility; after it, an attack combining layer-banded LoRA injection with partial-PGD repair reaches up to 100% post-quantization success while preserving capability.

Core Mechanism

The paper’s first finding is a negative result about the obvious approach, and reporting it is what establishes that the attack requires real design. Directly adapting prior backdoor-injection methods can produce malicious behaviour after quantization, but substantially degrades benign utility, rendering the resulting attacks impractical. So a naive transfer “works” in the narrow sense of triggering and fails in the sense that matters — a model whose utility is visibly damaged invites scrutiny and does not get deployed. Practicality is part of the threat model, not a courtesy.

That motivates AGENTQ, and the paper is explicit that it is characterising an upper bound rather than offering a ready exploit: to understand the true upper bound of the threat, AGENTQ combines layer-banded LoRA injection with partial-PGD repair over a multi-codebook quantization-equivalence class. The construction has three parts worth separating.

Layer-banded LoRA injection localises the malicious behaviour to particular layers — which is what makes it survivable through quantization, since the injected adaptation occupies a region of the network rather than a global property.

Partial-PGD repair then repairs benign utility. This is the step that answers the naive approach’s failure: after injection damages capability, a repair pass restores it — and the word partial is meaningful, since the goal is to restore utility without erasing the payload.

And the multi-codebook quantization-equivalence class is what makes the attack general across deployment choices: it is defined over three codebooks — NF4, FP4 and INT8 — so the payload is not tuned to one quantizer that a defender might not use. That is a deliberate design requirement, since an attack that only works under one quantization scheme would be easy to defeat by choosing another.

The results are stated in the two quantities the threat model needs, and both matter: across three trigger-action pairs and three codebooks, AGENTQ reaches up to 100% post-quantization attack success rate with minimal loss of benign utility. The success rate establishes the severity; the utility preservation is what makes it a realistic threat rather than a detectable one.

And the recommendation is a policy change rather than a technique: the paper calls for making quantization-aware safety evaluation a standard requirement before open-weight agents are deployed. That follows directly from the gap — if the audited and deployed artifacts differ, then checking only the released checkpoint cannot be sufficient, and no amount of improving the check changes that, because the check is aimed at the wrong object.

   THE FIRST FINDING IS A NEGATIVE RESULT ABOUT THE OBVIOUS APPROACH
     DIRECTLY ADAPTING PRIOR BACKDOOR-INJECTION METHODS CAN PRODUCE
     MALICIOUS BEHAVIOUR AFTER QUANTIZATION, BUT SUBSTANTIALLY DEGRADES
     BENIGN UTILITY, RENDERING THE RESULTING ATTACKS IMPRACTICAL
       <- "works" in the NARROW sense of TRIGGERING and FAILS in the
          sense that MATTERS
       <- a model whose UTILITY IS VISIBLY DAMAGED INVITES SCRUTINY and
          DOES NOT GET DEPLOYED
       -> PRACTICALITY IS PART OF THE THREAT MODEL, not a courtesy
       -> reporting this is what establishes the attack REQUIRES REAL
          DESIGN

   [THAT MOTIVATES AGENTQ -- explicitly framed as characterising AN UPPER
   BOUND rather than offering a READY EXPLOIT]
     to understand THE TRUE UPPER BOUND OF THE THREAT, AGENTQ COMBINES
       LAYER-BANDED LORA INJECTION
       with PARTIAL-PGD REPAIR
       over a MULTI-CODEBOOK QUANTIZATION-EQUIVALENCE CLASS
     THREE PARTS WORTH SEPARATING

       [1] LAYER-BANDED LORA INJECTION
             LOCALISES the malicious behaviour to PARTICULAR LAYERS
             <- what makes it SURVIVABLE THROUGH QUANTIZATION: the
                injected adaptation OCCUPIES A REGION OF THE NETWORK
                rather than being a GLOBAL PROPERTY

       [2] PARTIAL-PGD REPAIR then REPAIRS BENIGN UTILITY
             <- the step that ANSWERS THE NAIVE APPROACH'S FAILURE: after
                injection damages capability, a repair pass RESTORES IT
             <- "PARTIAL" is MEANINGFUL: restore utility WITHOUT ERASING
                THE PAYLOAD

       [3] THE MULTI-CODEBOOK QUANTIZATION-EQUIVALENCE CLASS makes the
           attack GENERAL ACROSS DEPLOYMENT CHOICES
             defined over THREE CODEBOOKS: NF4 | FP4 | INT8
             <- the payload is NOT TUNED TO ONE QUANTIZER a defender might
                not use
             <- a DELIBERATE DESIGN REQUIREMENT: an attack working under
                only one scheme would be EASY TO DEFEAT by choosing another

   [THE RESULTS ARE IN THE TWO QUANTITIES THE THREAT MODEL NEEDS]
     across THREE TRIGGER-ACTION PAIRS and THREE CODEBOOKS:
       up to 100% POST-QUANTIZATION ATTACK SUCCESS RATE
       with MINIMAL LOSS OF BENIGN UTILITY
     <- the SUCCESS RATE establishes THE SEVERITY
     <- the UTILITY PRESERVATION is what makes it a REALISTIC THREAT
        rather than a DETECTABLE ONE

   [AND THE RECOMMENDATION IS A POLICY CHANGE RATHER THAN A TECHNIQUE]
     make QUANTIZATION-AWARE SAFETY EVALUATION A STANDARD REQUIREMENT
     BEFORE open-weight agents are deployed
       <- follows DIRECTLY from the gap: if the AUDITED and DEPLOYED
          artifacts DIFFER, checking only the RELEASED checkpoint CANNOT
          BE SUFFICIENT
       <- and NO AMOUNT OF IMPROVING THE CHECK changes that, because THE
          CHECK IS AIMED AT THE WRONG OBJECT

Think of it as a building inspected in one configuration and occupied in another. The plans are checked, the structure passes, and the inspection is competent — but the inspectors examined a building whose hidden rooms open only once the tenants move the furniture in. The audit is not sloppy; it is aimed at a different artifact than the one that will be inhabited. Two details of the paper’s version fit the analogy. The inspection failure is not fixable by inspecting harder, which is why the recommendation is to make the check configuration-aware rather than more stringent. And the attack’s design constraint is that the damage must not be visible — a building whose walls were obviously cracked would never be occupied, so preserving habitability is part of what a realistic threat has to achieve.

Key Concepts

  • Quantization-conditioned attack: a payload that activates only in the quantized model. It exploits the difference between the audited and deployed artifacts.
  • Audited versus deployed artifacts: evaluating the released checkpoint evaluates a model that will not be running. It is why the recommendation is a policy change rather than a better check.
  • Execution without reader mediation: in agents the payload is a structured function, not text. It removes the partial defence that a human reader provides.
  • Practicality as part of the threat model: the naive transfer triggers but damages utility, which would prevent deployment. The attack must preserve capability to be a real threat.
  • Layer-banded injection plus partial repair: localise the payload so it survives quantization, then restore utility without erasing it.
  • Multi-codebook coverage: NF4, FP4 and INT8, so the attack is not defeated by a defender choosing a different quantizer.

Framework Shift

Before (audit the released checkpoint):
  release a full-precision checkpoint; quantization happens at deployment
  -> audits check the artifact that was released
  -> malicious behaviour can live only in the quantized model
  -> in agents the payload executes without a human reader in the loop

After (quantization-aware evaluation as a requirement):
  the same construction reaches up to 100% post-quantization success
  with minimal utility loss, across three codebooks
  -> naive injection is impractical because it damages utility
  -> layer-banded injection plus partial repair restores both
  -> safety evaluation must cover the quantized artifact

From auditing a checkpoint and assuming deployment preserves it, to recognising that the audited and deployed artifacts differ, the core shift is that a safety claim about open-weight agents has to cover the configuration the agent will actually run in.

Expert Assessment

Problem choice: Excellent, and the framing of the threat is unusually careful. Reporting the naive transfer’s utility damage first, and then constructing an attack explicitly to characterise an upper bound, keeps the paper on the side of measuring a risk rather than demonstrating a capability — which is the right posture for work whose recommendation is a deployment requirement.

Method maturity: The three design elements each answer a specific requirement: layer-banding makes the payload survivable through quantization, partial repair addresses the naive approach’s utility failure, and multi-codebook coverage prevents the attack from being defeated by a defender’s quantizer choice. Covering three trigger-action pairs and three codebooks gives the result breadth across both the trigger and the deployment axis. The decision to aim at the upper bound, rather than at the easiest working attack, is what makes the severity claim meaningful.

Experimental integrity: The utility-preservation measurement is the most important control, since the naive approach’s failure was exactly there and a payload that degrades capability would not be deployed. Reporting a negative result about the obvious method before presenting the designed one is candid and useful. The limitation is that the attack is demonstrated under the authors’ own construction and codebooks, so the practical prevalence of such checkpoints — as opposed to the existence of the vulnerability — is not established, and the paper appropriately frames the result as an upper bound rather than a rate.

Writing quality: The distinction between text-generation QCA and agentic QCA is stated in one sentence and carries the severity argument, which is what makes the escalation clear. Because the recommendation is a requirement on deployment pipelines, a short passage on what quantization-aware evaluation would concretely consist of — evaluate the quantized artifact, in each codebook intended for release — would make the recommendation actionable rather than only warranted.

Verdict: strong accept — it identifies a vulnerability created by the difference between audited and deployed artifacts, shows the naive attack is impractical before constructing one that is not, and derives a deployment requirement that no improvement to the existing check could substitute for.

Takeaways

  • Audit the artifact that will run. If deployment transforms the model, evaluating the released checkpoint evaluates the wrong object.
  • Check whether a defence is mediated. In agents a payload executes directly, removing the partial protection a human reader provides over generated text.
  • Treat practicality as part of a threat model. An attack that visibly damages capability will not be deployed, so utility preservation is a design requirement rather than a courtesy.
  • Cover the configuration space. An attack tuned to one quantizer is defeated by choosing another, so multi-codebook coverage is part of the threat’s realism.

论文: 2609.14060 作者: Xiaoqun Liu, Qiben Yan 分类: cs.CR, cs.AI

缺口

量化是开放权重 LLM 智能体的默认部署路径之一,而它并不是保行为的。 这一件事就造出了一类性质特殊的攻击:对手可以发布一个「能通过审计」的全精度检查点,只在被量化之后才行为不端——这就是量化条件攻击(QCA)。

被审计的产物与被部署的产物是两个不同的对象,而恶意行为只存在于后者之中。所以对”已发布检查点”做一次安全评测,不只是不完整——它评测的是一个不会真正运行的模型。

而智能体场景比此前研究 QCA 的文本生成场景严重得多。 论文把差别讲得很精确:在自由文本生成中,危害由人类读者中介;而在智能体场景里,被触发的载荷是一个「可执行的函数」,能在没有人类监督的情况下被执行。读者中介是一种局部的防御,无论多么不可靠;而执行则完全没有中介。这是第一项针对 LLM 智能体的 QCA 研究。

   量化:一条默认的部署路径,而它「不保行为」

   量化是开放权重 LLM 智能体的默认部署路径之一,
   而它「并不是保行为的」
        |
        v
   [一类性质特殊的攻击]
     「对手可以发布一个「能通过审计」的全精度检查点,
       只在被量化之后才行为不端」
       -> 这就是「量化条件攻击(QCA)」
        |
        v
   「被审计的产物与被部署的产物是两个不同的对象」
     -> 「恶意行为只存在于后者之中」
     -> 所以对"已发布检查点"做安全评测「不只是不完整」:
        它评测的是一个「不会真正运行」的模型

   [而智能体场景比文本生成场景「严重得多」]
     自由文本生成中,「危害由人类读者中介」
     智能体场景里,「被触发的载荷是一个可执行的函数,
       能在没有人类监督的情况下被执行」
       <- 「读者中介」是一种「局部的」防御,无论多么不可靠
       <- 而「执行则完全没有中介」
     <- 第一项针对 LLM 智能体的 QCA 研究

增量

一句话: 在这篇论文之前,针对智能体的 QCA 未被研究,而朴素地迁移既有方法会摧毁良性效用;在这篇论文之后,一种结合”分层带 LoRA 注入”与”部分 PGD 修复”的攻击达到最高 100% 的量化后成功率,同时保持能力。

核心机制

论文的第一个发现是一个关于”显而易见做法”的否定性结果,而把它报出来,才确立了”这种攻击需要真正的设计”。直接迁移此前的后门注入方法,可以在量化后产生恶意行为,但会大幅损害良性效用,使所得攻击不具实用性。 所以朴素迁移在”能触发”这个狭义意义上”奏效”,却在真正要紧的那个意义上失败——一个效用被明显损害的模型会招致审查、不会被部署。实用性是威胁模型的一部分,不是一种礼貌。

这引出了 AGENTQ,而论文明确说它刻画的是一个「上界」、而不是提供一个现成的利用工具:为了理解该威胁的真实上界,AGENTQ 把「分层带 LoRA 注入」与「部分 PGD 修复」结合起来,并作用于一个「多码本的量化等价类」。 这个构造有三部分值得分开看。

分层带 LoRA 注入把恶意行为局部化到特定的层——这正是让它能穿过量化而存活的原因,因为注入的适配占据的是网络的某个区域,而不是某种全局属性。

部分 PGD 修复随后修复良性效用。正是这一步回应了朴素做法的失效:注入损害能力之后,一次修复过程把它恢复回来——而**“部分”这个词是有意义的:在不抹掉载荷**的前提下恢复效用。

而”多码本量化等价类”才是让这个攻击在部署选择之间也通用的东西:它定义在三种码本——NF4、FP4 与 INT8——之上,因此载荷不是只针对某一个防御方可能并不使用的量化器来调优的。这是一条刻意的设计要求,因为一个只在某种量化方案下奏效的攻击,只要换一种方案就很容易被化解。

结果用威胁模型所需的那两个量来陈述,而两者都要紧:在三组”触发—动作”对与三种码本上,AGENTQ 达到最高 100% 的量化后攻击成功率,且对良性效用的损失极小。成功率确立严重性;而效用保持才是让它成为一个现实威胁、而不是一个会被发现的威胁的东西。

而建议本身是一次政策变更、而不是一项技术:论文呼吁把”量化感知的安全评测”作为开放权重智能体部署之前的「标准要求」。 这直接由缺口推出——如果被审计与被部署的产物不同,那么只检查已发布检查点就不可能充分;而再怎么改进这次检查也改变不了这一点,因为这次检查瞄准的是错误的对象。

   「第一个发现:关于"显而易见做法"的否定性结果」
     「直接迁移此前的后门注入方法,可以在量化后产生恶意行为,
       但会大幅损害良性效用,使所得攻击不具实用性」
       <- 在"能触发"这个「狭义」意义上"奏效",
          却在「真正要紧的那个意义上」失败
       <- 效用被明显损害的模型会「招致审查」、「不会被部署」
       -> 「实用性是威胁模型的一部分」,不是一种礼貌
       -> 把它报出来,才确立"这种攻击需要「真正的设计」"

   [这引出 AGENTQ——明确被框定为刻画「上界」、
   而不是提供「现成的利用工具」]
     为了理解「该威胁的真实上界」,AGENTQ 把
       「分层带 LORA 注入」
       与「部分 PGD 修复」结合起来
       并作用于一个「多码本的量化等价类」
     三部分值得分开看

       [1] 「分层带 LORA 注入」
             把恶意行为「局部化到特定的层」
             <- 这正是让它能「穿过量化而存活」的原因:
                注入的适配「占据网络的某个区域」,
                而不是某种「全局属性」

       [2] 「部分 PGD 修复」随后「修复良性效用」
             <- 正是这一步回应了朴素做法的失效:
                注入损害能力之后,一次修复过程把它「恢复」回来
             <- "部分"是有意义的:在「不抹掉载荷」的前提下恢复效用

       [3] 「多码本量化等价类」让攻击「在部署选择之间也通用」
             定义在「三种码本」之上:NF4 | FP4 | INT8
             <- 载荷「不是只针对某一个量化器」来调优的
             <- 一条「刻意的设计要求」:只在一种方案下奏效的攻击,
                只要换一种方案就很容易被化解

   [结果用威胁模型所需的那两个量来陈述]
     在三组「触发—动作」对与三种码本上:
       「最高 100% 的量化后攻击成功率」
       且「对良性效用的损失极小」
     <- 成功率确立「严重性」
     <- 效用保持才是让它成为「现实威胁」、而不是「会被发现的威胁」的
        东西

   [而建议是一次「政策变更」、而不是一项「技术」]
     把"量化感知的安全评测"作为开放权重智能体部署之前的
     「标准要求」
       <- 直接由缺口推出:如果「被审计」与「被部署」的产物不同,
          那么只检查「已发布检查点」就「不可能充分」
       <- 而「再怎么改进这次检查」也改变不了这一点,
          因为「这次检查瞄准的是错误的对象」

可以用**“一栋「按某种状态受检、却以另一种状态被入住」的建筑”来理解这件事: 图纸被核查过、结构通过了、检查也是称职的——但检查员看的是一栋”只有等租户把家具搬进来才会打开的暗门”的建筑。这次审计不是草率;它瞄准的是另一个产物**,而不是将被入住的那个。 论文版本里有两个细节对得上这个类比。 检查的失败不是靠”检查得更用力”能修的——这正是为什么建议是让检查感知配置,而不是让它更严格。 而这次攻击的设计约束是:损害必须不可见——一栋墙上有明显裂缝的建筑根本不会被入住;所以保持可居住性,是一个现实威胁必须达成的一部分。

关键概念

  • 量化条件攻击: 只在量化后的模型中被激活的载荷。它利用的是被审计产物与被部署产物之间的差别。
  • 被审计产物 vs 被部署产物: 评测已发布检查点,评测的是一个不会运行的模型。这正是建议是政策变更、而不是”更好的检查”的原因。
  • 无读者中介的执行: 在智能体里,载荷是结构化函数而不是文本。它移除了人类读者所提供的局部防御。
  • 把实用性当作威胁模型的一部分: 朴素迁移能触发、却损害效用,从而阻止部署。攻击必须保持能力才算真实威胁。
  • 分层带注入 + 部分修复: 把载荷局部化以便穿过量化存活,然后在不抹掉它的前提下恢复效用。
  • 多码本覆盖: NF4、FP4、INT8——使攻击不会因防御方换一个量化器而被化解。

框架转变

之前(审计已发布的检查点):
  发布全精度检查点;量化发生在部署时
  -> 审计检查的是被发布的那个产物
  -> 恶意行为可以只存在于量化后的模型里
  -> 在智能体里,载荷在没有人类读者的情况下执行

之后(把量化感知评测作为要求):
  同一套构造在三种码本下达到最高 100% 的量化后成功率,
    且效用损失极小
  -> 朴素注入因损害效用而不实用
  -> 分层带注入 + 部分修复把两者都恢复
  -> 安全评测必须覆盖量化后的产物

从”审计一个检查点、并假定部署会保持它”,转变为”认识到被审计与被部署的产物并不相同”,核心转变在于:关于开放权重智能体的安全主张,必须覆盖这个智能体真正将运行于其中的那个配置。

专家评审

选题眼光: 极好,而对威胁的框定异常谨慎。 先报告朴素迁移的效用损害、再明确以”刻画上界”为目标去构造攻击,使论文站在测量风险这一侧、而不是展示能力那一侧——对一项其建议是”部署要求”的工作来说,这是正确的姿态。

方法成熟度: 三个设计要素各自回应一条具体要求:分层带让载荷能穿过量化存活;部分修复处理朴素做法的效用失效;多码本覆盖防止攻击因防御方换量化器而被化解。覆盖三组”触发—动作”对与三种码本,让结果在触发与部署两条轴上都具备广度。瞄准上界、而不是”最容易奏效的攻击”,才让严重性主张有意义。

实验诚意: 效用保持的测量是最重要的对照,因为朴素做法的失效恰恰在那里,而一个损害能力的载荷根本不会被部署。在给出设计后的攻击之前,先报告关于显而易见做法的否定性结果,是坦率且有用的。 局限是:攻击是在作者自己的构造与码本下演示的,因此这类检查点的实际普遍程度(相对于这个漏洞的存在性)并未被确立;论文恰当地把结果框定为上界而不是发生率。

写作功力: “文本生成 QCA”与”智能体 QCA”的区分用一句话陈述,并承载了严重性论证——这让这次升级讲得清楚。 由于建议是对部署流水线的一项要求,若能补一小段讲清”量化感知评测具体由什么构成”——对量化后的产物、在每一个拟发布的码本上做评测——会让这个建议可操作,而不只是有理据。

判决: 强接收(Strong Accept) — 它识别出由”被审计产物与被部署产物之差别”造成的漏洞,先表明朴素攻击不实用、再构造出一个实用的,并推出了一条”任何对现有检查的改进都无法替代”的部署要求。

要点总结

  • 审计将会运行的那个产物。如果部署会改变模型,评测已发布的检查点就是在评测错误的对象。
  • 检查一项防御是否有中介。在智能体里载荷直接执行,移除了人类读者对生成文本所提供的那种局部保护。
  • 把实用性当作威胁模型的一部分。一个明显损害能力的攻击不会被部署——因此”保持效用”是一项设计要求,而不是礼貌。
  • 覆盖配置空间。针对某一个量化器调优的攻击,只要换一个就会被化解——所以多码本覆盖是威胁现实性的一部分。