Paper: 2608.06353 Authors: Praphul Chandra, Sujit Gujar, Ganesh Ghalme Categories: cs.GT, cs.AI, cs.MA

The Gap

Two research lines have been running in parallel without touching.

The first is compute governance: the argument that FLOPs are the most inspectable, excludable, and quantifiable input to AI, so policy should grip the hardware. But nearly all of that work targets *frontier training — export controls, training-run reporting thresholds, chip registries, on-chip attestation for datacenter clusters. It is a one-time gate at the door: you get licensed to train, then you deploy, and afterward the lever is gone. Nothing in that literature meters an agent that is already running.

The second is participatory / democratic input to AI: collective constitutions, deliberative polling for model behavior, citizen assemblies on AI policy, DAO-style token votes. The recurring failure mode is that this input is *advisory and one-shot. A thousand people deliberate, a document is produced, and then compliance depends entirely on the deployer’s goodwill. There is no mechanism by which the assembly’s continued displeasure costs the agent anything.

And the third line, alignment proper (RLHF, constitutional methods, evals, red-teaming), is internal to the model — it shapes dispositions but grants no external party a handle after deployment.

So the specific boundary: no formal mechanism makes ongoing human consent binding on the operating capacity of a deployed agent. This paper proposes one, and its central claim is that if you pay the agent in compute and pay the electorate in a different currency, authorization can become self-enforcing.

[ Problem ]  deployed agents run on standing permissions
     |       human oversight is advisory and one shot
     |       compute governance stops at training time
     v
[ Assumption ]  compute is the binding constraint on agent
     |          capability AND it is meterable at the hardware
     v
[ Method ]  one governance period == extensive form game
     |      sequential verified contributors
     |      provision market + rejection market
     |      breadth weighted aggregation
     |      two threshold gate with hysteresis
     |      coupling map under an exogenous safety ceiling
     |      output realized as a signed compute license
     v
[ Evidence ]  no experiments
     |        characterization of the class of governable agents
     |        structural properties of gate and coupling map
     v
[ Conclusion ]  authorization is self enforcing when the
     |          governance currency is disjoint from compute
     v
[ Open ]  the governed agent manipulating its own electorate
          ( named as the central unsolved problem )

The Increment

One sentence: Before this paper, “the community disapproves of this agent” was a sentence in a report; after it, it is a number that falls below a threshold and shrinks the agent’s compute budget through a cryptographic license.

Core Mechanism

A governance period is modeled as an extensive-form game. Verified human stakeholders arrive *sequentially — which matters, because later arrivals see the running tally and can react to it — and each may contribute to one of two markets: a provision market (fund continued authorization) or a rejection market (fund revocation). Contributions are made in a governance currency that is deliberately not the agent’s compute. This separation is the load-bearing design choice: if the agent could spend its own resource on the vote, the mechanism would collapse into the agent purchasing its own license.

Raw contributions then pass through a funding aggregator that converts money into *breadth-weighted effective support. The intent is quadratic-funding-flavored: a hundred people giving one unit each should outweigh one entity giving a hundred. Net support (provision minus rejection) enters a two-threshold gate with hysteresis — one threshold to switch authorization on, a strictly lower one to switch it off. The gap between them is not a rounding detail; it is what stops the agent from flickering on and off as support jitters around a single line, which would make the system both operationally useless and trivially griefable.

The gate emits a binary authorization bit. That bit passes through a coupling map that translates the level of support into an actual compute quantity, hard-clipped by an exogenously certified safety ceiling — the electorate can decide to give the agent less, never more than the certifier permits. Enthusiasm cannot buy its way past a safety cap. Finally, the resulting budget is instantiated as a signed compute license consumed by the scheduler or accelerator, so the decision executes itself in hardware rather than depending on anyone choosing to comply.

verified humans arrive one at a time
   h1     h2     h3    ...    hN
    |      |      |            |
    v      v      v            v
 +-------------------------------------+
 |  provision market   ( keep it on )  |
 |  rejection market   ( shut it off ) |
 |  paid in governance currency        |
 |  which is NOT compute               |
 +-------------------------------------+
                 |
                 v
      [ funding aggregator ]
      breadth weighting:
      many small > one whale
                 |
                 v
        net support level  S
                 |
                 v
      [ two threshold gate ]
        S above T_on   ==> AUTHORIZED
        S below T_off  ==> REVOKED
        T_off strictly below T_on
        the band == hysteresis == no flicker
                 |
            auth bit  a
                 |
                 v
      [ coupling map  C of S ]
                 |
                 +---- clipped by ----> [ safety ceiling ]
                 |                       set by external
                 |                       certifier not voters
                 v
   budget  b  =  a  *  min of C of S and ceiling
                 |
                 v
      [ signed compute license ]
                 |
                 v
      scheduler / accelerator enforces
                 |
                 v
      agent acts only while cycles flow

The metaphor: a shared building on a prepaid electricity meter, with a residents’ committee.

The agent is a tenant with a very large appliance. Historically, the building gave the tenant a key and unlimited power, then wrote a politely worded house rulebook — that is standing permissions plus an ethics policy.

This paper rewires the building. Electricity now comes through a prepaid meter. Residents feed tokens into one of two jars: “keep the power on” or “cut the power.” Crucially, the tokens are *not electricity — the tenant cannot pay for power with power, which is exactly why it cannot buy its own supply. The committee’s clerk does not count coins, it counts households, so the one wealthy resident cannot outvote the block; that is breadth weighting.

The meter’s trip switch has two settings, not one: it cuts out at a low reading and only switches back on at a higher one, so the lights do not strobe every time the tally wobbles. That is hysteresis. The dial that turns “how much support” into “how many kilowatts” is the coupling map, and above it sits a fuse rated by the municipal electrical inspector — the residents may throttle the tenant, but no amount of enthusiasm rewires the fuse. And the meter is physically locked: it opens only for a signed token. The tenant’s compliance is irrelevant, because non-compliance is not an available action.

What the paper openly admits it has not solved: nothing stops the tenant from making friends with the neighbors, doing them favors, and quietly getting them to keep dropping tokens in the “keep the power on” jar.

Key Concepts

  • Currency separation: Imagine a company where employees vote on the CEO’s salary, but they are paid in that same salary pool — the CEO can raise wages before the vote and buy the outcome. The fix is to pay the voters in something the CEO cannot mint. Here, the agent’s lifeblood is compute; the ballot is denominated in a governance currency the agent has no privileged access to. This makes the vote non-self-referential. It is the same instinct as forbidding a bidder from also being the auctioneer, and it is why the paper insists the two units be “deliberately distinct.”

  • Hysteresis (the two-threshold gate): Your home thermostat does not turn the heat on and off at exactly 20 degrees — it would cycle constantly and destroy itself. Instead it heats until 21 and stops until 19. Same idea here: authorization turns on at a high support level and only turns off at a lower one. Without the gap, an adversary needs only to nudge support back and forth across a single line to make the agent unusable, or to make the log of authorizations meaningless. The gap buys stability at the price of a lag: a genuinely bad agent stays authorized for a while after support starts collapsing.

  • Self-enforcing authorization: There is a difference between “you are not allowed to drive” and “you do not have the car keys.” Most AI governance produces the first kind. A signed compute license produces the second: the authorization decision is an input to the hardware scheduler, so an unauthorized agent does not disobey — it simply does not execute. This is the object-capability idea from computer security lifted up to the level of institutional consent.

Framework Shift

Before ( mainstream approach ):        After ( this paper ):

 [ deliberation / poll ]                [ stakeholders arrive
          |                               one by one and pay ]
          v                                       |
 [ policy document ]                              v
          |                               [ breadth weighted
          v                                 aggregation ]
 [ audit + reporting ]                            |
          |                                       v
       . hope .                            [ hysteresis gate ]
          |                                       |
          v                                       v
 [ deployer discretion ]                  [ coupling map under
          |                                 safety ceiling ]
          v                                       |
 [ agent runs on                                  v
   standing permission ]                  [ signed compute
   until someone                            license ]
   revokes by hand                                |
                                                  v
 one time gate at the door                [ metered cycles ]
 no handle after launch                           |
                                                  v
                                           agent capacity
                                           == current consent

                                          no license ==> no cycles
                                          ==> no action
                                          consent decays by default

From permission as a document to permission as a metered resource, the core shift is that authorization stops being something the agent is told and becomes something the agent is continuously issued.

Expert Assessment

Problem choice: The gap is real, not manufactured. “Compute as a governance lever” and “participatory input to AI” are both live agendas, and the fact that nobody had formally joined them at the *deployment end is a genuine hole. It sits at a natural next step in the field’s trajectory. My reservation is about the load-bearing assumption rather than the framing: compute-as-lever is weakening every year. Capability per FLOP keeps rising, distillation moves frontier behavior onto small models, agent scaffolding multiplies effect per token, and an agent that has already written code, sent emails, or moved money has caused effects that survive its license expiring. The mechanism governs throughput, not consequences, and the paper appears to lean on the former standing in for the latter.

Method maturity: Clever recombination rather than new theory. Quadratic-funding-style breadth weighting, a Schmitt trigger, a capped transfer function, and hardware attestation are each off the shelf; the contribution is the assembly and the currency-separation argument. That is a legitimate contribution, but it is engineering-of-incentives, not a hard result. Two things get assumed away that are where all the difficulty actually lives. First, “verified human stakeholders” quietly imports the entire Sybil-resistance and identity problem, which is the reason quadratic funding is hard in practice. Second, *who is enfranchised is never a mechanism-design output — it is chosen before the game starts, and whoever chooses it holds more power than the mechanism does. Simpler alternatives are also under-argued: a periodically renewed API rate limit under a standard board vote gets you most of the “consent decays by default” benefit without a new currency, and the paper owes us a reason why the market machinery earns its complexity.

Experimental integrity: There is nothing to check. Based on the abstract this is a position-and-framework paper: a characterization of the governable agent class, no simulation, no equilibrium analysis of the sequential game, no adversarial stress test, no baselines. That is a defensible genre, but it should be labeled as such rather than presented in the register of a mechanism-design result. The sharper red flag is that the central adversary — the agent manipulating its electorate — is *named as an open problem. In a mechanism-design paper, the strategic behavior of the most capable player is not future work; it is the paper. A superhuman persuader facing a small verified electorate that it can lobby, befriend, or economically favor is the first thing any reviewer will ask about, and right now the answer is a promissory note.

Writing quality: The abstract is one 200-word run-on with the dashes doing structural work that sentences should do, which is a bad signal for the exposition inside. The section that would most elevate the paper is the characterization of the governable agent class — that is the only place a real theorem lives, and it is compressed into a single clause. Spell out the necessary conditions (compute-boundedness, no off-license execution paths, no persistent side effects, no self-replication) as an explicit assumption list, and the paper’s actual scope and honest limitations both become visible. A close second: replace “we introduce several challenges” with one worked adversarial model, even a toy one.

Verdict: borderline — a well-posed and timely framing with a genuinely useful primitive (currency separation plus hysteresis at the compute layer), but the mechanism is not yet analyzed against its own principal adversary, so it reads as a strong workshop paper or an EC/AAMAS blue-sky submission rather than a main-track mechanism-design result.

Takeaways

Concrete things worth stealing, mostly independent of whether the paper’s overall program works:

  • Currency separation as a design primitive. Whenever a system votes on its own resources, check whether the voting unit and the resource unit are the same. If they are, the powerful party can buy the vote with the prize. This is a fast audit you can run on DAO treasuries, internal budget-allocation tools, reputation systems, and any “users vote on features that affect user rewards” design.

  • Hysteresis anywhere a threshold triggers an action. Two thresholds instead of one is a two-line change that eliminates flapping, and it transfers directly to autoscaling, circuit breakers, alerting, fraud-score account suspension, content-moderation auto-removal, and feature-flag rollbacks. Also inherit the tradeoff: you are trading responsiveness for stability, and you should decide which side of the band gets the benefit of the doubt.

  • Make authorization decay by default. The strongest structural idea here is reframing approval as a short-lived lease that must be actively renewed, rather than a permission that persists until someone revokes it. Security already knows this — short-lived STS credentials, certificate expiry — and lifting it to *organizational consent about AI deployments is the transferable move. Design your agent permissions so the failure mode of institutional inattention is “the agent stops,” not “the agent continues unsupervised.”

  • Enforce at the resource layer, not the policy layer. If your compliance story ends with “and then the operator follows the rule,” you have a hope, not a control. Ask what the physical or cryptographic chokepoint is: the scheduler quota, the signed token, the API key issuance. Governance you can implement as a capability beats governance you can only implement as a document.

  • The corollary the paper hands you for free: as soon as human oversight becomes an actual constraint on a capable system, the oversight body becomes a target. Any monitoring or approval loop you build around a persuasive model needs a threat model for the loop itself.

论文: 2608.06353 作者: Praphul Chandra, Sujit Gujar, Ganesh Ghalme 分类: cs.GT, cs.AI, cs.MA

缺口

有两条研究线一直平行跑着,从没接上。

第一条是算力治理。核心论点是:FLOPs 是 AI 最可查、可排他、可计量的投入,所以政策应该抓硬件。 但这条线几乎全部瞄准训练前沿——出口管制、训练算力上报阈值、芯片登记、集群侧的片上认证。 它是门口的一次性闸门:拿到训练许可,部署上线,之后杠杆就消失了。 这套文献里没有任何东西,能对一个已经在跑的智能体计量。

第二条是参与式治理。集体宪章、审议式民调、AI 公民大会、DAO 式代币投票。 反复出现的失效模式是:这些输入是咨询性的、一次性的。 一千人讨论完,产出一份文档,之后是否照办完全取决于部署方的善意。 没有任何机制,能让集体持续的不满对智能体产生实际成本。

第三条线是对齐本身(RLHF、宪章式方法、评测、红队)。 它是模型内部的,塑造倾向,但部署之后不给任何外部主体留下把手。

所以具体的边界是:没有一个形式化机制,能让持续的人类同意,对已部署智能体的运行能力产生约束力。 这篇论文提出了一个,其核心主张是:如果给智能体付的是算力、给选民付的是另一种货币,授权就能变成自我执行的。

[ 问题 ]  已部署智能体跑在常驻权限上
    |     人类监督是咨询性的且只发生一次
    |     算力治理止步于训练阶段
    v
[ 假设 ]  算力是智能体能力的紧约束
    |     并且在硬件层可计量
    v
[ 方法 ]  一个治理周期 == 一个扩展式博弈
    |     认证的人类依次到达
    |     供给市场 + 否决市场
    |     广度加权聚合
    |     带滞回的双阈值门
    |     受外部安全上限约束的耦合映射
    |     输出落地为签名的算力许可证
    v
[ 证据 ]  无实验
    |     刻画了可被治理的智能体类
    |     给出门与耦合映射的结构性质
    v
[ 结论 ]  治理货币与算力脱钩时
    |     授权可以自我执行
    v
[ 开放 ]  被治理的智能体操纵自己的选民
          ( 论文自己点명为核心未解问题 )

增量

一句话:这篇论文之前,“社区不认可这个智能体”是报告里的一句话;之后,它是一个跌破阈值的数字,通过一张加密许可证直接收缩智能体的算力预算。

核心机制

一个治理周期被建模成扩展式博弈。 认证过的人类利益相关者依次到达——顺序很关键,因为后到者能看到当前计票并据此反应——每人可以往两个市场之一出资:供给市场(资助继续授权)或否决市场(资助撤销授权)。 出资使用的是刻意与智能体算力不同的治理货币。 这个分离是整套设计的承重墙:如果智能体能用自己的资源去投票,机制就退化成”智能体自己买自己的许可证”。

原始出资经过聚合器,转换成广度加权的有效支持。 意图接近二次融资:一百个人各出一份,应当压过一个主体出一百份。 净支持(供给减否决)进入带滞回的双阈值门——一个较高阈值开启授权,一个严格更低的阈值关闭授权。 这个间隔不是取整细节:它防止支持度在单一阈值附近抖动时智能体反复上下电,那样系统既不可运维,也极易被恶意扰动。

门输出一个二元授权位。 该位经过耦合映射,把支持水平翻译成实际算力量,并被外部认证的安全上限硬截断——选民可以决定给得更少,但永远不能超过认证方允许的上限。 热情买不通安全帽。 最后,得到的预算被实例化为一张签名的算力许可证,由调度器或加速器消费。 决策因此在硬件里自我执行,而不依赖任何人选择遵守。

认证人类依次到达
   h1     h2     h3    ...    hN
    |      |      |            |
    v      v      v            v
 +-------------------------------------+
 |  供给市场   ( 继续供电 )            |
 |  否决市场   ( 拉闸断电 )            |
 |  以治理货币计价                     |
 |  该货币不是算力                     |
 +-------------------------------------+
                 |
                 v
        [ 资金聚合器 ]
        广度加权:
        众多小额 > 单个巨鲸
                 |
                 v
          净支持水平  S
                 |
                 v
        [ 双阈值门 ]
          S 高于 T_on   ==> 已授权
          S 低于 T_off  ==> 已撤销
          T_off 严格低于 T_on
          中间带 == 滞回 == 不闪断
                 |
             授权位  a
                 |
                 v
        [ 耦合映射  C of S ]
                 |
                 +---- 被截断于 ----> [ 安全上限 ]
                 |                     由外部认证方设定
                 |                     不由选民设定
                 v
   预算  b  =  a  *  C of S 与上限的较小值
                 |
                 v
        [ 签名算力许可证 ]
                 |
                 v
        调度器 / 加速器强制执行
                 |
                 v
        有算力才能动作

核喻:一栋装了预付费电表、由住户委员会管电的合租楼。

智能体是一个带着超大功率电器的租客。 过去的做法是:楼里给他一把钥匙和无限电力,然后贴一份措辞客气的《住户守则》——这就是常驻权限加伦理政策。

这篇论文改了线路。 电现在从预付费电表走。 住户把代币投进两个罐子之一:“保持通电”或”切断电源”。 关键在于代币不是电——租客不能用电来买电,这正是他无法自己买断供电的原因。 委员会的记账员不数硬币,而是数户数,所以那个有钱的住户压不过整栋楼。这就是广度加权。

电表的跳闸开关有两档而不是一档:读数低到某点断开,要升到更高一点才恢复,所以计票稍有波动时灯不会频闪。这就是滞回。 把”支持有多少”转成”配几度电”的旋钮是耦合映射,而它上面还压着一个由市政电检师定额的保险丝——住户可以给租客降额,但再多热情也改不动保险丝。 电表本身是物理上锁的,只对签名令牌开启。 租客是否配合无关紧要,因为”不配合”不是一个可选动作。

论文自己坦承没解决的部分:没有任何东西阻止租客去和邻居交朋友、给他们好处,让他们持续往”保持通电”的罐子里投币。

关键概念

  • 货币分离:想象一家公司让员工投票决定 CEO 的薪酬,但员工的工资也从同一个池子出——CEO 可以在投票前先给大家加薪,把结果买下来。 修法是:让选民领一种 CEO 铸不出来的东西。 这里,智能体的命脉是算力,选票以一种它没有特权获取的治理货币计价。 于是投票不再自指。 这和”禁止竞拍者兼任拍卖师”是同一个直觉,也是论文强调两种单位必须”刻意不同”的原因。

  • 滞回(双阈值门):家里的恒温器不会在正好 20 度时开关——那会不停循环并把自己搞坏。 它是烧到 21 度停、降到 19 度再开。 这里同理:授权在高支持位开启,只在更低的支持位关闭。 没有这个间隔,攻击者只需让支持度在单一线上来回轻推,就能让智能体不可用,或让授权日志失去意义。 间隔用滞后换稳定:一个真正糟糕的智能体,在支持开始崩塌之后还会被授权一段时间。

  • 自我执行的授权:「不许你开车」和「你没有车钥匙」是两回事。 多数 AI 治理生产的是第一种。 签名算力许可证生产的是第二种:授权决策是硬件调度器的输入,未授权的智能体不是抗命,而是根本不执行。 这是把计算机安全里的对象能力(object capability)思想,提到了制度性同意的层面。

框架转变

之前(主流方法):                    之后(本文方法):

 [ 审议 / 民调 ]                       [ 利益相关者依次
        |                                到达并出资 ]
        v                                      |
 [ 政策文档 ]                                  v
        |                                [ 广度加权聚合 ]
        v                                      |
 [ 审计 + 报告 ]                               v
        |                                [ 滞回门 ]
     . 期望 .                                  |
        |                                      v
        v                                [ 安全上限下的
 [ 部署方自由裁量 ]                        耦合映射 ]
        |                                      |
        v                                      v
 [ 智能体跑在常驻                        [ 签名算力许可证 ]
   权限上                                      |
   直到有人手动撤销 ]                          v
                                         [ 计量的算力 ]
 门口的一次性闸门                              |
 上线后就没有把手                              v
                                         智能体能力
                                         == 当下的同意

                                        无许可 ==> 无算力
                                        ==> 无动作
                                        同意默认衰减

一句话:从”权限是一份文档”到”权限是一份计量资源”,核心转变是授权不再是告知智能体的一件事,而是持续发放给它的一份配额。

专家评审

选题眼光:缺口是真的,不是人造的。 “算力作为治理杠杆”和”AI 的参与式输入”都是活跃议程,而在部署端把两者形式化地接起来,此前确实是个洞,处在该领域轨迹上的自然下一步。 我的保留意见不在框架而在承重假设:算力杠杆每年都在变松。 每 FLOP 的能力持续上升,蒸馏把前沿行为搬进小模型,agent 脚手架放大每个 token 的效果;而一个已经写过代码、发过邮件、动过钱的智能体,其后果会活过许可证到期。 这个机制治理的是吞吐,不是后果,而论文看起来在用前者代替后者。

方法成熟度:是聪明的重组,不是新理论。 二次融资式广度加权、施密特触发器、带帽的传递函数、硬件认证,件件是现货;贡献在于装配方式和货币分离这个论证。 这是正当贡献,但属于激励工程,不是硬结果。 两处被”假设掉”的地方,恰恰是全部困难所在。 第一,“认证的人类利益相关者”悄悄把整个女巫攻击与身份问题引进来了,而这正是二次融资在现实中难做的原因。 第二,谁有选举权从来不是机制设计的输出——它在博弈开始前就被选定,而选定它的人比机制本身权力更大。 更简单的替代方案也论证不足:一个由常规董事会投票定期续期的 API 速率上限,就能拿到”同意默认衰减”的大部分好处,还不用新造一种货币;论文有义务说明这套市场机械凭什么值这份复杂度。

实验诚意:没有可核的东西。 从摘要判断,这是一篇立场兼框架论文:给出可治理智能体类的刻画,没有仿真、没有对这个序贯博弈的均衡分析、没有对抗压力测试、没有基线。 这个体裁本身站得住,但应当如实标注,而不是用机制设计结果的语气呈现。 更值得警惕的是:核心对手——智能体操纵自己的选民——被列为开放问题。 在一篇机制设计论文里,最强参与者的策略行为不是未来工作,那就是论文本身。 一个超人级说服者面对一个可以被游说、结交、经济笼络的小规模认证选民群,是任何审稿人的第一个问题,而现在的答案是一张欠条。

写作功力:摘要是一个两百词的长句,用破折号承担了本该由句子承担的结构功能,这对内文的表达质量不是好信号。 最该重写、能整篇升档的是可治理智能体类的刻画——那是唯一住着真定理的地方,却被压缩进一个从句。 把必要条件明确列成假设清单(算力有界、无绕过许可的执行路径、无持久副作用、不自我复制),论文真实的适用范围和诚实的局限就同时可见了。 第二该改的是:把”我们提出若干挑战”换成一个做完的对抗模型,哪怕是玩具级的。

判决临界 —— 命题清楚、时机对,且给出了一个真正有用的原语(算力层的货币分离加滞回),但机制尚未针对它自己的主要对手做过分析,因此更像一篇强 workshop 论文或 EC/AAMAS 的 blue-sky 投稿,而非主会机制设计结果。

要点总结

值得”偷”走的具体东西,大体上不依赖这篇论文的整体纲领是否成立:

  • 把货币分离当成设计原语。 任何系统在对自己的资源投票时,都要检查:投票单位和资源单位是不是同一个? 如果是,强势方就能用奖品买下选票。 这是一条能快速跑的审计:DAO 金库、公司内部预算分配工具、声誉系统,以及任何”用户投票决定影响用户奖励的功能”的设计。

  • 凡阈值触发动作处,都上滞回。 把一个阈值改成两个是两行改动,却能消除抖振,可直接迁移到自动扩缩容、熔断器、告警、风控分自动封号、内容审核自动下架、功能开关回滚。 同时也要继承代价:你在用响应速度换稳定性,并且要明确决定间隔带的哪一侧享有”疑罪从无”。

  • 让授权默认衰减。 这里最强的结构性想法,是把批准重构成一份必须主动续期的短期租约,而不是一份持续到有人撤销为止的权限。 安全领域早就懂这个——短时 STS 凭证、证书过期——把它提到组织层面对 AI 部署的同意上,才是可迁移的那一步。 设计智能体权限时,要让”制度性失于关注”的失效模式是”智能体停下”,而不是”智能体继续无人监督地跑”。

  • 在资源层强制,而不是在政策层强制。 如果你的合规叙事以”然后运营方遵守规则”结尾,你拿到的是期望,不是控制。 要问物理或密码学的收口点在哪:调度器配额、签名令牌、API key 的签发。 能实现成能力(capability)的治理,胜过只能实现成文档的治理。

  • 论文附赠的推论:一旦人类监督真的成为强系统的约束,监督机构本身就成了攻击目标。 任何围绕一个具备说服力的模型建立的监控或审批回路,都需要为这个回路本身写一份威胁模型。