Paper: 2609.14065 Authors: Gabriele Farina, Juan Carlos Perdomo Categories: cs.LG, cs.GT
The Gap
When algorithmic predictions inform people’s decisions, the models we deploy are performative and actively shape the data we see. A model predicting which applicants will default changes who is approved, which changes who defaults — so the distribution the model is trained on is partly a consequence of the model.
That creates a question about the target itself: if different predictive models induce different distributions, is it possible to efficiently learn a prediction rule that is optimal for the distribution that it induces? The solution concept is performative stability — a predictor that is optimal for its own induced distribution rather than for a distribution that existed before it.
And the learning problem is unlike supervised learning in a structural way: the learner must deploy different predictors and observe their induced distributions. In supervised learning the distribution is fixed and data can be collected without consequence. Here every observation requires deploying a model into the world, so the number of deployments is the scarce resource — and it is the resource any algorithm must minimise.
PREDICTIONS THAT SHAPE THE DATA THEY PREDICT
when algorithmic predictions INFORM PEOPLE'S DECISIONS,
THE MODELS WE DEPLOY ARE PERFORMATIVE AND ACTIVELY SHAPE THE DATA
WE SEE
-> a model predicting which applicants will default changes WHO IS
APPROVED, which changes WHO DEFAULTS
-> the distribution the model is trained on is PARTLY A CONSEQUENCE
OF THE MODEL
|
v
[A QUESTION ABOUT THE TARGET ITSELF]
IF DIFFERENT PREDICTIVE MODELS INDUCE DIFFERENT DISTRIBUTIONS,
is it possible to EFFICIENTLY LEARN a prediction rule that is
OPTIMAL FOR THE DISTRIBUTION THAT IT INDUCES?
-> the solution concept is PERFORMATIVE STABILITY:
a predictor OPTIMAL FOR ITS OWN INDUCED DISTRIBUTION rather
than for a distribution that EXISTED BEFORE IT
[AND THE LEARNING PROBLEM IS STRUCTURALLY UNLIKE SUPERVISED LEARNING]
THE LEARNER MUST DEPLOY DIFFERENT PREDICTORS AND OBSERVE THEIR
INDUCED DISTRIBUTIONS
<- in supervised learning the distribution is FIXED and data can
be collected WITHOUT CONSEQUENCE
<- here EVERY OBSERVATION REQUIRES DEPLOYING A MODEL INTO THE
WORLD
-> so the NUMBER OF DEPLOYMENTS is THE SCARCE RESOURCE
-> and it is the resource ANY ALGORITHM MUST MINIMISE
The Increment
One sentence: Before this paper, finding a performatively stable predictor required many deployments or assumptions about how predictions shape distributions; after it, a procedure reaches stability in nearly the minimum number of deployments with no such assumptions.
Core Mechanism
The main result is an algorithmic procedure that, in the high-accuracy regime, finds a performatively stable model in nearly the minimum number of model deployments without making any assumptions regarding how predictions shape distributions.
Three parts of that sentence are each load-bearing, and the third is the strongest.
“Nearly the minimum number of model deployments” is the right optimality target, because deployments are the scarce resource. Not sample complexity in the abstract — deployments specifically — which is what makes the result about the actual cost of learning in this setting.
“Without making any assumptions regarding how predictions shape distributions” is what separates this from earlier work. Prior approaches achieved stability under conditions on the performative effect — for instance that it is small or well-behaved. Assuming the feedback loop is mild is assuming away part of the problem, since the interesting cases are where it is not.
And the improvement is exponential: the procedure succeeds at finding a randomized performatively stable predictor using exponentially fewer model deployments than prior approaches. The qualifier randomized is substantive rather than incidental — it is the relaxation that makes the efficiency possible, and it sets up the second contribution.
That second contribution is a derandomization result, and its structure is the frankest part of the paper: it shows how this randomized notion of stability can be derandomized into a single predictor satisfying the prior deterministic notion if one is willing to assume that the loss is well-conditioned and that performative effects are weak, as in early work in this area. Read as a trade-off statement, this is precise about what was gained and what it cost. The efficient procedure is randomized; converting it to a single deterministic predictor requires bringing back the assumptions that the main result avoided. The paper does not claim the derandomization is free — it states the price, which is the honest form for a result of this shape.
And the technical source is named: the results come from building on an underexplored technical connection between performative stability and expected variational inequalities. Naming the connection is useful in itself, since it suggests the machinery may transfer to other performative problems — the connection was available and unexploited.
THE MAIN RESULT
an ALGORITHMIC PROCEDURE that, IN THE HIGH-ACCURACY REGIME,
finds a PERFORMATIVELY STABLE MODEL
IN NEARLY THE MINIMUM NUMBER OF MODEL DEPLOYMENTS
WITHOUT MAKING ANY ASSUMPTIONS REGARDING HOW PREDICTIONS SHAPE
DISTRIBUTIONS
THREE PARTS, EACH LOAD-BEARING, THE THIRD STRONGEST
[1] "NEARLY THE MINIMUM NUMBER OF MODEL DEPLOYMENTS"
<- the RIGHT OPTIMALITY TARGET, because DEPLOYMENTS ARE THE
SCARCE RESOURCE
<- NOT sample complexity in the abstract -- DEPLOYMENTS
SPECIFICALLY
-> which makes the result about THE ACTUAL COST OF LEARNING
IN THIS SETTING
[2] "WITHOUT MAKING ANY ASSUMPTIONS REGARDING HOW PREDICTIONS SHAPE
DISTRIBUTIONS"
<- this SEPARATES IT FROM EARLIER WORK: prior approaches
achieved stability UNDER CONDITIONS ON THE PERFORMATIVE
EFFECT (that it is SMALL or WELL-BEHAVED)
<- assuming the FEEDBACK LOOP IS MILD is ASSUMING AWAY PART
OF THE PROBLEM
-> since THE INTERESTING CASES ARE WHERE IT IS NOT
[3] AND THE IMPROVEMENT IS EXPONENTIAL
the procedure finds a RANDOMIZED performatively stable
predictor using EXPONENTIALLY FEWER model deployments than
prior approaches
<- the qualifier "RANDOMIZED" is SUBSTANTIVE rather than
incidental: it is the RELAXATION that makes the EFFICIENCY
POSSIBLE
-> and it SETS UP THE SECOND CONTRIBUTION
[THE SECOND CONTRIBUTION: DERANDOMIZATION, AND ITS PRICE IS STATED]
it shows how this RANDOMIZED notion of stability can be
DERANDOMIZED INTO A SINGLE PREDICTOR satisfying the PRIOR
DETERMINISTIC NOTION
IF one is willing to assume that
THE LOSS IS WELL-CONDITIONED
PERFORMATIVE EFFECTS ARE WEAK, as in EARLY WORK IN THIS AREA
<- read as a TRADE-OFF STATEMENT, this is PRECISE about WHAT WAS
GAINED AND WHAT IT COST
<- the EFFICIENT procedure is RANDOMIZED; converting it to a
SINGLE DETERMINISTIC predictor REQUIRES BRINGING BACK THE
ASSUMPTIONS THE MAIN RESULT AVOIDED
<- the paper does NOT claim the derandomization is FREE -- it
STATES THE PRICE, the honest form for a result of this shape
[AND THE TECHNICAL SOURCE IS NAMED]
the results come from building on an UNDEREXPLORED TECHNICAL
CONNECTION BETWEEN PERFORMATIVE STABILITY AND EXPECTED VARIATIONAL
INEQUALITIES
<- naming the connection is useful IN ITSELF: it suggests the
machinery MAY TRANSFER to OTHER PERFORMATIVE PROBLEMS
<- the connection was AVAILABLE and UNEXPLOITED
Think of it as learning where to set a subsidy when the setting changes behaviour. You want a level that is right for the economy your level creates — and every trial means actually setting the subsidy and living with the consequences, which is why the number of trials is the cost that matters. Two properties of the procedure are what you would want. It does not require you to assume in advance that people barely respond to the subsidy, because assuming a mild response is assuming away the case you are worried about. And when you eventually want one fixed number rather than a randomised policy, the conversion is available — but only if you accept assumptions that the efficient version did not need, which the paper states rather than glosses.
Key Concepts
- Performativity and induced distributions: the deployed model shapes the data, so the target is optimality for its own induced distribution. It is what makes the learning problem structurally different from supervised learning.
- Deployments as the scarce resource: every observation requires putting a model into the world. It is why “nearly the minimum number of deployments” is the right optimality target.
- Assumption-free stability: no conditions on how predictions shape distributions. Assuming a mild performative effect would exclude the cases the problem is about.
- Randomization as the enabling relaxation: the exponential improvement is for a randomized notion of stability, which is what makes the efficiency attainable.
- Derandomization with a stated price: recovering a single deterministic predictor requires well-conditioned loss and weak performative effects. Naming the cost rather than presenting derandomization as free is what keeps the trade-off legible.
- The connection to expected variational inequalities: the technical route, previously unexploited, which may transfer to other performative problems.
Framework Shift
Before (stability under assumptions, or many deployments):
learn a predictor optimal for its induced distribution
-> prior approaches assumed performative effects weak or small
-> or required many model deployments
-> the interesting regime, strong feedback, was assumed away
After (assumption-free, nearly minimal deployments):
a procedure finds a randomized performatively stable predictor
-> exponentially fewer deployments than prior approaches
-> no assumptions about how predictions shape distributions
-> derandomizing to a single predictor costs well-conditionedness
and weak performative effects
From assuming the feedback loop is mild so that stability becomes tractable, to reaching stability without such assumptions in nearly the minimum number of deployments, the core shift is that the efficiency of learning can come from relaxing the solution concept rather than from restricting the problem.
Expert Assessment
Problem choice: Excellent, and the framing of the cost is the contribution: recognising that deployments, not samples, are the scarce resource turns the question into one with a well-defined optimality target. And refusing to assume mild performative effects keeps the problem in the regime that motivates it.
Method maturity: The main result is strong on the three axes that matter — the optimality target is the right one, the assumption set is empty, and the improvement is exponential. The derandomization contribution is well handled: presenting it as requiring the assumptions the main result avoided, with the specific conditions named, is the honest form for a trade-off. Naming the technical connection to expected variational inequalities is a contribution in itself, since it points the machinery at other performative problems.
Experimental integrity: This is a theory contribution, so the reading is of the high-accuracy regime and the solution concept. The regime qualifier is real — the result holds asymptotically, and behaviour in the low-accuracy regime is not addressed — and the randomized solution concept is a genuine relaxation, meaning a practitioner wanting a single deployed model is in the derandomization case with its assumptions. The paper states both boundaries, which keeps the result’s scope clear rather than overclaiming applicability.
Writing quality: The abstract separates the assumption-free randomized result from the conditional derandomization, which is exactly the structure a reader needs to see what applies to their situation. Because the practical question is how many deployments “nearly the minimum” means, a worked instance with the deployment count compared against a prior method would make the exponential claim concrete.
Verdict: strong accept — it reaches performative stability without assumptions about the feedback loop in nearly the minimum number of deployments, states the price of derandomization rather than hiding it, and opens a technical connection likely to transfer.
Takeaways
- Count the right resource. When every observation requires deploying a model into the world, deployments rather than samples are the cost to minimise.
- Check whether an assumption removes the interesting case. Assuming performative effects are mild excludes exactly the feedback loops that motivated the problem.
- Ask what a relaxation costs. An efficient result for a randomized solution concept may require assumptions to convert into a single deployable model.
- Note an unexploited technical connection. Identifying one is a contribution, because the machinery may transfer to other problems in the same class.
论文: 2609.14065 作者: Gabriele Farina, Juan Carlos Perdomo 分类: cs.LG, cs.GT
缺口
当算法预测为人们的决策提供依据时,我们所部署的模型是「表演性的」,并且在主动塑造我们所见到的数据。 一个预测”哪些申请者会违约”的模型,会改变”谁被批准”,而后者又改变”谁会违约”——所以模型训练所依据的那个分布,部分是模型自身的后果。
这就产生了一个关于目标本身的问题:如果不同的预测模型会诱发不同的分布,那么是否可能「高效地」学出一个预测规则,使它对「它自己所诱发的那个分布」最优? 这个解概念就是表演性稳定(performative stability)——一个对它自身诱发分布最优的预测器,而不是对”在它之前就存在的某个分布”最优。
而这个学习问题在结构上与监督学习不同:学习者必须部署不同的预测器、并观察它们所诱发的分布。 在监督学习里分布是固定的,数据可以无后果地采集。而在这里,每一次观察都需要把一个模型部署进世界——于是**「部署次数」就是那个稀缺资源**,而它是任何算法都必须最小化的资源。
会塑造它所预测之数据的预测
当算法预测为人们的决策提供依据时,
「我们所部署的模型是表演性的,并在主动塑造我们所见到的数据」
-> 一个预测"哪些申请者会违约"的模型,会改变"谁被批准",
而后者又改变"谁会违约"
-> 模型训练所依据的分布,「部分是模型自身的后果」
|
v
[一个关于「目标本身」的问题]
「如果不同的预测模型会诱发不同的分布,
那么是否可能「高效地」学出一个预测规则,
使它对「它自己所诱发的那个分布」最优?」
-> 解概念是「表演性稳定」:
一个对它「自身诱发分布」最优的预测器,
而不是对"在它之前就存在的某个分布"最优
[这个学习问题在结构上与监督学习不同]
「学习者必须部署不同的预测器、并观察它们所诱发的分布」
<- 监督学习里分布是「固定的」,数据可以「无后果地」采集
<- 而这里「每一次观察都需要把一个模型部署进世界」
-> 于是「部署次数就是那个稀缺资源」
-> 而它是「任何算法都必须最小化」的资源
增量
一句话: 在这篇论文之前,找到表演性稳定的预测器需要大量部署、或需要对”预测如何塑造分布”作假设;在这篇论文之后,一套流程在近乎最少的部署次数内达到稳定,且不带这类假设。
核心机制
主结果是一套算法流程:在高精度区间内,它在「近乎最少的模型部署次数」内找到表演性稳定的模型,且「不对预测如何塑造分布作任何假设」。
这句话有三部分是承重的,而第三部分最强。
“近乎最少的模型部署次数”是正确的优化目标,因为部署次数才是稀缺资源。不是抽象意义上的样本复杂度,而是具体到部署次数——这才让这个结果关乎该场景下学习的真实成本。
“不对预测如何塑造分布作任何假设”才是把这项工作与此前工作分开的地方。既有做法是在对表演性效应的条件下取得稳定性的——比如”它很小”或”它表现良好”。而假设反馈回路是温和的,等于把问题的一部分假设掉,因为真正有意思的恰恰是它不温和的情形。
而改进是指数级的:该流程用相对既有方法指数级更少的模型部署次数,找到了一个随机化的表演性稳定预测器。 “随机化”这个限定是实质性的、不是顺带的——它正是让这份效率得以可能的那个松弛,而它也引出第二项贡献。
第二项贡献是一个去随机化结果,而它的结构是全文最坦率的部分:它表明这个随机化的稳定概念可以被去随机化为一个满足此前「确定性」概念的单一预测器——前提是愿意假设「损失是良态的」、且「表演性效应是弱的」,正如这一领域早期工作所做的那样。 当作一个取舍陈述来读,它对”得到了什么、付出了什么”讲得很精确。高效流程是随机化的;把它换成单一的确定性预测器,需要把主结果所回避的那些假设请回来。论文没有声称去随机化是免费的——它说出了代价,对一个这种形态的结果而言,这是诚实的写法。
而技术来源也被点名:这些结果来自**“建立在「表演性稳定」与「期望变分不等式」之间一处未被充分探索的技术联系之上”。** 点名这处联系本身就有用,因为它提示这套机械可能迁移到其他表演性问题上——那处联系一直是可用的,却未被利用。
主结果
一套算法流程:在「高精度区间」内,
在「近乎最少的模型部署次数」内找到表演性稳定的模型,
且「不对预测如何塑造分布作任何假设」
三部分都承重,第三部分最强
[1] "近乎最少的模型部署次数"
<- 正确的优化目标,因为「部署次数才是稀缺资源」
<- 不是抽象意义上的样本复杂度,而是「具体到部署次数」
-> 这才让这个结果关乎「该场景下学习的真实成本」
[2] "不对预测如何塑造分布作任何假设"
<- 这才把它与此前工作「分开」:既有做法是在
「对表演性效应的条件下」取得稳定性的
(比如"它很小"或"它表现良好")
<- 「假设反馈回路是温和的,等于把问题的一部分假设掉」
-> 因为「真正有意思的恰恰是它不温和的情形」
[3] 而改进是「指数级」的
该流程用相对既有方法「指数级更少」的模型部署次数,
找到了一个「随机化的」表演性稳定预测器
<- "随机化"这个限定是「实质性的」、不是顺带的:
它正是让这份效率「得以可能」的那个松弛
-> 而它也「引出第二项贡献」
[第二项贡献:去随机化,而它的「代价被说出」]
它表明这个「随机化」的稳定概念可以被「去随机化为」
一个满足此前「确定性」概念的「单一预测器」
「前提是」愿意假设
「损失是良态的」
「表演性效应是弱的」,正如这一领域早期工作所做的那样
<- 当作一个「取舍陈述」来读,它对"得到了什么、付出了什么"
讲得很「精确」
<- 高效流程是「随机化」的;把它换成「单一的确定性」预测器,
需要把主结果所回避的那些假设「请回来」
<- 论文没有声称去随机化是「免费」的——它「说出了代价」,
这对一个这种形态的结果而言是「诚实」的写法
[而技术来源也被点名]
这些结果来自"建立在「表演性稳定」与「期望变分不等式」之间
一处「未被充分探索的技术联系」之上"
<- 点名这处联系本身就有用:它提示这套机械
「可能迁移到其他表演性问题」上
<- 那处联系「一直是可用的,却未被利用」
可以用**“在「补贴额度本身会改变人们行为」的情况下,学出该把补贴定在哪里”来理解这件事: 你想要一个对你所创造出的那个经济体而言正确的额度——而每一次试,都意味着真的把补贴定下去、并承担后果**,这就是为什么试的次数才是要紧的成本。 这套流程有两条性质,正是你会想要的。它不需要你事先假设”人们对补贴几乎不反应”——因为假设反应温和,等于把你真正担心的那种情形假设掉。而当你最终想要一个固定的数值、而不是一个随机化政策时,转换是可以的——但只有在你接受”高效版本并不需要”的那些假设时才可以;而论文把这一点说出来,而不是含糊过去。
关键概念
- 表演性与诱发分布: 被部署的模型塑造数据,因此目标是对它自身诱发的分布最优。正是它让这个学习问题在结构上不同于监督学习。
- 以部署次数作为稀缺资源: 每一次观察都需要把一个模型放进世界。这正是”近乎最少的部署次数”是正确的优化目标的原因。
- 无假设的稳定性: 不对”预测如何塑造分布”施加条件。假设表演性效应温和,会把问题真正关于的那些情形排除掉。
- 以随机化作为使能的松弛: 指数级改进是针对”随机化的”稳定概念而言的,正是它让这份效率可达。
- 带代价陈述的去随机化: 要回到单一确定性预测器,需要良态损失与弱表演性效应。说出代价、而不是把去随机化说成免费,才让这个取舍保持可读。
- 与期望变分不等式的联系: 技术路线,此前未被利用,可能迁移到其他表演性问题。
框架转变
之前(在假设下求稳定,或需要大量部署):
学一个对它诱发分布最优的预测器
-> 既有做法假设表演性效应弱或小
-> 或者需要大量模型部署
-> 那个有意思的区间——强反馈——被假设掉了
之后(无假设、近乎最少部署):
一套流程找到随机化的表演性稳定预测器
-> 相对既有方法指数级更少部署
-> 不对"预测如何塑造分布"作假设
-> 去随机化为单一预测器的代价是"良态性 + 弱表演性效应"
从”假设反馈回路温和、以便让稳定性变得可处理”,转变为”在不带这类假设的情况下、以近乎最少的部署次数达到稳定”,核心转变在于:学习的效率可以来自「松弛解概念」,而不是来自「限制问题」。
专家评审
选题眼光: 极好,而对成本的框定就是贡献:认识到稀缺资源是部署次数而不是样本数,使这个问题有了定义明确的优化目标。而拒绝假设”表演性效应温和”,则让问题留在促使它被提出的那个区间里。
方法成熟度: 主结果在三条要紧的轴上都强——优化目标是对的、假设集是空的、改进是指数级的。 去随机化这项贡献处理得好:把它呈现为”需要主结果所回避的那些假设”、并点名具体条件,是对一个取舍而言诚实的形态。 点名与期望变分不等式的技术联系本身就是贡献,因为它把这套机械指向了其他表演性问题。
实验诚意: 这是一项理论贡献,因此该读的是高精度区间与解概念。区间限定是真实的——结果在渐近意义下成立,而低精度区间的行为未被涉及——随机化的解概念是一次真实的松弛,意味着想要部署单一模型的实践者处在”去随机化”那一侧、并继承其假设。论文把这两条边界都说了,这让结果的范围保持清楚,而不是过度声称适用性。
写作功力: 摘要把”无假设的随机化结果”与”有条件的去随机化”分开陈述——这正是读者需要看到的结构,以便判断哪一条适用于自己的处境。 由于实际问题是”近乎最少”究竟意味着多少次部署,若能给一个具体实例、把部署次数与既有方法作对比,会让”指数级”这个主张变得具体。
判决: 强接收(Strong Accept) — 它在近乎最少的部署次数内、不带关于反馈回路的假设达到了表演性稳定,说出了去随机化的代价而不是把它藏起来,并打开了一处很可能迁移的技术联系。
要点总结
- 数对资源。当每一次观察都需要把一个模型部署进世界时,要最小化的是部署次数,而不是样本数。
- 检查一条假设是否把有意思的情形排除掉了。假设表演性效应温和,恰恰排除了促使这个问题被提出的那种反馈回路。
- 问一句一次松弛的代价是什么。一个针对”随机化解概念”的高效结果,可能需要额外假设才能转换成一个可部署的单一模型。
- 留意未被利用的技术联系。识别出它本身就是贡献,因为这套机械可能迁移到同一类问题中的其他问题上。