Paper: 2608.21325 Authors: Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos, Anabela C. Areias, Maya D’Eon, Fabíola Costa, Ricardo Rei, Nuno M. Guerreiro Categories: cs.CL
The Gap
People are already using language models for emotional support, at scale, whether or not the models were designed for it. But almost nothing is known about how these models actually conduct a therapeutic interaction. The gap is not one of outcome measurement — whether users feel better is hard to attribute and slow to measure. It is a gap of mechanism: nobody had a vocabulary for describing what a model does, turn by turn, as a therapist.
That absence has a cost, and it is a specific one. Without a shared set of categories, you cannot compare a model’s behaviour to a human clinician’s, because “supportive response” is not a behaviour — it is an impression. You also cannot audit a model for a systematic stylistic bias, and you cannot steer it, because steering requires naming the thing you want to change.
[REALITY] users already use LLMs for emotional support
|
v
[WHAT WE MEASURE] helpfulness / safety / user satisfaction
|
v
[WHAT WE CANNOT SAY] what the model actually DOES
- "supportive response" is an impression, not a behaviour
- no shared categories -> no comparison to clinicians
- no names -> no auditing, no steering
|
v
[GAP] no vocabulary for the move-by-move conduct of
a model-led therapeutic interaction
The Increment
One sentence: Before this paper, there was no vocabulary for what a model does in a therapy conversation and therefore no way to compare it to a clinician or steer it; after it, a validated ten-move ontology turns therapeutic conduct into a measurable distribution, and handing that ontology to the model as tools halves its deviation from human behaviour.
Core Mechanism
The ontology is the contribution, and its construction is what earns it credibility. Ten therapeutic moves are defined as compact, function-based categories grounded in the MULTI-60 inventory — a therapist’s repertoire, not a generic taxonomy of sentiment. “Function-based” is the key word: moves are classified by what they are trying to accomplish in the interaction, which is what makes them applicable to model output without asking whether the model intended anything.
Validation is two-stage. Five licensed psychologists annotated interactions, establishing expert agreement on the categories; the ontology was then scaled using a judge-based approach that was verified to match that expert agreement. So the measurement instrument is anchored to clinicians before it is applied at scale — the ordering matters, because it means the later findings rest on a category scheme that practitioners of the discipline endorsed.
Applying it produces three findings about frontier models, and they are behavioural rather than evaluative.
Models over-use inquiry at up to three times the human rate. The default therapeutic reflex of these models is to ask another question. That is a recognisable conversational tic, and at three times the human rate it is a strong stylistic signature rather than a subtle tendency.
They neglect psychoeducation. Explaining, teaching, normalising — giving the person something to take away — is under-represented relative to clinician behaviour.
They are strongly context-anchored. This is the most nuanced finding. Models carry forward strategies a human clinician initiated, but rarely initiate those strategies themselves. The model is capable of the move; it just does not reach for it unprompted. That distinction between capacity and propensity is only visible because the ontology makes the moves nameable.
The intervention follows directly. Exposing the ontology to the model as a set of tools — making the ten moves concrete, callable options rather than latent tendencies — roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapists by 7 to 9 percentage points, with no fine-tuning at all.
ONTOLOGY CONSTRUCTION
MULTI-60 inventory
|
v
10 therapeutic moves (function-based, compact)
|
v
5 licensed psychologists annotate --> expert agreement
|
v
judge-based scaling verified to MATCH expert agreement
|
v
apply to real counseling transcripts + model-led sessions
|
v
FINDINGS
- inquiry over-used: up to 3x human rate
- psychoeducation neglected
- context-anchored: follows human-initiated strategies,
rarely initiates them
|
v
INTERVENTION: expose ontology as tools
-> mean deviation from human distribution roughly HALVED
-> turn-level alignment +7-9 points
- no fine-tuning
Think of it as a chess coach reviewing a transcript with a notation for tactics. Watching two games, you can say “one player was more aggressive” — an impression, unfalsifiable, unactionable. Give the coach a notation that names the actual tactical motifs, and three things become possible at once: you can count how often each motif appears in each game, you can compare a student to a grandmaster and say “you play pins three times as often and never play a fork”, and — the important one — you can put that list of motifs in front of the student as a checklist. The last step is what the paper’s intervention does, and the result is that the student starts playing the moves they were always capable of, simply because the moves now have names.
Key Concepts
- Function-based therapeutic moves: categories defined by what a move tries to accomplish in the interaction rather than by its surface form or the speaker’s intent. This is what makes the scheme applicable to model output, where intent is not observable, and it is what allows comparison across human and model transcripts.
- Context anchoring: the observed tendency to continue a strategy a human initiated while rarely initiating it oneself. It separates capability from propensity, and it is a failure mode that aggregate helpfulness scores cannot detect, because a model that follows well still looks helpful.
- Ontology-as-tools: the intervention of exposing the category scheme to the model as callable tools. It is a form of prompting, but a structurally different one from instruction-tuning: rather than teaching the behaviour, it makes an existing capability selectable, which is why it works without any weight updates.
Framework Shift
Before (impression-level evaluation):
conversation -> "helpful? safe? empathetic?"
output: a rating
- cannot compare to a clinician (no shared units)
- cannot see systematic bias
- cannot steer: nothing to name
After (move-level measurement):
conversation -> distribution over 10 therapeutic moves
output: a measurable profile
- compare model vs clinician distributions
- audit for specific over/under-use
- steer by exposing the moves as tools
(no fine-tuning, 7-9 pt alignment gain)
From asking whether a model’s support feels helpful, to counting which therapeutic moves it makes and comparing that distribution to a clinician’s, the core shift is replacing an impression with a repertoire you can name, count, and hand back to the model.
Expert Assessment
Problem choice: Excellent, and it takes the right half of the problem. Outcome evaluation of AI therapy is a swamp — user-reported wellbeing is confounded by everything. Choosing instead to measure conduct — what the model does, move by move — makes the question tractable, and it happens to be the question a supervising clinician would actually ask.
Method maturity: The two-stage validation is the paper’s methodological backbone and it is done properly: experts first, judge-model second, with the judge explicitly verified against the expert baseline. The intervention is elegant in its economy — exposing the ontology as tools is a prompting change, and getting a 7-9 point alignment gain without touching weights suggests the capability was already present and merely unprompted.
Experimental integrity: The findings are stated as distributions against a human reference, which is the right frame, and the comparison set includes real counseling transcripts rather than only synthetic sessions. The residual concern is the well-known one for judge-based measurement at scale: the judge was validated to match expert agreement on the annotated sample, but how that agreement holds across the full distribution — particularly on the under-used moves where data may be sparse — is not fully characterised.
Writing quality: The finding that models are context-anchored rather than incapable is the paper’s most interesting observation, and it is presented with the right emphasis. What would lift the paper is a concrete example: one short transcript excerpt annotated with its move sequence, human and model side by side, would make the three findings immediately legible to a clinician reader in a way the distribution statistics do not.
Verdict: strong accept — it builds a validated instrument for a domain that was being evaluated by impression, and then shows the instrument itself is the intervention.
Takeaways
- Measure conduct, not just outcomes, when evaluating any model that talks to people. Naming the moves is what makes comparison, auditing and steering possible at all.
- Distinguish capability from propensity. A model that rarely initiates a strategy may be perfectly capable of it — the fix is making the capability selectable, not training it in.
- Exposing a category scheme as tools can be a cheaper intervention than fine-tuning, and here it was equally effective. If a behaviour is absent but nameable, try naming it first.
- Watch for the inquiry reflex in any conversational agent. Asking another question is the path of least resistance, and at three times the human rate it becomes a style rather than a choice.
论文: 2608.21325 作者: Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos, Anabela C. Areias, Maya D’Eon, Fabíola Costa, Ricardo Rei, Nuno M. Guerreiro 分类: cs.CL
缺口
人们已经在把大语言模型当作情感支持来用,规模很大,而且不管这些模型当初是否为这个目的设计过。但关于这些模型究竟如何进行一场治疗性互动,我们几乎一无所知。 这个缺口不在于结果测量——用户是否感觉更好,既难以归因,也难以快速测量。缺的是机制:此前没有人拥有一套词汇,去逐轮描述一个模型作为”治疗者”到底做了什么。
这种缺席有代价,而且代价非常具体。 没有一套共享的类别,你无法把模型的行为与人类临床医师的行为做对比,因为”支持性回应”不是一种行为,而是一种印象。 你也无法审计模型是否存在系统性的风格偏好,更无法引导它——因为引导的前提,是先能把你想要改变的东西命名出来。
[现实] 用户已经在把大模型用于情感支持
|
v
[我们现在测的] 有用性 / 安全性 / 用户满意度
|
v
[我们说不出的是] 模型到底做了什么
- "支持性回应"是印象,不是行为
- 没有共享类别 -> 无法与临床医师对比
- 没有命名 -> 无法审计,无法引导
|
v
[缺口] 缺少一套词汇,用来描述模型主导的治疗互动中
逐轮发生的行为
增量
一句话: 在这篇论文之前,我们既没有描述模型在治疗对话中行为的词汇,也就无从与临床医师对比、更无从引导它;在这篇论文之后,一套经过验证的十类「动作」本体把治疗行为变成了可度量的分布,而把这套本体作为工具交给模型,能将其与人类行为的偏离减半。
核心机制
本体本身就是贡献,而它的构建方式才是可信度的来源。 十个治疗动作被定义为紧凑的、以功能为基础的类别,根基是 MULTI-60 量表——那是治疗师的技能清单,而不是某种泛化的情绪分类法。 “以功能为基础”是关键:动作按它在互动中试图达成什么来分类。正因如此,这套类别可以直接套用到模型输出上,而无需去追问模型是否”有意为之”。
验证分两步。 先由五位执业心理学家对互动进行标注,确立专家层面的一致性;随后用评判模型(judge)的方法扩展到大规模,并验证其与专家一致性相匹配。 也就是说,这把量尺在被大规模使用之前,先被锚定在临床医师身上——这个顺序很重要,因为它意味着后续的发现建立在学科从业者认可的类别体系之上。
应用之后,得到了关于前沿模型的三条发现,而它们都是行为层面的,而非评价层面的。
模型使用探询(inquiry)的比例最高可达人类的三倍。 这些模型的默认治疗反射,就是再问一个问题。这是一种可识别的对话口头禅;而当它达到人类三倍的频率时,它已经从”细微倾向”升级为强烈的风格签名。
它们普遍忽视心理教育(psychoeducation)。 解释、科普、正常化——即给当事人一些可以带走的东西——相比临床医师的行为显著偏少。
它们强烈依赖上下文锚定(context-anchored)。 这是最细腻的一条。模型会继承人类临床医师发起的策略,但极少自己发起这些策略。模型具备这个动作的能力,它只是不会主动去取用。 这种”能力”与”倾向”之间的区分,只有在本体把动作变得可命名之后才看得见。
干预手段随之直接给出。 把本体以一组工具的形式暴露给模型——让十个动作成为具体、可调用的选项,而不是潜藏着的倾向——就能把与人类动作分布的平均偏离大致减半,并把逐轮与人类治疗师的吻合度提升 7 到 9 个百分点,而且完全没有做任何微调。
本体的构建
MULTI-60 量表
|
v
10 类治疗动作(以功能为基础、紧凑)
|
v
五位执业心理学家标注 --> 专家一致性
|
v
评判模型扩展到大规模,并验证其与专家一致性相匹配
|
v
应用到真实咨询记录 + 模型主导的会话
|
v
发现
- 探询被过度使用:最高达人类三倍
- 心理教育被忽视
- 上下文锚定:会跟随人类发起的策略,
但极少自己发起
|
v
干预:把本体作为工具暴露出去
-> 与人类分布的偏离平均减半
-> 逐轮吻合度 +7~9 个点
- 无需微调
可以用**“给棋类教练配一套战术记号”**来理解这件事: 看两盘棋,你只能说”一方更爱进攻”——这是印象,既不可证伪,也无法据以行动。 而一旦给教练一套能点名具体战术母题的记号,三件事会同时变得可能:你可以数出每种母题在各局里出现了几次;你可以把学生和大师对比,说出”你的牵制(pin)下得比人家多三倍,而且从来不下叉攻(fork)“;还有最关键的那一步——你可以把这份母题清单摆到学生面前当检查表用。 最后这一步,正是本文干预所做的事;结果是:学生开始下出他本来就一直会的招,仅仅因为这些招终于有了名字。
关键概念
- 以功能为基础的治疗动作: 按动作在互动中试图达成什么来分类,而不是按表层形式或说话人的意图来分类。这使该体系能够套用到模型输出上(因为意图不可观测),也正是它能够在人类与模型文本之间做对比的原因。
- 上下文锚定(context anchoring): 观察到的一种倾向——会延续人类发起的策略,却极少自己发起。它把”能力”与”倾向”区分开来,而且是一种聚合的”有用性”分数根本无法察觉的失效模式,因为一个只会很好地跟随的模型,看起来依然很有帮助。
- 本体即工具(ontology-as-tools): 把类别体系作为可调用工具暴露给模型的干预方式。它是提示的一种,但在结构上不同于指令微调:它不是去教这个行为,而是让一个已有的能力变得可被选中,这也是它无需更新权重就能生效的原因。
框架转变
之前(印象层面的评估):
对话 -> "有帮助吗?安全吗?有共情吗?"
输出:一个评分
- 无法与临床医师对比(没有共同单位)
- 看不出系统性偏好
- 无法引导:没有东西可命名
之后(动作层面的测量):
对话 -> 10 类治疗动作上的分布
输出:一份可度量的画像
- 对比模型与临床医师的分布
- 审计具体的过度使用 / 使用不足
- 通过把动作作为工具暴露来引导
(无需微调,吻合度 +7~9 个点)
从”问一句模型的支不支持算不算有帮助”,转变为”数清它做了哪些治疗动作、并把这份分布与临床医师对比”,核心转变在于:把一种印象,换成一套你能命名、能清点、还能交还给模型的技能库。
专家评审
选题眼光: 极好,而且选对了问题的另一半。 对 AI 心理治疗做结果评估是一片沼泽——用户自报的身心状态被太多因素混淆。转而测量行为本身——模型逐轮做了什么——把这个题目变得可处理,而且它恰好是一位督导临床医师真正会问的问题。
方法成熟度: 两段式验证是全文的方法学脊梁,而且做得规范:先专家、后评判模型,并明确验证评判模型与专家基线的一致性。干预手段的”省”很漂亮——把本体作为工具暴露,本质是一次提示层面的改动;在不碰权重的前提下拿到 7 到 9 个点的吻合度提升,说明该能力本就存在,只是没有被激发。
实验诚意: 结论以”相对人类参考的分布”形式给出,这是正确的框架;对比集合里包含真实咨询记录,而不是只有合成会话。残留的疑虑是”大规模评判模型测量”的老问题:评判模型在被标注的样本上被验证为与专家一致,但这份一致性在完整分布上——尤其在使用不足、数据可能稀疏的那几个动作上——能保持到什么程度,并未被完整刻画。
写作功力: “模型是上下文锚定、而非没有能力”是全文最有意思的观察,而论文给了它恰当的强调。真正能让论文更上一层楼的是一段具体示例:一段带动作序列标注的短对话片段,人类与模型并排展示。对临床读者来说,那会比分布统计立刻可读得多。
判决: 强接收(Strong Accept) — 它为一个人人都在用印象评估的领域造出了一把经过验证的量尺,并进一步证明:这把量尺本身就是干预手段。
要点总结
- 评估任何会与人对话的模型时,请测量行为,而不只是结果。把动作命名出来,是对比、审计与引导得以成立的前提。
- 区分「能力」与「倾向」。一个极少主动发起某策略的模型,很可能本来就具备这个能力——解法是让这个能力可被选中,而不是把它训进去。
- 把类别体系作为工具暴露出去,可能比微调更省,而在这项工作里两者效果相当。如果某个行为缺席但它可被命名,先试试命名它。
- 警惕任何对话智能体的「追问反射」。再问一个问题是最省力的路径,而当它达到人类三倍的频率时,它就从”选择”变成了”风格”。