Common Ground: Shared Understanding Before Autonomous Work

common-ground is a skill that aligns the model and me on what a task is for before the model does the work. The model explains the task in its own words, adds its own judgment, and tests that understanding against the intended use. At the end, the model writes a brief that the work can start from. Until 2026-10-04, the skill had the name deposition.
This essay explains the parts of the skill and the reason for each part. Each part comes from a problem that I saw in real sessions.
The problem: a correct result for a different purpose
The skill exists because a model can complete my request correctly and still miss my purpose. The model does the work, produces a full result, and passes its checks. Then I read the result and see that the model organized the work around a different goal. Sometimes I did not say my intention. Sometimes I did not know it yet.
This essay has its own example. I asked an agent for “another blog” about the updated skill, and I told it to publish without questions. The agent searched the web, wrote the essay, made the image, and published. The essay was good, but about 40% of it repeated my earlier essay on deposition. Nothing asked what the second essay was for. This merged essay replaces both.
Recent models make the problem larger. Anthropic's guide to Opus 5.5 tells users to give the whole task in one message and to name the finish line. Then the user lets the model work. Early testers ran coding tasks for hours with little oversight. OpenAI's skills guidance for GPT-6 Astra also asks authors to state decision boundaries and the condition for completion.
When I worked with the model turn by turn, a misunderstanding cost one turn. I read the reply, saw the drift, and corrected it. In a long autonomous run (in this essay, a long run), I am not there. The model makes hundreds of small choices under its reading of the goal. I see all of the choices at the end. So the cost of a misunderstanding increases with the length of the run, and the chance to correct it moves to the start.
Why the name is “common ground”
The name comes from psycholinguistics. Herbert Clark and Susan Brennan's paper “Grounding in communication” (1991) describes conversation as a joint activity. The speaker and the listener both work to confirm that the listener understood. They continue until both believe that the understanding is sufficient for the current purpose. Clark and Brennan call that threshold the grounding criterion. The mutual knowledge that the two people build is their common ground.
Three parts of the paper apply to the skill.
- Grounding goes in two directions. The old name,
deposition, described one person who asks questions and one person who answers for the record. I never wanted that procedure. The model gives its interpretation and its judgment. I accept it, change it, or reject it. Each of us needs evidence that the other understood. - The criterion depends on the purpose. If I will read the next reply anyway, a rough shared picture is sufficient, because I can correct it in the next turn. If the next step is a long run, the model must act alone for hours on the understanding. So the criterion is higher before a long run.
- The medium controls how grounding can occur. Face-to-face conversation lets people see and hear each other at the same time. A listener can frown in the middle of a sentence. A long run removes these signals. Written text keeps two properties: a reader can read it again, and a writer can correct it before sending. So the understanding for a long run must be a written text that the run can read again.
The model explains the task in its own words
The first step is a restatement. Lauren (@poteto) recommends this practice in the second part of her pstack guide. The part has the title “The art of supervising someone smarter than you.” She asks the agent to restate the problem in its own words. The restatement shows a misunderstanding before the agent writes code. Also, the agent did not copy her assumptions, so it can find a view that she did not consider.
A restatement must take a position, or it gives me nothing to check. Suppose I ask for a first presentation on a collaborator's spatial transcriptomics data: basic quality control (QC) and a slide deck, with no deep biological analysis. A weak reply repeats the request: “You want basic QC and a slide deck, without deep biological analysis.” The reply is accurate, but I cannot learn anything from it.
A useful reply states a judgment: “This meeting must show the collaborators that we received and understood their data. The QC metrics are evidence for that. They are not the center of the talk.” The judgment can be wrong. But I can react to it without a slide plan of my own. I can say if it is close to my meaning.
The skill calls the result a shared mental model. The model connects the intended use to the constraints, the choices, the consequences, and the evidence of success. Then it shows how those connections control its judgment.
An earlier version of the skill told the model to explain four fields: goals, acceptance, verification, and boundaries. The list became a form. The model filled in each field and stopped. A real task can depend on something that no field names. So the skill now asks for the connected picture, and the task decides which parts need attention.
The model adds its own judgment
The model must also give judgment that I did not ask for. A model that only makes my stated requirements precise cannot help with requirements that I did not state. The skill tells the model to state assumptions and tacit judgments that affect the task. The model also brings knowledge, alternatives, or a better framing when they change what is worth doing. For each idea, the model says what the idea changes and what evidence supports it.
Tacit knowledge needs a concrete prompt. I see two cases:
- I have a judgment that I cannot say yet. For example, I know that the slides are wrong, but I cannot say why. A candidate statement helps, because I can recognize or reject it.
- I already have the knowledge, but I did not connect it to the task. The model can propose the connection.
In both cases, a contrasting example or a short sketch is easier to answer than an abstract question. The model does not claim to know what I want. My answer decides. When a blind spot needs a longer exploration, the companion skill known-unknowns does that work, and the result goes back into the shared mental model.
The model asks only for my judgments
The model must ask me only for judgments that belong to me and that can change the direction. Every question has a cost. Clark's related principle of least collaborative effort says that partners keep the total work of understanding small. A question about a reversible choice moves work from the model to me, and the question adds little understanding.
The skill started as a fork of a questioning skill in Matt Pocock's skill collection. The early version of his grill-me skill asks one question at a time and gives a recommended answer for each question. I often accepted one recommendation, then the next. I confirmed each choice, but my purpose did not become clearer. If the framing of the questions depends on a goal that nobody said, the answers carry that goal forward.
In later sessions, I saw the opposite fault. The model asked about points that the record already settled, or about choices that I can undo in a minute. Now the skill sets these rules:
- Before the model asks, it checks if earlier work or the record already answers the question.
- The model finds discoverable facts itself.
- The model decides reversible, low-risk points itself and lists them as defaults that I can change.
- The model asks me about purpose, priorities, and acceptable costs.
Before a long run, the model writes a brief
Before a long run, the model must close the discussion with a brief. The brief is one message that the run can start from. It contains these items:
- Goal: what the result is for.
- Finish line: the behavior that counts as done, and the check that shows it.
- Settled decisions: what we agreed, so that the run does not open the decision again.
- Defaults: the reversible choices that the model made, so that I can change them.
- Bounded uncertainties: the open questions, and how much each one can affect the result.
The model then confirms the brief with me and starts the run from it. My standing instruction makes one exception: if the conversation already settled the brief, the model does not ask again. A second confirmation of a settled brief adds a turn and no understanding.
The finish line prevents a known failure. Anthropic's report on harnesses for long-running agents describes a later session that sees progress and declares the job done. A finish line written as a behavior and a check gives the run a test. Without the test, the run has only a feeling of progress.
Anthropic uses the same idea between two agents. In their harness for long-running application development, the generator and the evaluator agree on a sprint contract before they write code. The contract says what the generator will build and how the evaluator will verify it. Common Ground makes the same agreement between the model and me, while I am still present.
Two rules for the brief come from runs that went wrong:
- Mark each number as an estimate. If a plan says “about twenty tests” as a rough scale, the executor treats the number as a requirement. So each count or duration in the brief has the word “estimate.”
- Size the scope from real use. The model asks what evidence of real use supports each limit and mechanism. In my sessions, the largest misalignments came from scope that matched worst cases or the feature list of a reference product.
A passed check does not prove fitness for purpose
A finish line is a check, and a check supports only a specific claim. The tests pass. The page builds. These facts do not show that the result is useful to the person who reads it.
The preview image for this essay is an example. The first image met every written requirement: a 5:2 panorama, pure white paper, ink hatching, no text. But one figure held the staff at shoulder height, and the staff looked like a rifle. No check could find that fault. I saw it when I looked at the image as a reader. The second image fixed the pose.
So before the model closes the discussion, it walks through one plausible use of the result. It starts with the part that we understand least. It asks: can the work meet every stated requirement and still fail the person who uses it?
The walk-through must reach the content. In the presentation example, an agreement on the sequence “QC, then figures, then slides” does not decide what the slides say. When the model walks through what the collaborators see on each slide and what they conclude, the walk-through shows if we agree on the deliverable. For an essay, a short synopsis and outline with the audience and the argument marked as provisional do the same work. I see what the model will write before the full draft depends on it.
The finish line goes into the brief after the walk-through. For exploratory work, we agree on the question and a time to review the findings. We do not need the answer in advance.
Agreement and authorization are different
The model must keep a judgment about the task apart from permission to act. The model can propose a new framing without permission. To change files, increase the scope, or make a new deliverable, the model needs authorization.
When I invoke the skill explicitly, I ask for discussion before the dependent work. After I authorize the work, the model completes it without more requests for permission. An open question stops only the actions that depend on it. For example, if the wording of a button is settled and the page title is open, the model changes the button and keeps the title question.
The word “yes” caused friction. If the model asks whether its interpretation is correct and I say yes, I confirm an understanding. If the model proposes a specific change in an authorized task and I say yes, the model should do the change. Models often treated every yes as permission to start, or they asked again for permission that I already gave.
Two rules decrease this friction:
- The model says whether it checks an interpretation or proposes an action. It reads my reply against that and against the authorization that already applies.
- I can use two optional signals.
agreeconfirms the understanding and starts no new work.proceedauthorizes the settled next action. Ordinary words, such as “yes, make those changes,” also work.proceeddoes not settle a scope that is still unclear.
During the run, the shared mental model controls the choices. If new evidence changes an important premise, the model opens the discussion again. If I clarify the goal during the work, the model examines the earlier choices again. A new constraint added to the old plan is not sufficient.
I revise the skill from evidence
I change the skill when a session shows a specific cause. For each failure, I find out which of three causes applies. The guidance had a gap or was unclear. The model ignored an instruction. Or the necessary context did not reach the model. If I add a rule after every bad result, the skill becomes a heavy process.
The sessions showed causes that I did not expect:
- In one tool, the skill text reached the model cut at 8,000 characters. The rules about authorization never arrived. A correctly installed skill is not always a completely read skill.
- In another session, the model had the full text and still asked for permission twice. It read “pause to resolve a disagreement” as “start the approval again.”
- In the presentation work, the model kept its earlier figure choices after the purpose became clearer. It improved the titles but did not examine the figures again. The cause was not missing instructions. The model did not continue to use the understanding that we built.
In September 2026, I cut the skill to about one third of its earlier length. I removed generic method tutorials, repeated completeness checklists, and a fixed two-round approval script. A review by a second model agreed with most cuts. The review also found cuts that removed real constraints, such as a deliverable without a statement of its content. I put those constraints back.
The October changes (the new name, the brief, and the trigger before a long run) are three days old. I have the design reasons and the sessions that caused the changes. I do not have a measured result yet.
How to use the skill
Common Ground is in my skill collection. The skill instructions and the source notes are public.
For a new idea, I start with this request:
Use common-ground to help me think this through. I am not sure yet what good looks like. Look into the context, tell me in your own words what this is for and what would make it useful, and point out what I might assume without saying it. Then we decide the next step.
Before a long run, I start with this request:
Before you start, use common-ground. Tell me in your own words what this run is for and what would make the result useful to me. Ask only what you cannot settle yourself. Then give me the brief that you will run from, with the finish line and the check that shows it.
Common Ground:自主工作之前,先形成共同理解

common-ground 是一项技能。模型动手之前,它让模型和我先就任务的用途达成一致。模型用自己的话解释任务,提出自己的判断,再用预期用途检验这份理解。最后,模型写一份简报(brief),工作从这份简报开始。2026-10-04 之前,这项技能叫 deposition。
本文说明这项技能的各个部分,以及每个部分的理由。每个部分都来自我在实际会话中遇到的问题。
问题:结果正确,用途却不对
这项技能要解决的问题是:模型可以正确完成我的请求,却没有达到我的目的。模型完成工作,交出完整的结果,检查也全部通过。我读结果时才发现,模型是围绕另一个目标组织这项工作的。有时是我没有说出意图,有时是我自己还没想清楚。
本文自己就是一个例子。我让一个智能体就更新后的技能“再写一篇博客”,并要求它不提问、直接发布。智能体搜索了资料,写完文章,生成配图,然后发布。文章本身不差,但大约 40% 的内容和我之前写 deposition 的那篇重复。整个过程中,没有任何一步问过:第二篇文章是做什么用的。现在这篇合并后的文章替换了那两篇。
新模型让这个问题变得更大。Anthropic 的 Opus 5.5 使用指南告诉用户在一条消息里交代整个任务,并写明终点线(finish line),然后放手让模型工作。早期测试者让它在很少监督的情况下连续完成几个小时的编码任务。OpenAI 的 GPT-6 Astra 技能指南也要求作者写明决策边界和完成条件。
以前我和模型一轮一轮地工作,一次误解只损失一轮。我读回复,发现偏差,马上纠正。在长时间自主运行(long autonomous run,下文简称长时运行)中,我不在场。模型按照它对目标的理解,做几百个小选择。我到最后才一次看到全部选择。所以,误解的代价随运行时长增加,纠正的机会则移到了开始之前。
为什么叫“共同基础”
这个名字来自心理语言学。Herbert Clark 和 Susan Brennan 在 1991 年的论文《Grounding in communication》中,把对话看作一种共同行动。说话的人和听的人都要确认,听的人已经理解。双方一直确认,直到都相信这份理解足以满足当前目的。Clark 和 Brennan 把这个门槛叫作共同基础标准(grounding criterion)。双方由此建立的相互知识,就是共同基础(common ground)。
论文中有三点适用于这项技能。
- 共同基础是双向建立的。 旧名字
deposition的意思是证词录取:一方提问,另一方作答并记录在案。我从来不想要这种流程。模型提出它的解读和判断,我接受、修改或否定。双方都需要证据,证明对方已经理解。 - 标准取决于目的。 如果我反正会读下一条回复,大致相同的理解就够了,因为我可以在下一轮纠正。如果下一步是长时运行,模型要独自依据这份理解工作几个小时。所以,长时运行之前的标准更高。
- 媒介决定共同基础怎样建立。 面对面交谈时,双方能同时看到、听到对方。听的人可以在一句话说到一半时皱眉。长时运行去掉了这些信号。书面文字保留了两个特性:读者可以重读,作者可以在发出前修改。所以,长时运行需要的理解应写成文字,让运行过程可以随时重读。
模型用自己的话解释任务
第一步是复述。Lauren(@poteto)在 pstack 指南第二部分中推荐这种做法。这一部分的标题是“The art of supervising someone smarter than you”。她要求智能体用自己的话复述问题。复述能在智能体写代码之前暴露误解。而且,智能体没有照搬她的假设,所以它可能提出她没有考虑过的看法。
复述应表明立场,否则我没有可以检查的东西。假设我请模型为合作者的空间转录组数据准备第一次汇报:做基本的质量控制(QC),做一套幻灯片,不做深入的生物学分析。一个弱的回复只是重复请求:“你要基本的 QC 和一套幻灯片,不做深入的生物学分析。”这个回复准确,但我从中得不到任何信息。
一个有用的回复会给出判断:“这次会议要让合作者看到,我们已经收到并理解了他们的数据。QC 指标是这一点的证据,不是汇报的中心。”这个判断可能是错的。但我可以直接回应它,不需要自己先想好幻灯片结构。我只需要说,它离我的意思近不近。
这项技能把这种结果叫作共享心智模型(shared mental model)。模型把预期用途和约束、选择、后果、成功的证据连起来,再说明这些联系怎样决定它的判断。
这项技能的一个早期版本要求模型解释四项:目标、验收、验证、边界。这份清单变成了一张表格。模型填完每一项就停下了。实际任务可能取决于四项之外的东西。所以,现在的技能要求模型给出联系起来的整体图景,由具体任务决定哪些部分需要关注。
模型提出自己的判断
模型还应提出我没有要求的判断。一个只会把我说出的要求变精确的模型,帮不了我那些还没说出的要求。这项技能要求模型指出会影响任务的假设和默会判断。如果某种知识、替代方案或更好的框架会改变值得做的事,模型就应提出来。对每个想法,模型要说明它会改变什么,以及有什么证据支持它。
默会知识需要一个具体的引子。我遇到过两种情况:
- 我有一个判断,但还说不出来。比如,我知道幻灯片不对,却说不清为什么。这时一个候选说法有帮助,因为我能认出它,或者否定它。
- 我已经有这方面的知识,但没有把它和当前任务联系起来。模型可以提出这个联系。
两种情况下,一个对照例子或一张简短草图,都比抽象的问题更容易回答。模型不会声称知道我想要什么,由我的回应来决定。如果某个盲点需要更长的探索,配套技能 known-unknowns 负责这项工作,结果再回到共享心智模型中。
模型只问应由我做的判断
模型只应问我两类条件同时满足的问题:判断应由我来做,而且答案会改变方向。每个问题都有成本。Clark 还有一条原则,叫最小协作努力(least collaborative effort):双方会把达成理解的总工作量压到最小。问一个可逆的选择,是把工作从模型转给我,换来的理解却很少。
这项技能最初来自 Matt Pocock 技能合集中的一项提问技能。grill-me 的早期版本每次只问一个问题,并为每个问题给出推荐答案。我常常接受一个推荐,再接受下一个。每个选择我都确认了,我的目的却没有变得更清楚。如果提问的框架取决于一个没人说出的目标,这些回答就会把这个目标一直带下去。
后来的会话中,我看到了相反的问题。模型会问记录中早已定下的事,或者一分钟就能撤销的选择。现在这项技能规定:
- 提问之前,模型先检查之前的工作或记录是否已经回答了这个问题。
- 能查到的事实,模型自己查。
- 可逆、低风险的选择,模型自己决定,并列为我可以修改的默认选择。
- 模型只就目的、优先级和可以接受的代价问我。
长时运行之前,模型写一份简报
长时运行之前,模型应以一份简报结束讨论。简报是一条消息,运行可以直接从它开始。简报包含以下几项:
- 目标:结果是做什么用的。
- 终点线:怎样的行为算完成,以及证明它的检查。
- 已定决策:双方已经商定的内容,运行中不再重新讨论。
- 默认选择:模型自己做的可逆选择,列出来方便我修改。
- 有界的不确定性:仍然开放的问题,以及每个问题对结果的影响范围。
模型先和我确认简报,再从简报开始运行。我的常设指令有一个例外:如果对话已经确定了简报,模型不再询问。已经确定的简报,再确认一次只会多一轮对话,不会增加理解。
终点线能防止一种已知的失败。Anthropic 关于长时间运行智能体的运行框架(harness)的报告描述过这种情况:后来的一个会话看到已有进展,就宣布任务完成。把终点线写成“行为加检查”,运行就有了一项测试。没有这项测试,运行只能依靠一种“进展不错”的感觉。
Anthropic 在两个智能体之间也用了同样的思路。在他们的长时间应用开发框架中,生成者和评估者在写代码之前先商定一份冲刺合同(sprint contract)。合同写明生成者要做什么,以及评估者怎样验证。Common Ground 在模型和我之间做同样的约定,而且是在我还在场的时候。
简报的两条规则来自出过问题的运行:
- 每个数字都标明是估计。 计划中写“大约二十个测试”,本意只是大致规模,执行者却会把它当作要求。所以,简报中的每个数量或时长都标明“估计”。
- 按实际使用确定范围。 模型会问:有什么实际使用的证据,支持每项限制和机制?在我的会话中,最大的偏差都来自范围定错了:范围是按最坏情况定的,或者照搬了参照产品的功能清单。
检查通过,不等于适合实际用途
终点线是一项检查,而一项检查只能支持一个具体的结论。测试通过了,页面构建成功了,这些事实不能说明结果对读它的人有用。
本文的配图就是一个例子。第一张图满足了所有写明的要求:5:2 全景,纯白纸面,钢笔排线,没有文字。但其中一人把手杖举在肩膀高度,看起来像一支步枪。没有哪项检查能发现这个问题。我是以读者的眼光看图时才发现的。第二张图改正了这个姿势。
所以,模型结束讨论之前,要设想一次合理的实际使用,并从双方最不清楚的部分开始。模型要问:这项工作有没有可能满足所有写明的要求,却仍然辜负使用它的人?
这种设想要深入到内容本身。在汇报的例子中,双方同意“先做 QC,再做图,再做幻灯片”,并不能决定幻灯片说什么。模型逐页设想合作者会看到什么、会得出什么结论,才能看出双方对交付物本身是否一致。写文章也一样。一份简短的提要和大纲,标明读者和论点都是暂定的,就能让我在完整初稿依赖它之前,看到模型打算写什么。
终点线要在这一步之后才写进简报。对于探索性工作,双方商定要研究的问题和回顾结果的时间即可,不必事先知道答案。
一致不等于授权
模型应把“对任务的判断”和“行动的许可”分开。模型可以不经许可提出新的框架。但修改文件、扩大范围或制作新的交付物,需要授权。
我明确调用这项技能时,是在要求先讨论,再做依赖讨论结果的工作。我授权之后,模型就完成工作,不再为每个细节请求许可。一个未决问题只暂停依赖它的操作。例如,按钮文字已经定下,页面标题还没定,模型就先改按钮,标题问题留着。
“是”这个字曾经造成摩擦。如果模型问它的解读对不对,我说“是”,我确认的是一份理解。如果模型在已授权的任务中提出一项具体修改,我说“是”,模型就应去做。模型过去常常把任何“是”都当作开始的许可,或者再次请求我已经给过的许可。
两条规则减少了这种摩擦:
- 模型说明它是在核对一种解读,还是在提出一项行动,再结合这一点和已有的授权来理解我的回复。
- 我可以用两个可选信号。
agree确认理解,不开始新的工作。proceed授权已经确定的下一步。普通的话也可以,比如“好,就这样改”。proceed不能解决仍不清楚的范围。
运行过程中,共享心智模型指导各项选择。如果新证据改变了一个关键前提,模型就重新打开讨论。如果我在工作中途澄清了目标,模型要重新检查之前的选择,而不是只在旧计划上加一条新约束。
我根据证据修改技能
只有会话显示出具体原因时,我才修改技能。每次失败,我都先判断是三种原因中的哪一种:指导有缺口或不清楚;模型忽略了某条指令;必要的情境没有传给模型。如果每次结果不好就加一条规则,技能会变成一套沉重的流程。
会话显示的原因出乎我的意料:
- 在一个工具中,技能文本传给模型时在 8,000 个字符处被截断。关于授权的规则根本没有到达模型。技能安装正确,不等于模型读完了它。
- 在另一个会话中,模型拿到了完整文本,仍然两次请求许可。它把“暂停以解决分歧”理解成了“重新开始审批”。
- 在汇报工作中,目的变清楚之后,模型仍然保留了之前的图表选择。它改进了标题,却没有重新考虑图表本身。原因不是指令不够,而是模型没有继续使用我们已经建立的理解。
2026 年 9 月,我把技能缩减到原来篇幅的三分之一左右。我删除了通用的方法教程、重复的完整性清单,以及固定的两轮审批脚本。另一个模型审阅后,同意大部分删减。它也指出,有些删减去掉了真正的约束,比如只列出交付物,却没有说明它应包含什么。我把这些约束加了回去。
10 月的改动(新名字、简报和长时运行之前的触发)才三天。我有设计上的理由,也有促成这些改动的会话,但还没有测量过效果。
如何使用这项技能
Common Ground 收录在我的技能合集中。技能说明和来源说明都已公开。
面对一个新想法,我这样开始:
用 common-ground 帮我把这件事想清楚。我还不确定好的结果是什么样。先了解情境,用你自己的话告诉我这件事是做什么用的、怎样才算有用,并指出我可能默认了却没说出来的东西。然后我们再决定下一步。
长时运行之前,我这样开始:
开始之前,先用 common-ground。用你自己的话告诉我,这次运行是为了什么,怎样的结果对我真正有用。只问你自己无法确定的问题。然后给我你将据以运行的简报,写明终点线和证明它的检查。