cankun.me

Common Ground: Shared Understanding Before Autonomous Work

Oct 7, 2026
Two people stand by a stone marker on an open plain; one points a staff toward a distant pole on the horizon
Two people stand by a stone marker on an open plain; one points a staff toward a distant pole on the horizon · View full resolution

common-ground is a skill that aligns the model and me on what a task is for before the model does the work. The model explains the task in its own words, adds its own judgment, and tests that understanding against the intended use. At the end, the model writes a brief that the work can start from. Until 2026-10-04, the skill had the name deposition.

This essay explains the parts of the skill and the reason for each part. Each part comes from a problem that I saw in real sessions.

The problem: a correct result for a different purpose

The skill exists because a model can complete my request correctly and still miss my purpose. The model does the work, produces a full result, and passes its checks. Then I read the result and see that the model organized the work around a different goal. Sometimes I did not say my intention. Sometimes I did not know it yet.

This essay has its own example. I asked an agent for “another blog” about the updated skill, and I told it to publish without questions. The agent searched the web, wrote the essay, made the image, and published. The essay was good, but about 40% of it repeated my earlier essay on deposition. Nothing asked what the second essay was for. This merged essay replaces both.

Recent models make the problem larger. Anthropic's guide to Opus 5.5 tells users to give the whole task in one message and to name the finish line. Then the user lets the model work. Early testers ran coding tasks for hours with little oversight. OpenAI's skills guidance for GPT-6 Astra also asks authors to state decision boundaries and the condition for completion.

When I worked with the model turn by turn, a misunderstanding cost one turn. I read the reply, saw the drift, and corrected it. In a long autonomous run (in this essay, a long run), I am not there. The model makes hundreds of small choices under its reading of the goal. I see all of the choices at the end. So the cost of a misunderstanding increases with the length of the run, and the chance to correct it moves to the start.

Why the name is “common ground”

The name comes from psycholinguistics. Herbert Clark and Susan Brennan's paper “Grounding in communication” (1991) describes conversation as a joint activity. The speaker and the listener both work to confirm that the listener understood. They continue until both believe that the understanding is sufficient for the current purpose. Clark and Brennan call that threshold the grounding criterion. The mutual knowledge that the two people build is their common ground.

Three parts of the paper apply to the skill.

  1. Grounding goes in two directions. The old name, deposition, described one person who asks questions and one person who answers for the record. I never wanted that procedure. The model gives its interpretation and its judgment. I accept it, change it, or reject it. Each of us needs evidence that the other understood.
  2. The criterion depends on the purpose. If I will read the next reply anyway, a rough shared picture is sufficient, because I can correct it in the next turn. If the next step is a long run, the model must act alone for hours on the understanding. So the criterion is higher before a long run.
  3. The medium controls how grounding can occur. Face-to-face conversation lets people see and hear each other at the same time. A listener can frown in the middle of a sentence. A long run removes these signals. Written text keeps two properties: a reader can read it again, and a writer can correct it before sending. So the understanding for a long run must be a written text that the run can read again.

The model explains the task in its own words

The first step is a restatement. Lauren (@poteto) recommends this practice in the second part of her pstack guide. The part has the title “The art of supervising someone smarter than you.” She asks the agent to restate the problem in its own words. The restatement shows a misunderstanding before the agent writes code. Also, the agent did not copy her assumptions, so it can find a view that she did not consider.

A restatement must take a position, or it gives me nothing to check. Suppose I ask for a first presentation on a collaborator's spatial transcriptomics data: basic quality control (QC) and a slide deck, with no deep biological analysis. A weak reply repeats the request: “You want basic QC and a slide deck, without deep biological analysis.” The reply is accurate, but I cannot learn anything from it.

A useful reply states a judgment: “This meeting must show the collaborators that we received and understood their data. The QC metrics are evidence for that. They are not the center of the talk.” The judgment can be wrong. But I can react to it without a slide plan of my own. I can say if it is close to my meaning.

The skill calls the result a shared mental model. The model connects the intended use to the constraints, the choices, the consequences, and the evidence of success. Then it shows how those connections control its judgment.

An earlier version of the skill told the model to explain four fields: goals, acceptance, verification, and boundaries. The list became a form. The model filled in each field and stopped. A real task can depend on something that no field names. So the skill now asks for the connected picture, and the task decides which parts need attention.

The model adds its own judgment

The model must also give judgment that I did not ask for. A model that only makes my stated requirements precise cannot help with requirements that I did not state. The skill tells the model to state assumptions and tacit judgments that affect the task. The model also brings knowledge, alternatives, or a better framing when they change what is worth doing. For each idea, the model says what the idea changes and what evidence supports it.

Tacit knowledge needs a concrete prompt. I see two cases:

  • I have a judgment that I cannot say yet. For example, I know that the slides are wrong, but I cannot say why. A candidate statement helps, because I can recognize or reject it.
  • I already have the knowledge, but I did not connect it to the task. The model can propose the connection.

In both cases, a contrasting example or a short sketch is easier to answer than an abstract question. The model does not claim to know what I want. My answer decides. When a blind spot needs a longer exploration, the companion skill known-unknowns does that work, and the result goes back into the shared mental model.

The model asks only for my judgments

The model must ask me only for judgments that belong to me and that can change the direction. Every question has a cost. Clark's related principle of least collaborative effort says that partners keep the total work of understanding small. A question about a reversible choice moves work from the model to me, and the question adds little understanding.

The skill started as a fork of a questioning skill in Matt Pocock's skill collection. The early version of his grill-me skill asks one question at a time and gives a recommended answer for each question. I often accepted one recommendation, then the next. I confirmed each choice, but my purpose did not become clearer. If the framing of the questions depends on a goal that nobody said, the answers carry that goal forward.

In later sessions, I saw the opposite fault. The model asked about points that the record already settled, or about choices that I can undo in a minute. Now the skill sets these rules:

  • Before the model asks, it checks if earlier work or the record already answers the question.
  • The model finds discoverable facts itself.
  • The model decides reversible, low-risk points itself and lists them as defaults that I can change.
  • The model asks me about purpose, priorities, and acceptable costs.

Before a long run, the model writes a brief

Before a long run, the model must close the discussion with a brief. The brief is one message that the run can start from. It contains these items:

  • Goal: what the result is for.
  • Finish line: the behavior that counts as done, and the check that shows it.
  • Settled decisions: what we agreed, so that the run does not open the decision again.
  • Defaults: the reversible choices that the model made, so that I can change them.
  • Bounded uncertainties: the open questions, and how much each one can affect the result.

The model then confirms the brief with me and starts the run from it. My standing instruction makes one exception: if the conversation already settled the brief, the model does not ask again. A second confirmation of a settled brief adds a turn and no understanding.

The finish line prevents a known failure. Anthropic's report on harnesses for long-running agents describes a later session that sees progress and declares the job done. A finish line written as a behavior and a check gives the run a test. Without the test, the run has only a feeling of progress.

Anthropic uses the same idea between two agents. In their harness for long-running application development, the generator and the evaluator agree on a sprint contract before they write code. The contract says what the generator will build and how the evaluator will verify it. Common Ground makes the same agreement between the model and me, while I am still present.

Two rules for the brief come from runs that went wrong:

  • Mark each number as an estimate. If a plan says “about twenty tests” as a rough scale, the executor treats the number as a requirement. So each count or duration in the brief has the word “estimate.”
  • Size the scope from real use. The model asks what evidence of real use supports each limit and mechanism. In my sessions, the largest misalignments came from scope that matched worst cases or the feature list of a reference product.

A passed check does not prove fitness for purpose

A finish line is a check, and a check supports only a specific claim. The tests pass. The page builds. These facts do not show that the result is useful to the person who reads it.

The preview image for this essay is an example. The first image met every written requirement: a 5:2 panorama, pure white paper, ink hatching, no text. But one figure held the staff at shoulder height, and the staff looked like a rifle. No check could find that fault. I saw it when I looked at the image as a reader. The second image fixed the pose.

So before the model closes the discussion, it walks through one plausible use of the result. It starts with the part that we understand least. It asks: can the work meet every stated requirement and still fail the person who uses it?

The walk-through must reach the content. In the presentation example, an agreement on the sequence “QC, then figures, then slides” does not decide what the slides say. When the model walks through what the collaborators see on each slide and what they conclude, the walk-through shows if we agree on the deliverable. For an essay, a short synopsis and outline with the audience and the argument marked as provisional do the same work. I see what the model will write before the full draft depends on it.

The finish line goes into the brief after the walk-through. For exploratory work, we agree on the question and a time to review the findings. We do not need the answer in advance.

Agreement and authorization are different

The model must keep a judgment about the task apart from permission to act. The model can propose a new framing without permission. To change files, increase the scope, or make a new deliverable, the model needs authorization.

When I invoke the skill explicitly, I ask for discussion before the dependent work. After I authorize the work, the model completes it without more requests for permission. An open question stops only the actions that depend on it. For example, if the wording of a button is settled and the page title is open, the model changes the button and keeps the title question.

The word “yes” caused friction. If the model asks whether its interpretation is correct and I say yes, I confirm an understanding. If the model proposes a specific change in an authorized task and I say yes, the model should do the change. Models often treated every yes as permission to start, or they asked again for permission that I already gave.

Two rules decrease this friction:

  • The model says whether it checks an interpretation or proposes an action. It reads my reply against that and against the authorization that already applies.
  • I can use two optional signals. agree confirms the understanding and starts no new work. proceed authorizes the settled next action. Ordinary words, such as “yes, make those changes,” also work. proceed does not settle a scope that is still unclear.

During the run, the shared mental model controls the choices. If new evidence changes an important premise, the model opens the discussion again. If I clarify the goal during the work, the model examines the earlier choices again. A new constraint added to the old plan is not sufficient.

I revise the skill from evidence

I change the skill when a session shows a specific cause. For each failure, I find out which of three causes applies. The guidance had a gap or was unclear. The model ignored an instruction. Or the necessary context did not reach the model. If I add a rule after every bad result, the skill becomes a heavy process.

The sessions showed causes that I did not expect:

  • In one tool, the skill text reached the model cut at 8,000 characters. The rules about authorization never arrived. A correctly installed skill is not always a completely read skill.
  • In another session, the model had the full text and still asked for permission twice. It read “pause to resolve a disagreement” as “start the approval again.”
  • In the presentation work, the model kept its earlier figure choices after the purpose became clearer. It improved the titles but did not examine the figures again. The cause was not missing instructions. The model did not continue to use the understanding that we built.

In September 2026, I cut the skill to about one third of its earlier length. I removed generic method tutorials, repeated completeness checklists, and a fixed two-round approval script. A review by a second model agreed with most cuts. The review also found cuts that removed real constraints, such as a deliverable without a statement of its content. I put those constraints back.

The October changes (the new name, the brief, and the trigger before a long run) are three days old. I have the design reasons and the sessions that caused the changes. I do not have a measured result yet.

How to use the skill

Common Ground is in my skill collection. The skill instructions and the source notes are public.

For a new idea, I start with this request:

Use common-ground to help me think this through. I am not sure yet what good looks like. Look into the context, tell me in your own words what this is for and what would make it useful, and point out what I might assume without saying it. Then we decide the next step.

Before a long run, I start with this request:

Before you start, use common-ground. Tell me in your own words what this run is for and what would make the result useful to me. Ask only what you cannot settle yourself. Then give me the brief that you will run from, with the finish line and the check that shows it.

← Back to Writing