What did this correction mean?
- the agent’s output was wrong
- the instruction or workflow was wrong
- what the work required had changed
- insufficient information at task time
The reference set kept the correction, but not the explanation.
- The agent records the lease’s effective date as January 1.
- After review it becomes March 1. The reason field says “review comment”: it records that a change was requested, not why. Lime marks where expert judgment met the agent’s work.
- Four different explanations could lie behind the same correction, and each calls for different expertise.
- The reference set keeps March 1. The explanations fade; only the question marks remain.
An agent filled in a lease workpaper. It recorded the lease’s effective date as January 1.
After review, the effective date became March 1. The reason on file is “review comment,” and the item is marked resolved. That field records that a reviewer asked for the change. It doesn’t record why. The person who made the edit may simply have been responding to the reviewer’s note.
So what did the correction mean? Was the agent’s output wrong? Was the instruction or workflow wrong? Had what the work required changed? Or was there simply not enough information available at the time the task ran?
Those possibilities call for different expertise, and the person who entered the correction may not be the one who can tell them apart.
Later, the corrected workpaper becomes part of a reference set for improving the agent. The reference set keeps the accepted answer, March 1. It doesn’t tell us which of those things happened.
This example is a composite, but the situation is ordinary. My team monitors agentic systems in regulated expert workflows. I’m using financial audit as the example, though the same thing happens across professional services. An agent produces a workpaper; later we receive the corrected version as a reference set. In between, the work moved through questions, corrections, requests for more evidence, review notes, and sometimes changes to the procedure itself.
Five releases across production audit platforms · 80+ detection methods
When people evaluate agents in expert work, they usually ask whether the agent completed the task. That’s a reasonable question. But a correction like this one raises a different question: when experts correct an agent’s work, what does the organization learn from it?
What I want to propose is a way of improving not only the agents but the organization around them at the same time. That means judging agents not just on how well they complete tasks, but on how well they help people orient themselves to complex work. This is really an argument about incentives: about getting organizations, and the knowledge they collect, into the loop.
The route succeeded. Did the readings survive?
- The chart. Every sounding, contour and buoy was established by someone on an earlier passage.
- The same passage reduced to its result. The coastline is accurate and the route succeeded, but the readings are gone.
I sometimes use a navigational chart with clients, because it helps explain what gets lost when we keep only the end result. Every reading on a chart was established by someone, on an earlier passage: where the water was deep enough, where it wasn’t. Those readings help another person locate themselves and decide what’s safe, even when they take a different route. That is what I mean by keeping what the organization learned while doing the work.
And let’s be clear: getting to the correct output is a useful goal. For anyone improving an agent, it looks like a very good hill to climb. There’s an output, a corrected answer, and a way to measure progress.
But a successful route tells us that someone got through. The readings tell us more about how to navigate the next passage. If we keep only the corrected document, we may improve this task while losing what made the work acceptable. I call this loss expert signal evaporation.
Guides · Hutchins 1995 · Hollan, Hutchins & Kirsh 2000 · Star & Strauss 1999
The explanation might be far from the correction
- Corrections are easiest to see, and count, in the near corner, where expert judgment meets the agent’s output.
- A working hypothesis, not a measurement: meaning rises as capture falls.
- Start with the people accountable for quality. Requirements come from the far corner.
- Capture begins where the work happens. The observation is routed to the reviewer, partner or methodology team who can interpret it.
So what would we have to observe to keep those readings? In our work we distinguish five kinds of observation:
- 0OutputWas the answer correct?
- 1ExecutionWhat did the agent do?
- 2CoordinationDid the handoffs hold together?
- 3CommunicationWhat changed when the work met expert judgment?
- 4Fitness for purposeCan the organization stand behind this?
The first three are relatively accessible. We can compare answers, inspect execution, and often see technical handoffs. What gets harder is seeing what changed when the work met expert judgment: why the expert intervened, and later, whether the organization could actually stand behind the work. Those connections don’t appear automatically in an agent trace. We have to design for them.
Usually the agent is closest to one task and one practitioner. In the figure, moving up, we move through the organization: practitioner, reviewer, partner, methodology. Moving across, we move through the kinds of response the work may need, from a retry or a correction, through overrides and escalations, to working around the system or quietly withdrawing trust in it.
The near corner is easiest to see. We can count corrections there. But a correction alone doesn’t tell us which explanation is right. In the lease example, the effective date changed because a partner had spoken with the client. That information wasn’t known when the agent ran, and it wouldn’t ordinarily have reached the practitioner. The correction showed up at the practitioner’s level; its explanation sat with the partner. The person who suggests the correction may not have access to the system, and the person who enters it may not know why. If the agent’s view ends with its assigned task, that distinction disappears from what we can learn.
Our working hypothesis is that meaning rises as capture falls. The figure is illustrative, not a measurement.
Route the observation to the judgment
What if we started at the far corner? Ask the people accountable for quality what they would need to see to stand behind this work. That tells us what evidence to preserve where the work is happening.
Then keep the observation connected to the correction, and route it to the reviewer, partner, or methodology team who can interpret it. A local trace cannot make that judgment by itself. And the practitioner shouldn’t have to diagnose a workflow problem just because the correction happened at their desk.
Requirements come from the far corner. Evidence capture begins close to the work.
Guides · Majors et al. 2026 · Simon 1962 · Jaques 1990 · Malone, Laubacher & Dellarocas 2010
The record stays with the work
- Agent attemptwhat it tried to do
- Evidence usedwhat it relied on
- Expert interventionwhere judgment entered, by role
- Local contextthe circumstances of the work
- Dispositionwhat happened next
Recorded by role and authority, not by worker identity
- Even an agent that cannot finish the task can show what it attempted and the evidence it used.
- The design question: can that attempt be reliably connected to the expert intervention?
- …and to the local context and what happened next, so someone else can interpret it.
- The record names the role or authority behind a judgment, not the person.
Even when an agent can’t finish this task reliably, it might still be useful as an instrument. It can show what it attempted and the evidence it used.
The design question is whether we can reliably connect that attempt to the expert intervention, the local context, and what happened next.
That gives someone else a record they can interpret, even if the agent didn’t produce the final acceptable workpaper. The same agent doesn’t have to carry the whole task through the organization. The record has to stay with the work.
An instrument helps establish what happened and where judgment entered. It doesn’t decide which judgment is authoritative or what the organization should change. The record supports that judgment; it does not settle it.
And the same record could be used to monitor workers. We need to know the relevant role or authority behind a judgment (a reviewer, a partner, the methodology team) without making named workers the subject of measurement. The purpose of the record, and who can see it, have to be designed alongside the instrument.
An agent can be useful at observing the work before it can finish the task reliably.
Guides · Boston et al. 2026 · Star & Strauss 1999
What keeps causing this correction?
Possible explanations · the team’s reading, not the record’s
- the agent’s output→ repair the agent
- the instruction or workflow→ change the instruction
- what the work required→ revise how the work is reviewed
- insufficient information at task time→ wait for the evidence before the task runs
one linked record · lime: the expert intervention · grey: its lineage
- Linked records accumulate. Each keeps the attempt, evidence, intervention, context and disposition together. Nothing is sorted.
- A quality team investigates, reading each record alongside the work it came from.
- The team, not the record, attributes each correction to a possible explanation. Counts are illustrative.
- Each explanation calls for a different response.
Design proposal · not collected data
Now go back to the corrected lease. Suppose we kept the linked record each time a correction like it came up: the attempt, the evidence, the intervention, the local context, and what happened next. Over time, those records accumulate.
Keeping them doesn’t explain them. The record doesn’t classify the cause. It keeps the evidence together so the people with the right expertise can.
A quality team could then investigate, reading each record alongside the work it came from.
The investigation might find that the recurring problem is the agent’s output, the instruction or workflow, a change in what the work requires, or insufficient information at task time.
Those explanations lead to different responses: repair the agent, change the instruction, revise how the work is reviewed, or wait for the evidence before the task runs at all.
The next person inherits more than a corrected effective date. They inherit a better chart.
This is a design proposal. We don’t already have a dataset that answers the question.
Guide · Stamatis 1995
The next chart
- A correct output can be reached while navigating more blindly.
- With the readings kept, the next chart shows where it is safe to go, and more than one route through.
We can reach a correct output while navigating more blindly, discarding the readings that would help us choose the next route.
The corrections and questions already happen. Agents give us a chance to keep them connected to the work, so people can decide what to change and where to go next.
None of this needs an agent that can do perfect work. It needs an instrument that can observe reliably, get what it observes to someone who can judge it, and help change the next chart, without turning the instrument on the workers.
After the work is done, is the organization more capable of doing the next one?
Where this stands
- Observed
- In our monitoring work, corrected outputs became reference sets that kept the accepted answer but not its explanation, and technical traces were easier to capture than expert responses.
- Proposed
- Agents as instruments that keep a linked record of the work, by role, and route it to the people who can interpret it.
- Open
- What agents can observe reliably, how observations reach the right level of judgment, how to observe work without surveilling workers, and whether linked records improve organizational learning over time.
Our guides
Each note says what the source contributes and where the essay goes beyond it.
- Marisa Ferrara Boston · 2026
Agents Are Instruments, Not Pilots: Observing Expert Work beyond the Agent’s Task. Talk, Collective Intelligence 2026.
The talk this essay adapts, including the five kinds of observation and the term expert signal evaporation.
- Boston, Hanson, Georgala, Hudgens & Frase · 2026
Monitoring Agentic Systems Before They’re Reliable. Workshop on Agentic Software Engineering (AgenticSE), co-located with ACM CAIS 2026, San Jose, CA (non-archival). arXiv:2606.02494.
Shows monitoring producing useful structural findings before a system is reliable, using rule-based and statistical monitors on a synthetic 220-run testbed. It reports no production data and does not test agents as observers.
- Hollan, Hutchins & Kirsh · 2000
Distributed Cognition: Toward a New Foundation for Human-Computer Interaction Research. ACM Transactions on Computer-Human Interaction 7(2), 174–196. doi:10.1145/353485.353487.
- Edwin Hutchins · 1995
Cognition in the Wild. Cambridge, MA: MIT Press.
Ship navigation as cognition distributed across people, instruments, charts and procedures.
- Elliott Jaques · 1990
In Praise of Hierarchy. Harvard Business Review 68(1), 127–133.
Organizational levels of accountability. The practitioner-to-methodology levels are our mapping, not his.
- Majors, Fong-Jones & Miranda, with Parker · 2026
Observability Engineering: Achieving Production Excellence, 2nd ed. Sebastopol, CA: O’Reilly Media.
Background on what technical telemetry and traces capture. The comparison with expert judgment is ours.
- Malone, Laubacher & Dellarocas · 2010
The Collective Intelligence Genome. MIT Sloan Management Review 51(3), 21–31.
Who contributes, why, and how contributions are combined. Routing observations to the right judgment extends their framing.
- Herbert A. Simon · 1962
The Architecture of Complexity. Proceedings of the American Philosophical Society 106(6), 467–482. JSTOR 985254.
Organizations as hierarchic, nearly decomposable systems.
- D. H. Stamatis · 1995
Failure Mode and Effect Analysis: FMEA from Theory to Execution. Milwaukee, WI: ASQC Quality Press.
Background on failure mode and effects analysis: investigating recurring causes before choosing a response.
- Star & Strauss · 1999
Layers of Silence, Arenas of Voice: The Ecology of Visible and Invisible Work. Computer Supported Cooperative Work 8(1–2), 9–30. doi:10.1023/A:1008651105359.
How work becomes visible or invisible. Recording the role behind a judgment, rather than the person, is our design position.