Agents Are Instruments, Not Pilots

Observing expert work beyond the agent’s task

Marisa Ferrara Boston · Adapted from a talk at Collective Intelligence 2026 · About a five-minute read

1 · A corrected workpaper

What did this correction mean?

Workpaper · Lease — effective date
agentJanuary 1
after reviewMarch 1
reasonreview commentrecords that a change was requested, not why
statusresolved
Possible explanations
  • the agent’s output was wrong
  • the instruction or workflow was wrong
  • what the work required had changed
  • insufficient information at task time

The reference set kept the correction, but not the explanation.

Field-grounded composite

An agent filled in a lease workpaper. It recorded the lease’s effective date as January 1.

After review, the effective date became March 1. The reason on file is “review comment,” and the item is marked resolved. That field records that a reviewer asked for the change. It doesn’t record why. The person who made the edit may simply have been responding to the reviewer’s note.

So what did the correction mean? Was the agent’s output wrong? Was the instruction or workflow wrong? Had what the work required changed? Or was there simply not enough information available at the time the task ran?

Those possibilities call for different expertise, and the person who entered the correction may not be the one who can tell them apart.

Later, the corrected workpaper becomes part of a reference set for improving the agent. The reference set keeps the accepted answer, March 1. It doesn’t tell us which of those things happened.

This example is a composite, but the situation is ordinary. My team monitors agentic systems in regulated expert workflows. I’m using financial audit as the example, though the same thing happens across professional services. An agent produces a workpaper; later we receive the corrected version as a reference set. In between, the work moved through questions, corrections, requests for more evidence, review notes, and sometimes changes to the procedure itself.

Five releases across production audit platforms · 80+ detection methods

When people evaluate agents in expert work, they usually ask whether the agent completed the task. That’s a reasonable question. But a correction like this one raises a different question: when experts correct an agent’s work, what does the organization learn from it?

What I want to propose is a way of improving not only the agents but the organization around them at the same time. That means judging agents not just on how well they complete tasks, but on how well they help people orient themselves to complex work. This is really an argument about incentives: about getting organizations, and the knowledge they collect, into the loop.

2 · Expert signal evaporation

The route succeeded. Did the readings survive?

A nautical chart: soundings, depth contours, buoys, rocks and a compass rose, with a route drawn through them to a vessel. The same passage with every reading removed: only the coastline, an island, the route and the vessel remain. The chart The result

I sometimes use a navigational chart with clients, because it helps explain what gets lost when we keep only the end result. Every reading on a chart was established by someone, on an earlier passage: where the water was deep enough, where it wasn’t. Those readings help another person locate themselves and decide what’s safe, even when they take a different route. That is what I mean by keeping what the organization learned while doing the work.

And let’s be clear: getting to the correct output is a useful goal. For anyone improving an agent, it looks like a very good hill to climb. There’s an output, a corrected answer, and a way to measure progress.

But a successful route tells us that someone got through. The readings tell us more about how to navigate the next passage. If we keep only the corrected document, we may improve this task while losing what made the work acceptable. I call this loss expert signal evaporation.

Guides · Hutchins 1995 · Hollan, Hutchins & Kirsh 2000 · Star & Strauss 1999

3 · Where expert signal is hardest to see

The explanation might be far from the correction

Organizational scope × expert responseIllustrative
Organizational scope rises from practitioner to reviewer, partner and methodology. Expert responses run from retry and correct to override, escalate, work around and withdraw trust. Responses near the corner are captured in system traces; further out they need review context; furthest out they are visible only over time, if at all. captured in system traces needs review context visible only over time, if at all ORGANIZATIONAL SCOPE ↑ practitioner(evidence) reviewer(expertise) partner(relationships) methodology(quality) retry correct reclassify override escalate workaround withdrawtrust EXPERT RESPONSE → WHERE EXPERT JUDGMENT AND SYSTEM OUTPUT MEET Hypothesis:meaning rises as capture falls. requirements comefrom the far corner

So what would we have to observe to keep those readings? In our work we distinguish five kinds of observation:

  1. 0OutputWas the answer correct?
  2. 1ExecutionWhat did the agent do?
  3. 2CoordinationDid the handoffs hold together?
  4. 3CommunicationWhat changed when the work met expert judgment?
  5. 4Fitness for purposeCan the organization stand behind this?

The first three are relatively accessible. We can compare answers, inspect execution, and often see technical handoffs. What gets harder is seeing what changed when the work met expert judgment: why the expert intervened, and later, whether the organization could actually stand behind the work. Those connections don’t appear automatically in an agent trace. We have to design for them.

Usually the agent is closest to one task and one practitioner. In the figure, moving up, we move through the organization: practitioner, reviewer, partner, methodology. Moving across, we move through the kinds of response the work may need, from a retry or a correction, through overrides and escalations, to working around the system or quietly withdrawing trust in it.

The near corner is easiest to see. We can count corrections there. But a correction alone doesn’t tell us which explanation is right. In the lease example, the effective date changed because a partner had spoken with the client. That information wasn’t known when the agent ran, and it wouldn’t ordinarily have reached the practitioner. The correction showed up at the practitioner’s level; its explanation sat with the partner. The person who suggests the correction may not have access to the system, and the person who enters it may not know why. If the agent’s view ends with its assigned task, that distinction disappears from what we can learn.

Our working hypothesis is that meaning rises as capture falls. The figure is illustrative, not a measurement.

Route the observation to the judgment

What if we started at the far corner? Ask the people accountable for quality what they would need to see to stand behind this work. That tells us what evidence to preserve where the work is happening.

Then keep the observation connected to the correction, and route it to the reviewer, partner, or methodology team who can interpret it. A local trace cannot make that judgment by itself. And the practitioner shouldn’t have to diagnose a workflow problem just because the correction happened at their desk.

Requirements come from the far corner. Evidence capture begins close to the work.

Guides · Majors et al. 2026 · Simon 1962 · Jaques 1990 · Malone, Laubacher & Dellarocas 2010

4 · Agents as instruments

The record stays with the work

One linked record
  1. Agent attemptwhat it tried to do
  2. Evidence usedwhat it relied on
  3. Expert interventionwhere judgment entered, by role
  4. Local contextthe circumstances of the work
  5. Dispositionwhat happened next

Recorded by role and authority, not by worker identity

Even when an agent can’t finish this task reliably, it might still be useful as an instrument. It can show what it attempted and the evidence it used.

The design question is whether we can reliably connect that attempt to the expert intervention, the local context, and what happened next.

That gives someone else a record they can interpret, even if the agent didn’t produce the final acceptable workpaper. The same agent doesn’t have to carry the whole task through the organization. The record has to stay with the work.

An instrument helps establish what happened and where judgment entered. It doesn’t decide which judgment is authoritative or what the organization should change. The record supports that judgment; it does not settle it.

And the same record could be used to monitor workers. We need to know the relevant role or authority behind a judgment (a reviewer, a partner, the methodology team) without making named workers the subject of measurement. The purpose of the record, and who can see it, have to be designed alongside the instrument.

An agent can be useful at observing the work before it can finish the task reliably.

Guides · Boston et al. 2026 · Star & Strauss 1999

5 · The same correction, seen again

What keeps causing this correction?

Linked records over timeDesign proposal · not collected data
Retained · not yet interpreted A quality team investigates, record by record

Possible explanations · the team’s reading, not the record’s

  • the agent’s output→ repair the agent
  • the instruction or workflow→ change the instruction
  • what the work required→ revise how the work is reviewed
  • insufficient information at task time→ wait for the evidence before the task runs

one linked record · lime: the expert intervention · grey: its lineage

Design proposal · not collected data

Now go back to the corrected lease. Suppose we kept the linked record each time a correction like it came up: the attempt, the evidence, the intervention, the local context, and what happened next. Over time, those records accumulate.

Keeping them doesn’t explain them. The record doesn’t classify the cause. It keeps the evidence together so the people with the right expertise can.

A quality team could then investigate, reading each record alongside the work it came from.

The investigation might find that the recurring problem is the agent’s output, the instruction or workflow, a change in what the work requires, or insufficient information at task time.

Those explanations lead to different responses: repair the agent, change the instruction, revise how the work is reviewed, or wait for the evidence before the task runs at all.

The next person inherits more than a corrected effective date. They inherit a better chart.

This is a design proposal. We don’t already have a dataset that answers the question.

Guide · Stamatis 1995

6 · The readings kept, and more than one route

The next chart

The passage reduced to its result: coastline, island, route and vessel, with no readings. The chart again with all its readings, and a second, dashed route alongside the first. The result alone The next chart

We can reach a correct output while navigating more blindly, discarding the readings that would help us choose the next route.

The corrections and questions already happen. Agents give us a chance to keep them connected to the work, so people can decide what to change and where to go next.

None of this needs an agent that can do perfect work. It needs an instrument that can observe reliably, get what it observes to someone who can judge it, and help change the next chart, without turning the instrument on the workers.

After the work is done, is the organization more capable of doing the next one?

Guide · Malone, Laubacher & Dellarocas 2010

Where this stands

Observed
In our monitoring work, corrected outputs became reference sets that kept the accepted answer but not its explanation, and technical traces were easier to capture than expert responses.
Proposed
Agents as instruments that keep a linked record of the work, by role, and route it to the people who can interpret it.
Open
What agents can observe reliably, how observations reach the right level of judgment, how to observe work without surveilling workers, and whether linked records improve organizational learning over time.

Our guides

Each note says what the source contributes and where the essay goes beyond it.