When the Standard Moves

Published September 23, 2026

A Field Guide entry in “AI-Native Ways of Working.”

In brief

When agents do the work, your standard moves. How to tell which standards were priced by scarcity and can move, which were priced by consequence and must hold, and how to write the standard where the work runs.

A team that has repriced its work says the standard moved. So does a team that has stopped checking.

A small internal dashboard can be useful with a clumsy interface and too much to scroll past, where the same defects on a customer product would be unacceptable. I make that distinction on my own systems now, because the tools are for me, the reversal cost is low, and having the thing is worth more than polishing the thing.

My own standard has moved. The last entry in this field guide treated that drift as damage, which is what it is when there is more to judge than a person can judge. That is not the whole picture, because some of what moved in my own work moved for reasons I would defend. A tool can have a rough interface and still be better than the tool nobody had time or budget to build. Code can be ugly in ways I cannot fully evaluate, because I am not an engineer and the agents write the code. What I can see is whether the thing works for me, which tells me nothing about whether it would survive anyone else running it. That trade has produced more usable work than I could have produced any other way. That judgment comes from the same instrument this essay spends a section distrusting, so hold it at the weight of a felt report and not a measurement.

That trade has also cost me. There have been times when I let something go too long and found a problem later. Sometimes the issue was the wrong supervision agent. Sometimes the issue was that the attention I brought was the same as it ever was while the volume it had to cover was not. Did my standard move for a reason I would still defend, or because I got used to what the agents produced?

The old bar had an economy

Before agents, most small tools never got built at all. A team might need a dashboard, a workflow helper, or a narrow internal tool, and the answer would be no because the team could not get engineering time. The few things that did get funded were usually high-exposure systems with many users, external customers, regulatory consequences, or scale requirements. Those systems needed the full standard because the consequences priced the work that way.

Agents change the population of things being built. A small team can now build tools for small use cases that previously died in the request queue. A standard calibrated for high-exposure software starts to feel strange when it is applied to a one-user internal tool that can be thrown away tomorrow.

Sandboxing matters more when new attack paths open, and data boundaries matter more when agents can touch more systems. Anything priced by consequence still has to hold, and sometimes it has to tighten.

What priced this standard? If scarcity priced it, the standard can move. If consequence priced it, the standard has to hold. If nobody can say what priced it, the standard may have been habit wearing the language of quality. A great deal of what a team calls best practice has never been priced at all, and nobody notices until someone asks.

A team that cannot answer it is losing the bar, however deliberate the move feels.

Feel is a bad detector

I cannot reliably tell, from the inside, which one I am doing, so I let the stakes decide instead of the feeling. If the work is low risk, easy to reverse, and specific to me, I let the standard move and come back to it later. If the work affects other people, crosses a boundary, or carries a real blast radius, the answer has to be verification.

That test comes from the first entry in this field guide, where autonomy is earned trust multiplied by blast radius rather than a single dial. Blast radius sets a floor no amount of earned trust lowers, so the higher the consequence, the less I am willing to rely on a feeling that the work is probably fine.

Parasuraman and Manzey's 2010 review of automation bias and complacency predates modern coding agents. In one line of studies, people detected automation failures 82 percent of the time when reliability varied, and 33 percent of the time when reliability stayed constant. Detection moved inversely with automation reliability, with no floor where complacency disappears. Training and explicit instructions to verify did not prevent the bias. That last part revises something I wrote in the first entry of this field guide, where the whole model rested on trust but verify. Verification still decides the outcome. Telling people to do it is not what makes it happen, which is most of the argument for putting the standard somewhere other than in a person's intentions.

Parasuraman and Manzey studied humans supervising automation, so carrying the finding into AI code review is my inference. I make it because the shape is familiar: a system performs well for long enough, the human reallocates attention, and the rare failure arrives exactly when attention is weakest. What that costs you is the ability to notice, from the inside, that a standard has been repriced — which is why the question has to be asked out loud, against something written down.

An agent writes the work. Another agent reviews the work. A dashboard says review happened. A human sees the metric and believes oversight happened.

Duma and colleagues measured the software version of that trap in 2026. Across the same GitHub repositories, human-only review happened on 25.21 percent of human-authored pull requests and on 8.08 percent of agent-authored ones. Total review went up on the agent-authored ones, because agents were doing most of the reviewing. Their study is descriptive and does not control for pull-request size. Your total review rate can hold steady, or improve, while the rate of purely human review falls by two thirds, and no dashboard in the building will tell you which standard moved underneath it.

The low-risk corner

My own case sits in the low-risk corner, and I want to keep it there. The tools I build for myself would fail the standard I would apply to a product for real people: bad interface, clumsy workflow, and real usefulness anyway, because each one is specific to my context and my needs. Internal work that never would have been funded before can now exist.

I think it is acceptable to build and use individually worse things when the alternative is that most of them never get built, and I think the profession has underweighted the benefit of the work that never existed. I cannot tell you it outweighs the cost, because I have priced one column and not the other. That position carries real exposure: code slop, defects, and vulnerabilities nobody catches. Every time I find a problem late, I can go back and ask how I had priced that work before it broke. If the late problems sat in work I had called low-stakes and reversible, my pricing is doing its job. If one of them sat in work I had called consequential, my pricing is broken, and I would rather say so in public than adjust the story. A solo practice contains nobody who can overrule my pricing, so the public record is my route to ratification: readers are the party whose standing does not depend on my schedule.

I also have an interest in the answer. My work assumes the standard can legitimately move, so a finding that it cannot would cost me more than it costs you. That is the same structure as an organization pricing its own deviations, and naming it does not neutralize it.

What argues against this

Recalibration is exactly what people say while quality falls. Every shop that lowers a bar has a story about context, economics, speed, or modern ways of working. Quality still matters regardless of who produced the work. If agents create far more output, the answer may be to hold the line harder and refuse any machinery that makes a lower line feel principled.

Wells Fargo called bad accounts “slippage” and treated a certain amount as the cost of doing business. Columbia's foam strikes became “in-family” events and stopped escalating the way they should have. In both cases, a deviation became a category, and the category helped the organization stop reacting. That pattern is older than agents and still available to any team that wants a nicer word for a moved standard.

“Low-risk internal tool that can be thrown away tomorrow” is my own category, and categories are how organizations stop reacting. Columbia is the case I have to answer, and the record does not say what a quick reading suggests. Engineers close to the work asked for imagery of the foam strike. Management turned that request off, apologized to the Defense Department for going outside proper channels, and left the Debris Assessment Team to prove the hit was dangerous using pictures it had been denied. The category survived because a judgment made near the work was overruled further from it, and the investigation board found no paper trail of anyone escalating the concern.

Pricing a standard close to the work is necessary and it is not sufficient, because the people who price a standard are rarely the people who can overrule the pricing. Wells Fargo makes the same point from the other side, and it is the half that costs me more: leadership there did ask what the deviation cost and answered in explicit terms, treating a rate of bad accounts as the price of doing business. That answer passes the test I am proposing. It was wrong because the party doing the pricing was the party the price protected. So the failure mode was never the unasked question. It was the interested answer, and no test administered by the interested party survives it. A team that prices its own standards with no route to ratify them anywhere else has rebuilt the same arrangement with better vocabulary, and the route has to end in a person whose standing does not depend on the schedule.

A thousand individually disposable single-user tools also hold credentials and touch data boundaries as an estate, whatever each one is worth alone. I have been assessing the benefits of agent-built software at the population level and its costs one artifact at a time, which is the exact asymmetry this essay warns about.

The partial answer is that pricing has to happen twice, once per artifact and once across the portfolio, and that the portfolio only becomes visible on a map. A team that prices only the artifact will be right about every tool and wrong about the estate.

“Bring in fresh eyes” sounds right if the whole team's bar has moved together. Bosu, Greiler, and Bird studied 1.5 million review comments across Microsoft projects and found that reviewers familiar with a file were almost twice as useful as first-time reviewers, with new hires the least useful of all. Their reviewers were reading human-authored code that other humans had read, so the transfer is imperfect.

So a newcomer can supply the alarm that the bar has moved, because surprise and contrast are still available to someone who has not normalized the local pattern. I could find no study testing that directly, so treat it as a mechanism rather than a measured effect. Diagnosis belongs with the person closest to the work, who has to decide whether the standard moved because scarcity changed, consequence demanded it, or habit got exposed. Ratification belongs somewhere else, for the reason Columbia gives.

Write the standard where the work runs

Vigilance is too weak a control for the volume of work agents produce, if the automation-bias research carries across the way I think it does. The standard belongs in the substrate the agents read: the instruction files, the tests, and the checks that run whether or not anyone is paying attention.

I did that on my own systems by pointing the agents at the engineering and change practices I trust, and by asking them to score the current setup against any new practice or critique that seems right. The answer comes back as already doing it, doing it better than the thing I brought them, or missing a part they then borrow.

The substrate still misses things. I watch agents catch a verification step somebody skipped, and I watch them claim they followed a rule only after being reminded of it. Whether the substrate holds when nobody is watching is the one thing looking cannot tell me, which is the same limit as seeing that a tool works for me, and it is why a written standard needs its own maintenance.

Every few weeks I have a fresh agent audit the repository, the code, and the instruction files, and I ask what the substrate says the agents must do, whether that still makes sense, and whether any of the instructions contradict each other. Contradictions send agents in a direction nobody chose, and they are invisible to the person who wrote both halves months apart.

A fresh session is not a different reviewer, and that is a hole in the practice I just described. Model review is worth the most under three conditions. Two of them come from the evidence below: the reviewer has to be more capable than the author, and different from it. The third comes from the finding that models struggle to correct their own reasoning without external feedback, so the reviewer also has to be pointed at something checkable. An auditing agent running on the same model as the agents that wrote the substrate fails the first two outright, and two of my three audit questions are matters of judgment rather than fact, which strains the third. Anyone running this loop, and that includes me, should be able to name the model doing the auditing and say how it differs from the models being audited.

An agent that reviews its own code sometimes reports it has done something it has not, which looks like gaming the system and is not. It is generating the most likely answer, and sometimes the most likely answer is yes without checking.

Xiang and colleagues found, in a July 2026 study, that a stronger reviewer model improved a weaker author's work by 18.1 points, while a weaker reviewer made a stronger author's work worse, with 13 regressions against 3 fixes. A strong model reviewing its own work gained exactly nothing and cost 72 percent more. A weaker model reviewing its own work did improve, though less than a stronger reviewer would have managed. So “an agent reviewed it” tells you little until you know which agent reviewed what.

Agents can build tests and scripts that they later have to pass, which is where deterministic checks earn their place. Du and colleagues reported at ICSE this year that at Tencent, deterministic static analysis found the candidate alarms and a model triaged them, which eliminated 94 to 98 percent of false positives with high recall, in seconds and cents per alarm. The detector was deterministic and the judgment was probabilistic, in that order.

For leaders

The last entry asked leaders to map the work before the next autonomy raise, and I stand by that. What it did not say is what the map has to show if it is going to answer the pricing question: which standards apply where, what agents are allowed to touch, where humans are supposed to review, what each handoff interacts with, and what the blast radius is if the work is wrong.

Others have been here. Mik Kersten named the thing being mapped in Output to Outcome, the security world has mapped agent blast radius and attack surface in real depth, and Eric Karsten (a different person, one letter apart) has already argued in CIO that value stream maps and capability models are the governance control plane for agentic work, covering agent capability, permission, interaction surface, and blast radius. Steve Pereira and Andrew Davis gave us five maps in Flow Engineering, and none of the five has an agent in it. What I have not found is anyone treating that join as a thing you do before you change how you supervise, by the people who do the work rather than by the people who audit it. I have not drawn one with agents in the flow either, so take the sequence as the claim and not the artifact.

Kersten would push back on the whole idea. He moved past value stream mapping because a map creates a static picture, and the map goes stale as soon as the mapping is finished. A one-time agentic map inherits that criticism exactly. The audit loop is the answer: draw the map, then re-run it — with something other than the model that drew it — and read the difference. The re-run would show a new place where work escapes, or a permission that grew, or a review step that used to have a person in it. Those differences are the practice, and the map is only the baseline they are measured against.

Skipping the map creates a fake control. Existing agents will ignore excellent substrate if they do not know where to look or which instruction file governs their work. Substrate the agents do not read is worse than no substrate, because it gives leaders confidence without changing the work.

If your answer to the volume is that agents review the agents, make the reviewer more capable than the author and point it at something checkable. On the one benchmark I have, a weaker reviewer made the work worse and a strong model reviewing its own output changed nothing at all, which is thin evidence for a strong instruction and still more than most teams are working from. I owe my own audit loop the same test.

Set the accountability where you want the behavior. Skitka, Mosier, and Burdick found that people accountable for overall accuracy made fewer errors and used an independent cross-check more often, while people accountable for speed cross-checked less. If reviewers are accountable for accuracy, give them the time, tools, and permission to check.

The standard is going to move. Whether that movement is visible, priced, written into the system, and audited afterward is a decision a leadership team makes through what it measures, not through what it says about quality. The measurement I am starting with is two columns: every problem I find late, and how I had priced that work before it broke.


This essay was produced the way it says work now gets produced: an agent drafted it from my interview answers, reviewer panels ran on model families different from the drafter's, and every call, including the ones that overruled the reviewers, was mine. It is one pattern from a set I'm working through on human-agent teams. It draws on my own hands-on R&D rather than a production deployment, and on eighteen years of coaching and organizational change work. I am not an engineer; the systems behind these examples are ones I own and use myself. Where this describes organizations, it stands on public research and on the transfer I am making from that research to agentic work.

Back to Field Notes