Consistency Isn't the Model's Job

Published October 8, 2026

A Field Guide entry in “AI-Native Ways of Working.”

In brief

AI models are built to vary. When you need the same result every time, have an agent build the script that does the work, and keep the model for the work that needs judgment.

The expectation

Engineers I coach keep bringing me the same complaint. They asked an AI agent to move one function, a small piece of code, from one file to another. It wrote five hundred lines of new code. They knew exactly what change they wanted. When they looked, the agent had added a lot and removed nothing. As far as they were concerned, it had failed to follow a simple instruction.

People who don't write code expect AI to do their work their way, too. They want a status report, a forecast or a change order to come back the same way every time, laid out and calculated the way they'd have done it. I use AI to draft nearly everything I work on, financial forecasts and change orders included, so I understand why.

A language model writes by choosing each next word from a range of likely words. Ask it the same question twice and you can get two different answers. It was built to work that way. If you need the same answer every time, you're asking the model to do the wrong job.

What deterministic means

When I call a tool deterministic, I mean it gives the same output every time it gets the same input. A calculator works that way. Put in the same numbers today or a year from now and you get the same answer. A spreadsheet with a bad formula is deterministic too. It gives you the same wrong total every time you open it.

Many AI tools let you set something called temperature, which controls how much the model varies its word choices. Turn it down and the model sticks closer to its most likely words. The answers vary less, but they still vary.

Researchers at Thinking Machines Labasked one model to tell them about the physicist Richard Feynman a thousand times, with the temperature at zero. They got eighty different answers. Every answer started out word for word the same. The first difference came partway in. Most said Feynman was born in Queens, New York, and a few said New York City. Either answer is acceptable. The team found the cause and changed how the model ran on their own servers, and after that, all thousand answers came out identical. That version ran slower, and it isn't a setting you can switch on in the AI tools most of us use.

Suppose a vendor did give you that setting. A model that gives the identical answer every time will also repeat its wrong answers exactly. Even then, I'd want a person to read the answer before anyone relied on it.

Should you use AI for the work at all?

If a script or a scheduled task could have done a piece of work reliably before AI came along, it almost certainly still should. An agent can run the script you already have, or write the one you always meant to build. I'd start with the work that can run the same way every time, and get that scripted before anything else.

Where you aren't sure, try the AI version and put its result next to what your current system or process produces. Run it more than once, since one good result tells you little about the next one. Look at whether the numbers match what you would have produced, and whether you'd send the result to a client without fixing it first. If the AI version is better, use it. If it only matches a script or automation you already have, you're paying on every run for a result you already had. If it falls short, look at whether a script could do that work instead. I'd rather you find this out on your own work than take my word for it.

Have the AI build a script

I'm not an engineer, and agents write all of my code. When I need something done the same way every time, I have an agent write a script that does it. From then on, the agents run the script.

My AI-Native Engineering Maturity Model, a guide I publish for engineering leaders, comes in two PDFs, a 28-page full edition and a 7-page excerpt. A script builds both of them from one source file that holds all the content. When something needs to change, an agent edits the source file and the script rebuilds both PDFs, laid out the same way every time. AI vendors measure use in tokens, small pieces of words, and an agent rewriting a 28-page document uses a lot of them. I had it built this way so the agents wouldn't burn tokens rewriting the whole document for every revision. It also keeps the two editions in step with each other and makes each new version easy to track.

I did the same with my writing. Drafting agents kept producing sentences like “The useful question is simple:”, which only announce what's coming instead of saying it. I was catching them by hand. I could have asked another agent to read every draft and judge it. I'd have paid tokens every time it read a draft, and it might not have judged the same way twice. An agent wrote a short script for me. It scores the first sentence of every paragraph the same way every time and sends back any draft with too many of those sentences. It only catches one kind of bad sentence, and only at the start of a paragraph. I still decide whether a draft is any good.

Code editors already move a function from one file to another the same way every time. An agent can use the editor's own move command to make the change and leave the rest of the code alone. That's what the engineers I coach wanted in the first place. When the agent writes new code, I'm still working out when to tell it how to do the job and when to just judge what comes back.

Every time an agent redoes work it has done before, it spends tokens on it again. Agents will write the same code over and over instead of writing it once and reusing it, and they'll repeat steps a script could run. A script costs tokens when an agent writes it and when an agent changes it. In between, running it costs almost nothing in tokens, because the script does the work. I'd rather spend tokens on work that needs a model.

An agent once built me a tool to track my to-dos and follow-ups. The agent that built it told me it could record them, and when I tried it quickly, it worked. Then I gave it a real message that listed several follow-ups, and it filed the whole list as a single item. I've learned to try anything an agent builds on real work before anyone relies on it. After that, I trust it until its inputs change. In my setup, agents usually review the scripts other agents build. For work other people rely on, I want the first real review done by a person who knows what the right result looks like.

Where the model belongs

Say you want a summary of ten news sources every day. A script can pull them and paste the pieces together, and the summaries can come out cut off halfway or badly formatted. The script also can't tell which details matter. I'd give that job to an agent that reads the sources and writes the summaries itself. It can pick out what's relevant and leave out the filler.

Someone I know runs tabletop role-playing games, where one player tells the story for everyone else at the table. He built a tool to help run his campaigns, and he built it so it could surprise him. The tool runs part of the story, which lets him play in his own campaign. If it told the same story every time, he'd have no use for it.

Before I'd hand a judgment call to a model, I'd split it. Which parts could be decided ahead of time as a rule? A script can handle those. For the parts that remain, which calls should a person make, and which can an agent make? When there's a lot to sort, a simpler tool can give each item a score first, for example how positive or negative a message sounds. The clear cases follow a rule, and only the unclear ones go to the model.

Take a weekly status report. A script can pull the totals and dates from wherever the team tracks its work, the same way every week. A model can write the paragraph at the top about what changed. It reads only what it needs for that paragraph, so it uses fewer tokens, and it never touches the totals.

In a 2024 study, software researchers built a system called LILACto read computer logs, the line-by-line records software keeps of what it did. It matched each line against patterns it had already seen and asked the model only about lines it hadn't. Each answer from the model became a new pattern, so the next line like it didn't need the model. The sets of logs averaged about 3.6 million lines each, and it called the model about 280 times per set. It was also more accurate than the rule-based tools it was compared against. Logs repeat themselves far more than most business work, so that study shows the best case.

When one agent's output is the next agent's input, a small error at the first handoff grows at every handoff after it. In a chain like that, I'd make the first step a script that checks the inputs before any agent reads them. That way a bad input gets caught once, at the start, instead of being passed down the chain.

The other side

Plenty of people who build with these models every day would disagree with me. Models keep getting better. With good instructions and the temperature turned down, they'd say, an agent can do repeated work directly and do it well enough. Scripts are extra work to build and maintain. They break when anything around them changes, and they follow their rules even when a case calls for judgment.

In a 2022 Stanford study, a large language model beat two established tools on messy data-cleaning jobs, such as filling in missing values, once it was shown a few worked examples. One tool repairs data using statistics. The other writes reformatting rules from examples. Without the worked examples, the model did worse than the rule-writing tool. With them, it beat both.

An agent doing repeated work directly spends tokens on every run, and because its output can vary, someone has to look over every run as well. What you save on those runs can pay for building the script and fixing it when it breaks. For work that repeats often, I'd make that trade.

When a website changes its layout or a file shows up in a new format, a script that worked yesterday fails today. When a script breaks, you can see that it broke and why, and fix it, often with an agent's help. That only works if someone finds out it broke. Give every script a named owner. Have it check its own inputs, and set an alert for when it doesn't run.

Models change under you too. A group of researchersgave GPT-4 the same set of easy coding problems in March and again in June of 2023. In March, more than half of its answers ran without anyone changing them. By June, one in ten did. The June version had started wrapping its code in extra text, and that caused most of the drop. The researchers warned that people can easily miss a shift like that when a model's output feeds a larger system. I'd rather fix a script that told me it stopped than miss a change like that.

When people say a script can't use judgment, I agree, and for most repeated work that's what I want. The totals in a status report shouldn't depend on anyone's judgment, a model's included. If a model is using judgment to add up a column, I've given it the wrong job. Where the work does need judgment, like the messy data in that Stanford study, I'd give it to the model, after taking out what a rule can handle.

For leaders

For each piece of AI work your teams repeat, ask how often it runs, what a wrong result costs and who it lands on, how stable its inputs are, and how many people or agents depend on it. If it runs often and other people depend on the result, and its inputs rarely change, have an agent build a script for it. When the inputs change every time, or the work needs someone to read the context, I'd give it to a model. I'd split off the parts a rule can handle first.

You can ask an agent what work it keeps repeating. If your team has code, you can also have an agent go through it and list what could be scripted or scheduled. Token spend is a reasonable place to start, because repeated work shows up there. I keep an eye on mine through a dashboard I had built.

When someone hands you a forecast or a report that AI had a hand in, ask what produced it. If a script did, ask who owns the script and when someone last looked at its inputs. If an agent wrote it this morning, ask who read it before it reached you.


This is one pattern from a set I'm working through on human-agent teams. It comes out of my own hands-on R&D, not a production deployment.

Back to Field Notes