ELU NOVA
← All field notes
The pilot-to-production gap

Agent washing: telling a real agentic system from a renamed chatbot

5 min read

A term with a definition problem

In a press release dated 25 June 2025, Gartner made two predictions that were reported together and should be read separately.

The first: more than 40% of agentic AI projects would be cancelled by the end of 2027, attributed to escalating costs, unclear business value and inadequate risk controls.

The second, which received less attention and is more immediately useful to a buyer: of the thousands of vendors describing their products as agentic AI, Gartner estimated only around 130 were building anything that genuinely warranted the label. The rest were, in Gartner's framing, engaged in "agent washing" — rebranding existing AI assistants, robotic process automation and chatbots without adding substantially agentic capability.

Anushree Verma, the Gartner analyst quoted in that release, made a further point that is easy to skim past: many use cases positioned as agentic today do not require an agentic implementation at all.

Two separate problems, then. Some of what is sold as agentic is not. And some of what genuinely is agentic should not have been, because a simpler mechanism would have been cheaper to build, cheaper to run, and considerably easier to audit.

What "agentic" actually claims

The word has been stretched to the point of uselessness in marketing material, so it is worth stating the distinction in operational terms.

A scripted automation executes a predetermined sequence. Read the queue, extract these fields, apply this rule, write to that system. If something falls outside the script, it stops or it escalates. The path is knowable in advance, which means it can be tested exhaustively and explained to an auditor line by line.

An agentic system is given a goal and some latitude in how to reach it. It decides which tools to call, in what order, and when the goal is met. The path is not fully knowable in advance. That latitude is the entire value proposition, and it is also the entire risk.

The practical consequence: an agentic system needs a control surface that a scripted automation does not. Boundaries on what it may touch. Limits on what it may do without approval. A decision log detailed enough to reconstruct why it did what it did. Rollback for actions already taken downstream.

If a vendor is offering you agency without offering you those four things, the interesting question is not whether the product is agentic. It is whether they have thought about what happens when it is wrong.

A buyer's checklist

Questions that a rebranded chatbot cannot answer well:

"Show me a run where the system chose a different path than it did in the demo, and explain why." Genuine agency produces path variance. If every run looks identical, the system is following a script, which may be entirely fine — but then you should be paying script prices and getting script-level auditability.

"What is the boundary, and what enforces it?" The answer should describe a technical control — scoped credentials, an allowlist of permitted actions, an approval gate before consequential writes. If the answer is that the prompt instructs the model not to do certain things, that is not a boundary. It is a request.

"What does the decision log capture, at what granularity?" You want: input, model version, tools called, the reasoning trace or its structured equivalent, confidence where available, the human reviewer if one was involved, timestamp, and the action taken. If the log captures only input and output, you cannot answer an auditor's question six months later.

"What is the per-transaction cost at our volume, including retries?" Agentic systems make multiple model calls per transaction, and the count varies with difficulty. Cost per transaction at pilot volume on easy cases is not the number you will pay. Gartner named escalating costs first among the three cancellation causes for a reason.

"How does it fail?" Ask for a specific failure and what happened next. A vendor who cannot produce one either has not run at production volume or is not being straight with you.

"What is your exception rate in production, and where does an exception go?" The exception path is where most of the real cost sits. A vendor without a production exception rate does not have production deployments.

The harder question: should this be an agent at all

This is the part of the conversation that vendors have no incentive to raise, and it is the one that most often saves money.

A large share of the workflows presented to us as agentic candidates are rules-based processes with a stable decision tree and a manageable set of exceptions. For those, a deterministic implementation is cheaper to build, dramatically cheaper to run, faster in execution, testable to completion, and trivial to explain to a regulator. It also does not get worse when a model version changes underneath it.

Agency earns its cost where the path genuinely cannot be enumerated — where the number of possible situations is large, the right next step depends on context that cannot be captured in a rule, and the cost of an occasional wrong turn is low relative to the value of handling the long tail at all.

That describes some enterprise workflows. It does not describe most of them.

We put this in writing during the discovery audit, in a section listing the workflows where we recommend against AI or against agency specifically. Clients quote that section back to us more often than any other part of the roadmap, which tells you something about how rarely they hear it.

What to do before the next vendor meeting

  1. Write down the decision the system would make and the latitude it would need. If you can enumerate the paths, you probably do not need an agent.
  2. Ask for the per-transaction cost at your volume, on your difficulty mix, including retries.
  3. Ask what enforces the boundary, and require a technical answer.
  4. Ask for the production exception rate and where exceptions go.
  5. Cost the deterministic alternative before you approve the agentic one. Sometimes it wins. That is a good outcome, not a failed project.

Gartner's 40% cancellation prediction runs to the end of 2027. The projects that survive it will mostly be the ones that were scoped honestly at the start.

Find out where AI actually pays off in your business.

A focused discovery audit with a senior principal. You leave with a prioritized, costed roadmap — whether or not you work with us.

contact@elunovatech.com · Response within one business day