June 17, 2026 ·
The question changed from ‘are agents real’ to ‘what gets agentized first’
NVIDIA, Microsoft, and Google all shipped enterprise agent platforms this quarter. The market settled the 'are they real' question — but the obvious way to answer 'which part first' is exactly the way that produces the six-month cleanup.
Somewhere in the last few weeks, the conversation flipped. Through 2025 the boardroom question was “are AI agents actually real, or is this another hype cycle?” In June 2026 — the month NVIDIA and ServiceNow announced governed autonomous agents for the enterprise, Microsoft put its agent control plane into general availability, and Google rolled its agent platform out as the successor to its older AI offering — the major platforms all started describing agents in the same language: systems with goals, memory, planning, tool use, and genuine autonomy. The market stopped asking whether agents are real. It started asking which part of the company gets agentized first.
That is the better question, and also the more dangerous one — because the obvious way to answer it is exactly the way that produces the six-month cleanup. “Which part first” tends to get answered by whoever is loudest, whichever process is most annoying this week, or wherever a vendor demo landed most impressively. None of those is the part of your business that should go first.
The two things that just made this easier — and the trap inside them
Two real shifts are behind the change in tone. Inference cost has fallen another sixty to seventy-five percent in a year, so the economics that made an agent fleet look expensive in 2024 no longer hold — the interesting move now is a routed fleet, where a cheap fast model handles the easy eighty percent and an expensive smart one takes the genuinely ambiguous decisions. And the interoperability layer matured, so wiring an agent into your existing tools is no longer a bespoke project per system. Building agents got cheaper and the plumbing got standard.
The trap is that “easier to build” and “safe to deploy into a system of record” are different statements, and the gap between them is where most agent projects still die. A control plane and a cheap routed fleet get you an agent that works in a demo. What separates that from an agent your operations team can live with on a Tuesday morning six months in is the unglamorous part: the evals that define what good looks like before any code is written, the observability that lets you replay any decision, the guardrails on what the agent is allowed to touch, and the runbook the on-call engineer reads when it misbehaves at 3am.
How to actually choose what goes first
The right first candidate is not the most painful process or the most impressive demo. It is the one that sits in a specific intersection: high enough volume that automation pays back, well-defined enough that you can write down what a correct outcome looks like, and forgiving enough that a wrong answer is recoverable rather than catastrophic while you build trust. Get one agent shipped into that intersection, prove it in production, and you have both the template and the organisational confidence to do the harder ones. Pick a high-stakes, ill-defined process first and you get an expensive failure that sets the whole programme back a year.
- Volume. Does this happen often enough that automating it returns more than it costs to build and maintain? Many annoying processes are simply too rare to be worth an agent.
- Definability. Can you write down what a correct outcome looks like, as an eval set, before building anything? If you cannot specify it, an agent cannot reliably do it — and you cannot tell when it regresses.
- Recoverability. When the agent gets one wrong — it will — is the cost an engineer rolling their eyes, or a regulatory finding? Start where mistakes are cheap and earn your way to where they are not.
- System depth. Does it touch a system of record? The moment an agent reads from or writes to the things your business runs on, the bar moves from “clever automation” to “production software with on-call.”
Choosing that first agent well, and building it so the second and third get easier rather than harder, is the core of an AI Agents engagement: production-grade agents with evals, observability, guardrails, and on-call runbooks, built on the SDKs your engineers already trust, and handed back to your team to own. We budget the demo at fifteen percent of the work and the part that keeps it running at the other eighty-five. It is not popular in pitch decks. It happens to be the difference between an agent that survives contact with a real operation and one that gets quietly switched off.
The honest counter-argument
There is a real case for moving faster and messier than this. If you wait for the perfect first candidate and the complete eval suite and the polished runbook, a more aggressive competitor ships three rough agents, learns from all of them, and is two quarters ahead of your careful first deployment. Speed has compounding returns, and over-engineering the first agent is its own failure mode. The reconciliation is to be fast on the things that are cheap to get wrong and disciplined on the things that are not — ship the low-stakes, well-defined agent quickly and learn from it, but do not let “move fast” talk you into pointing a half-built agent at your finance close or your customer records. The agentize-first decision is not “fast versus careful.” It is “fast where mistakes are recoverable, careful where they are not.”
What to do in the next 30 days
- List your candidate processes against the four tests. Volume, definability, recoverability, system depth. The right first agent usually scores well on the first three and low on the fourth.
- Write the eval set before the agent. If you cannot define a correct outcome for your chosen process, that is the work to do first — it is the spec.
- Resist the loudest process. The most painful or most-demoed candidate is rarely the right first one. Decide on the tests, not the volume of complaints.
- Plan for the eighty-five percent. Budget the observability, guardrails, and runbook from the start. The demo is the cheap part.
The market is right that the question has changed. “Which part gets agentized first” is the question of 2026. But the companies that will look smart in 2027 are not the ones who answered it fastest — they are the ones who answered it deliberately, shipped one agent into a place where it could prove itself, and built the discipline to do the next ones at increasing stakes. The platforms have made building easy. Choosing well, and building to last, is still the job.
Trying to work out which part of your business should get agentized first — and how to build it so it survives past the demo? Talk to Cravings about an AI Agents engagement. We help you pick the right first candidate, write the evals that define success, and ship a production-grade agent your team can own — with the observability and runbooks that keep it working six months in.