August 17, 2026 ·
The AI gap tripled in five months. It is not about access.
OpenAI's Enterprise Signals data puts frontier firms at 8.3x the output tokens per user of typical firms, up from 2.6x in January. The gap is delegation depth, not licences — and the fastest growth is in legal, not engineering.
OpenAI published its Enterprise Signals data in mid-August and one number in it should reorganise how you think about your AI programme. Classifying the top ten percent of companies by output tokens per monthly active user as “frontier” firms and the 45th to 55th percentile as “typical”, it found that as of June the frontier group generated 8.3 times as many output tokens per active user. In January that gap was 2.6 times. It tripled in five months.
The instinct is to read this as an adoption story — the leaders bought more licences, the laggards did not. That reading is wrong and it leads to the wrong remedy. Everyone in both groups has access to the same models at the same prices. The gap is not in what they bought. It is in how much work they are willing to hand over.
Tokens per user is a delegation metric in disguise
Output tokens per active user is an awkward measure of almost everything except one thing, which it measures rather well: how big a unit of work the average person hands to a model. Someone asking a chat window to tidy an email produces a few hundred output tokens. Someone handing an agent a task that runs for twenty minutes across a codebase, a ticketing system and a database produces orders of magnitude more. Same licence, same model, same person — completely different relationship with the tool.
The supporting numbers confirm this is what is being measured. As of June, Codex accounted for 64 percent of combined Codex and ChatGPT output tokens among enterprise customers. The majority of enterprise AI consumption is no longer conversational. It is delegated execution — jobs dispatched and collected rather than questions asked and answered. That is the behaviour that separates the two groups, and it is a behaviour, not a purchase.
Which is why the gap tripled so fast. Adoption curves do not triple in five months; they grind. Behaviour changes fast once the surrounding conditions permit it, and it does not change at all while they do not. The frontier firms did not out-buy anyone. They removed whatever was stopping their people from handing over the whole task instead of a fragment of it.
The 108x figure is the one to sit with
Buried in the same release is a breakdown that undercuts almost every assumption about where enterprise AI value lands. Since February, weekly active enterprise Codex users grew 108 times in legal, 41 times in sales and recruiting, and 26 times in marketing — against 5 times in engineering.
Engineering is the slowest-growing category for a coding tool. Not because engineers stopped adopting it, but because they adopted it in 2024 and the curve has flattened. The growth has moved to the functions nobody built the tool for. Legal teams are running structured work through a coding agent at a hundredfold rate not because lawyers learned to code, but because a coding agent is a general-purpose system for doing multi-step work against documents and data with a verifiable result. Contract review, obligation extraction, diligence checklists, clause comparison across a portfolio — this is the same shape of work as a refactor.
If your AI programme is still organised around engineering productivity, you are funding the flattest part of the curve. The tenfold-plus movements are happening in functions that have never had a serious automation budget and have no internal capability to build one. That is where the neglected value is, and that is where nobody in the organisation is asking for it — because those teams do not know that the tool their engineers use is the tool they need.
Why the laggards are stuck, and it is not enthusiasm
Every organisation on the wrong side of this gap that we have looked at was blocked by the same four things, in roughly this order.
- No permission to delegate. Nobody had said, in writing, which categories of work may be handed to a model without a human in the loop. In the absence of that statement people default to the safest possible unit of delegation, which is a single paragraph of text.
- No way to verify the output. Delegating a twenty-minute task is only rational if checking it takes less than twenty minutes. Without evals, tests, or a review surface, the check costs as much as the work, so people delegate only what they can eyeball.
- No access to the systems. A model that cannot read the contract repository or write to the case management system can only ever be a text generator. The frontier firms wired the tools into the systems where the work lives; the rest left a chat window at the edge of the business.
- No recovery path. Where a wrong action cannot be caught and reversed, the sane response is to never let the agent act. Reversibility, not model quality, is usually the binding constraint.
None of these is solved by better models or more seats, which is why another year of licence renewals will not move the number. They are diagnosed by looking at what a team actually does on a Tuesday and finding which of the four is stopping them — the substance of an AI readiness assessment, and the reason it examines your data, systems, and operating model rather than your tooling. The constraint is almost never the thing the leadership team assumes it is.
The honest counter-argument
Take the source seriously: OpenAI sells tokens, and this is a report demonstrating that the most successful companies consume many more tokens than everyone else. That is a sales chart with a research framing, and the causal arrow it implies is exactly the one that benefits the publisher. Token volume measures consumption, not value. A firm running an inefficient agent that retries four times and reasons at length about trivialities will post frontier-tier numbers while producing nothing, and a disciplined team with a well-routed pipeline that sends the easy eighty percent to a small cheap model will look like a laggard on this metric while getting more done for less. We have argued at length that routing down to smaller models is one of the highest-return moves available, and that advice makes your token count worse.
The industry breakdown cuts the same way. The largest gap is in information and technology at 11.7 times and the smallest is manufacturing at 5.3 times, which is roughly what you would expect if the metric partly tracked how text-shaped an industry’s work is rather than how effectively it uses AI. A manufacturer is not eight times worse at this than a software company; a manufacturer has less work that takes the form of tokens. So do not manage to this number. The defensible reading is narrower: a gap that triples in five months is measuring something that changes fast, and delegation depth changes fast while procurement does not. That is the signal worth acting on, and it survives the scepticism about the metric.
What to do in the next 30 days
- Measure your own delegation depth, not your seat count. Ask how long the longest task anyone routinely hands to a model runs for. If the answer is measured in seconds, licences are not your problem.
- Write down what may be delegated. One page, per function, naming the categories of work that can go to a model unsupervised, supervised, or not at all. Ambiguity here is what keeps people asking the model to rewrite sentences.
- Look outside engineering. Find the function with the most structured document work and no automation budget. On the published growth rates, that is where your next order of magnitude is.
- Fix verification before capability. Whatever you want to delegate, build the check first. Delegation scales exactly as far as your ability to confirm the result cheaply.
The uncomfortable implication of a gap that triples in five months is that it can triple again, and that the position you hold in it is mostly a function of decisions you have not made rather than technology you have not bought. The companies pulling away are not the ones with the best models. They are the ones that answered the delegation question early, in writing, and then built the verification to make the answer safe.
Suspect your organisation is on the wrong side of this gap but cannot say why? Talk to Cravings about an AI readiness assessment — two weeks, across your data, systems, team and operating model, ending in a ranked list of what is actually blocking delegation and what it costs to unblock.