The brief: An industrial manufacturer with 1,400 staff had bought AI assistant licences for 900 of them eighteen months earlier. Usage looked healthy — around 70% monthly active — and the board could not find a single number in the P&L that had moved. The CFO’s question was blunt: are we paying for a tool nobody needs, or are we using it badly? They asked us to answer that before the renewal.

What the assessment found

  • Usage was real and almost entirely trivial. We sampled two weeks of activity across six functions: the median task handed to the assistant took it under a minute to complete. Rewriting emails, summarising a document someone had already read, tidying meeting notes.
  • Nobody was delegating work. Not because they were sceptical — the enthusiasm was high — but because the four conditions that make delegation rational were all absent, and nobody had noticed that they were the constraint.
  • No stated permission. No function had written down what could be handed to a model unsupervised. In that vacuum people defaulted to the smallest safe unit, which is a paragraph of text.
  • No cheap verification. Checking a substantial output took as long as producing it, so delegating anything large was irrational even when it worked.
  • No system access. The assistant could not read the specification library, the tender portal, or the quality management system. It was a text box beside the work rather than inside it.
  • No recovery path. Where an error would reach a customer document uncaught, the only safe policy was to let the model advise and never act.

What we built

We deliberately did not start with a new platform. Three processes were chosen on volume, definability, and reversibility, and each got the four missing conditions rather than a new tool.

  • Tender response assembly. An agent with read access to the specification library, the previous tender archive, and the pricing sheet drafts the compliance matrix and the technical response, citing the source document for every claim. Bid managers edit rather than assemble.
  • Incoming specification review. Customer drawings and specifications are checked against manufacturing capability and the standards library, producing a flagged exceptions list. Engineers adjudicate the exceptions instead of reading three hundred pages to find them.
  • Supplier document triage. Certificates, test reports, and declarations of conformity are extracted, validated against what the purchase order required, and either filed or escalated.
  • An eval set per process, written first. Several hundred historical cases per process, re-adjudicated by the people who own the work, became the graded standard and the verification mechanism. This was the unlock: once a check costs minutes, delegating hours becomes rational.
  • A one-page delegation policy per function, signed by the function head, naming what may be done unsupervised, what requires review, and what may never be delegated. Boring, and the single most-cited artefact six months on.

What changed

  • Tender response preparation: 11 working days → 3, with bid managers reporting they now decline fewer opportunities on capacity grounds. Tender volume responded to rose 34% with the same team.
  • Incoming specification review: 6 hours of engineering time per specification → 40 minutes adjudicating a flagged exceptions list.
  • Supplier documents processed straight-through: 0% → 81%, with the remainder escalated for a documented reason.
  • Median delegated task duration across the three processes: under 1 minute → 22 minutes. This was the metric we managed to, because it is the one that tracks whether the tool is doing work or doing chores.
  • Licence renewal proceeded, at a reduced seat count. Around 300 of the 900 seats were genuinely unused and were cut; the saving funded the build.

What we left behind

Three processes running as delegated work with evals their own team maintains, a delegation policy per function that makes the next process easier to start, and a diagnostic the operations director now runs herself: for any candidate process, which of the four conditions is missing. Most of the time the answer is verification, and the fix is an eval set rather than a platform.

The finding worth repeating is that nothing here required a better model than the one they already had eighteen months earlier. The licences were not the problem and the enthusiasm was not the problem. The organisation had bought access to a capability and never granted itself permission to use it, then measured adoption instead of delegation and concluded from healthy-looking numbers that everything was fine.