Skip to content
AZGARD
perspectives

Same model, 8.3x the output. The gap is AI skills, written down

Frontier firms get 8.3x the output per user from the same models. OpenAI's data points at written-down AI skills and connected company data.

Angus McDonald · 13 Aug 2026 · 10 min read

Same model, 8.3x the output. The gap is AI skills, written down

OpenAI published its enterprise telemetry yesterday, and buried in it is the end of an argument I have been having with business owners all year.

Firms it classes as "frontier" now generate 8.3 times the output tokens per active user that typical firms do. In January that multiple was 2.6. The gap more than tripled in five months, and it did that while both groups sat in front of the same models on the same subscriptions. OpenAI puts it plainly: frontier firms "have access to the same AI models as other enterprises, but they are deepening their use faster".

So the question I get asked most often, which model should we standardise on, turns out to be close to irrelevant. What matters is the fortnight after you pick one.

Why do some companies get so much more out of AI than others

OpenAI answers this in its own report, and the answer has nothing to do with how clever the model is. Two features separate the groups. Skills, which are reusable written instructions the model loads when it needs them. And plugins, which connect it to the systems where company data lives.

At frontier firms, 19% of weekly active users use skills. At typical firms it is 3%. Plugins run 21% against 9%.

That 3% is the number I would put on a wall. In the overwhelming majority of businesses nobody has written down how the place makes its decisions in a form a model can pick up. Not because it is hard. It has never been anybody's job.

The industry breakdown says the same thing sideways: information and technology at 11.7x, manufacturing at 5.3x. Typical firms did not stand still either, growing 1.9x to 2.8x across the year. Genuine improvement, in a straight line, while the top of the market compounded.

Credit to OpenAI for putting the caveat in its own report: "Tokens are an imperfect measure of business value: a short response can be highly valuable, while a long one may add little." Worth holding onto. A firm burning eight times the tokens is not automatically producing eight times the value, and I would not sign off on a business case built on token counts. A 2.6x gap widening to 8.3x in five months is still not noise.

The engineering head start was a filing head start

The same report explains why engineering ran ahead of every other function, and this is the line I would print out. "Codebases give agents clear context, tests make outputs easier to verify." Knowledge work lags because its tasks "provide limited context, can be difficult to specify, and lack clear criteria for verifying the result."

Read that as a business diagnosis rather than a technical one. Engineering got there first because a codebase is a business that already wrote itself down: one place, one format, with an automated way to check whether the answer was right. Every other function in your company runs on rules that were never written anywhere, producing outputs nobody agreed how to grade.

Codex accounted for 64% of combined Codex and ChatGPT output tokens among enterprise customers as of June. The interesting part is what has happened since February. Enterprise Codex use grew 108x in legal, 41x in sales, 41x in recruiting and 26x in marketing, against 5x in engineering. None of those functions have a repo. They are getting there anyway, which is the encouraging read: the context problem is solvable outside a codebase, it just has to be solved deliberately.

Anthropic deleted 80% of its instructions and lost nothing

Now the argument being thrown at everything I have just written.

In a post from late July, Anthropic reported that it "removed over 80% of Claude Code's system prompt" for models like Claude Opus 5 and Claude Fable 5, with no measurable loss on its coding evaluations. It has been circulating hard this week, and the take riding on it is blunt: delete your instruction files, the model already knows.

The resolution sits in the same post. Skills earn their keep when they "encode particular opinions, knowledge, or best practices that are particular to you, your team, or product". Alongside it: "Avoid stating 'the obvious' things Claude should know by looking at your file system or your repo."

The 80% Anthropic cut was the generic 80%. Instructions telling a very capable model things it could work out by looking around.

Which means the frontier gap is not about volume of writing at all. It separates businesses that have written down the decision logic nobody outside the building could infer from businesses that have written nothing, or that have written forty pages of generic AI advice the model never needed. A fat prompt full of common knowledge costs you money and buys you nothing. Two pages naming your margin floor, your approval thresholds, the exceptions your best estimator applies without thinking, and what you do with a job you have never quoted before: that is an asset, and no vendor can hand it to you. It is the same forty pages I was banging on about when Meta made the model free, arrived at from the other direction.

What AI skills and plugins for business actually are

Strip the vendor vocabulary and there are two jobs.

A skill is a written instruction for a decision your business makes over and over. How a quote gets built, or when a discount needs sign-off. It sits in a file, the agent loads it when that situation comes up, and it improves every time somebody corrects it. The hard part is not technical: writing one forces you to state a rule you have been carrying around unstated, and occasionally to find that two people in your business hold different versions of it.

A plugin, or app, or connector, whatever this month's word is, is a live line into a system holding facts: your job management platform, your accounting file.

Both are things you build once and maintain, and neither shows up well in a demo. That is most of the reason so few businesses have them: nobody sells a quarterly retainer for a decision you wrote down yourself.

How to give AI agents access to company data safely

The access half of the gap picked up its own number this week. MIT Technology Review, working with Google Cloud, surveyed 300 data and technology executives and found that on average only about 45% of company data is reachable by agentic AI. The organisations it classes as data leaders get past 70%. Everybody else sits at 30% or below.

Then the finding I keep rereading. 100% of those data leaders say they trust the accuracy and decisions of their agents. Across the full sample it is about half.

The obvious reading is that confidence follows competence. I think the causation runs the other way. An agent working from 30% of your data is guessing at the rest, producing plausible answers that are wrong in ways nobody can trace back to a missing file. Distrust is the correct response to that agent. The leaders did not talk themselves into trusting theirs. Their agents can see the file.

Which is where the safety question gets asked backwards. In Upwork's SMB research, data security is the single biggest barrier to adoption at 27%, and I have never thought that concern was silly. But the failure I actually get called about is not a leak. It is an agent with access so thin that it improvises, and improvisation on a customer quote costs more than a permissions review would have. Scope access per system and per person, write down what the agent may not touch, and you are most of the way there.

The rest of the MIT numbers describe the market I walk into every week: 98% of those organisations use agentic AI, 73% on a limited number of use cases, and one in ten widely. 66% say legacy systems are what stops them scaling.

The strongest argument against writing anything down

Skan AI raised a $63M Series C this week, co-led by Cathay Innovation and Dell Technologies Capital, on the premise that enterprise work context is becoming an infrastructure layer the way CRM became the system of record for customer data. Its method is observation: it watches how work actually moves through an organisation. At one large US bank it says it tracked 11.2 million context switches across 1,500 finance staff and surfaced $37M in operational friction. Those figures are company-supplied, so hold them loosely.

The line that should sting comes from the investor. Simon Wu at Cathay says Skan is the only company he has seen building work context "from direct observation rather than from documentation or system logs". And Avinash Misra, the chief executive: "Everyone is obsessed with building a better car. We think the bigger opportunity is building a better navigation system."

I take that seriously because I made half of the argument myself a couple of days ago. Documentation lies. Your written policy and your practised policy drift apart, and the written one ages worse. Observation flatters nobody. It surfaces the workaround your team invented in March and never mentioned upstairs.

Where I get off the bus is what observation can produce. It answers "what happens here". An agent making a decision needs "what should happen here". Watch a sales team for a year and you will learn that they discount inconsistently. You will not learn what the discount should be, because that number is a judgement somebody has to make out loud.

Observation is a superb instrument for finding what needs writing down. Somebody still has to write it.

Why your AI pilot is not delivering productivity gains

The encouraging thing for a small business is that none of this is capital-intensive.

Upwork's research on AI in SMBs, published in June, covers 195 leaders at firms of 10 to 99 people. 62% are very or extremely confident handing high-stakes tasks to AI agents. 32% already describe agents as mission-critical. 74% report improved productivity, and then the sting: those improvements "remain under 25% for most".

Conviction is there. Access is there. Returns stalling in the low twenties inside a business that already believes and already has the tools is the signature of model access with nothing wired in behind it.

One more figure worth sitting with. Six months after adoption, OpenAI found early-career employees sending 13 more messages per week than executives. The people who know the firm's rules are barely touching the thing, and the people using it hardest have not learned the rules yet. A written skill is the only mechanism I know that closes that gap: the senior person's judgement, in a file, available to the graduate at nine on a Tuesday night.

What a small business needs before AI agents actually work

Roughly four things, none of which need a procurement process.

  • One workflow where the decisions carry real money. Not the one that irritates you most.
  • The rules that govern it, written where an agent can read them: current numbers, the exceptions, and an honest note on what the rules do not cover.
  • A connection into whichever system holds the facts, so the agent stops guessing at the 70% it cannot see.
  • A person's name against it, and a date to check it again.

Then delete everything you wrote that a good model could have worked out on its own. That step is what Anthropic's post is really about, and it is the difference between a skill and a prompt full of throat-clearing.

The firms pulling away are not smarter and they are not better funded. They wrote it down, and five months ago they were only 2.6x ahead. If you are waiting on the next model release to change your position, you are watching the wrong number, and buying another tool before you have mapped the work will not move it either.

FAQ

tags: skillscompany-brainadoptioncontext

Angus McDonald

Angus McDonald

Founder, Azgard

Builds and operates production AI systems for organisations that need results, not slide decks.