Meta made the model free. Your context is the bill.
Meta's Muse Glimmer runs a capable agent on one GPU. A 128K context window is rationed VRAM, so running AI agents locally rewards the tidy business.
Angus McDonald · 12 Aug 2026 · 9 min read

On this page
Meta put Muse Glimmer out on Monday and my feed did the thing it does. Open weights, Apache 2.0, thirty billion parameters, tuned for "always-on local agent workflows", small enough to sit on one consumer graphics card. If you run a small business and you have spent the past year being quoted per-seat pricing on cloud AI, that headline reads like the bill just went to zero.
The bill didn't go to zero. It moved somewhere you can't pay it with a credit card.
A 128K context window is a budget, not a spec-sheet line
Read Meta's launch post past the headline and the interesting part is the accounting. At full precision the model "would require over 55 GB of memory". Quantised to roughly four bits, the language model lands "under 20 GB", which Meta says "leaves enough headroom for the model's working memory (its 'KV cache'), the perception encoder... and the speculative decoding drafter to run simultaneously within a 24 GB or 32 GB envelope."
Headroom is the operative word. Your context window is whatever is left on the card once the weights have taken their cut. Every token of company context you feed the thing, a price list, a scope-of-works template, six months of job notes, is physical memory you took away from something else.
None of that is your problem on a cloud frontier model. You get a million-token window, you shovel a disorganised business into it, and you pay the difference in dollars. Crude, but it works, and for plenty of SMBs it has quietly been the plan: skip the filing, rent a bigger bucket.
Put the same agent on a card under your desk and the bucket is fixed at whatever you paid for. The mess now has a hardware price.
The first thing practitioners argued about was filing
The Hacker News thread hit 1,182 points in a day. Plenty of it is the usual release-day scrap over benchmarks and how it stacks up against Qwen and Gemma. The complaint that kept surfacing underneath all that was the context budget.
One commenter wrote off the 128K window as a complete non-starter for his team, because their average context runs well above it. Another called it "the biggest apparent limitation of this model". A third had already rope/yarn-scaled it to 256K and reported running 216K tokens deep on a single Radeon R9700, which is a lovely hack and also tells you precisely which problem he was solving.
Day one of the free local agent era, and the people planning to actually run the thing were arguing about how much of their business fits in the window.
Redis published its State of Context Engineering 2026 this month with the numbers that go alongside that: 97% of leaders believe context decides whether their agents work, and 4% have built for it. Another 73% say their agents fail more often from broken context than from broken models, and 81% are still sitting in the two earliest maturity stages. The finding I keep coming back to is that 83% think fresh context matters more than adding model parameters, which is a funny thing to believe in the same week everyone queued up to download more parameters.
How much company context fits in a 128K context window
128,000 tokens is around 96,000 words. Roughly a 300-page book.
So: could a competent stranger run one day of your business from a 300-page book you handed them? Not the whole business. One day. Quoting, scheduling, knowing what to charge, knowing which supplier to ring and what your margin floor is before you walk away from a job.
Most SMBs I have sat with do not have 300 pages of decision-bearing knowledge. They have maybe forty pages of it. What they have is 40,000 files. The pricing logic is real and it is perfectly legible to the person who invented it, it just lives across a spreadsheet, three email threads and one estimator's memory.
Which is why I think the small window is good news. It sends you looking for the forty pages that decide things and gets you to put them somewhere an agent can walk through. That is work worth doing whether or not you ever run a model on your own hardware.
What I find persuasive is that the size a human can keep true and the size that fits on the card come out about the same. Keeping a brain honest means holding the canon small enough that one person can reread all of it in an afternoon, which is an argument about ownership and staleness and has nothing to do with silicon. Meta's VRAM accounting arrives at roughly the same size from the other end. When a discipline problem and a hardware constraint agree on a number, I stop treating that number as a matter of taste.
Curated context beats complete context, even when you can afford both
The instinct when you hit a context limit is to fight the limit: scale the rope, buy a bigger card, wait for the local model with the million-token window. Worth knowing that the load-everything approach loses even when you can pay for it.
Oracle published research on agent memory the same day Glimmer landed. Their curated memory system scored 93.8% on LongMemEval while holding about 1,300 input tokens per request at turn 80. The comparison was flat conversation history: the full transcript, nothing dropped, about 13,900 tokens per request. Curation won 48, lost 13, tied 19. Their line for it is "bounded, relevant context beat complete, unfocused context".
Ten times fewer tokens and better answers, against a baseline that could see everything. That is the case for a company brain in one sentence, and it holds whether your window is 128K or a million.
Running AI agents locally for a small business is a filing problem
We describe AI capability to clients in four parts: what the system can know, do, reach and remember.
Glimmer, plus NVIDIA's local-agent push the following day, hands a small business do and reach for the price of a graphics card. Tool calling, multi-step execution, picking itself up after a failed step, all of it private and offline. That used to be the expensive half.
It hands you nothing on know and remember. Look at NVIDIA's own list of what this is for. "Private data processing: Read, summarize and act on local files, documents, emails and messages." "Long-running workflows: Break larger projects into steps, track progress and resume interrupted sessions with context intact."
That is a company-brain spec with no brain supplied.
Meta says it more plainly than anyone. An agent that "manages your schedule, drafts your messages, organizes your files, and learns how you work needs deep access to personal context." Agreed. And then the post ships the model, because the model is the part Meta can ship.
We talk with clients about five levels of AI maturity. What happened this week is that the middle levels got cheap and the top one, an agent working with your whole organisation's context behind it, didn't move at all.
Which is the same reason a model you downloaded is still a tool you bought before mapping the work. Free and local doesn't excuse you from knowing which job it is for.
What an agent-navigable business looks like
The signage manufacturer we work with is the cleanest example I have. Apex's decision-bearing context is small: how a job gets priced, what suppliers charge this month, the rules an estimator applies when a drawing is ambiguous, which enquiries are worth quoting at all. Written down properly, that is a short, structured set of documents an agent can navigate.
It did not start that way. It started as years of files, a spreadsheet with a tab per product type, and pricing rules that lived in one person's head. Most of our work with them has been converting the second thing into the first. None of it was about models. All of it is what makes a model useful, and it would fit inside a 128K window with room left over.
That is my answer to a question I get asked in different words every month: what do you need before an AI agent can use your company files? Not a platform. A short, current, owned set of documents that say how decisions get made, and somebody whose job it is to keep them true.
Where I think this gets oversold
Two things cut against the local story, and I would rather say them than have you find out in month three.
Hardware first. The Register measured about 12.2 tok/s on a DGX Spark, and reckons that without a dedicated GPU with enough memory "the best you can expect is around 6 to 14 tok/s". That is a model you can watch typing. Most Australian SMBs do not have an RTX 5090 under a desk and I am not going to tell a plumbing business to buy one.
Second, the cloud frontier models are still better at hard reasoning, and for most owners the privacy case for local weights is weaker than it feels. The enterprise agreement you already signed probably covers what you are actually worried about. Open weight model versus cloud AI is a real decision for a clinic sitting on health records; for most trade businesses it is a preference, not a requirement.
So this isn't a "go local" piece. The point is that the work you would do to make a local agent useful is the same work that makes a cloud agent useful, and it is the only part of the stack nobody can hand you for free. The Register reads Glimmer as positioned for "small-to-medium sized enterprises or enthusiasts". Fair enough. The ones it will work for are the ones who did the filing.
Where I would start this week
- Pick the workflow where the decisions are worth the most money, not the one that annoys you most.
- Write down the rules that govern it. Real rules, current numbers, one place.
- Give it an owner and a review date. A brain that has gone stale is worse than no brain, because people trust it.
- Then point whatever model you like at it, cloud or local, and watch where it gets confused. That is your next page to write.
Do that and the next open-weights release is a genuine upgrade for you. Skip it and you will have downloaded a very fast machine for reading a mess, decided local models are not there yet, and been wrong about why.
If you want a second opinion on which forty pages actually matter, that is what our discovery call is for.
FAQ

Angus McDonald
Founder, Azgard
Builds and operates production AI systems for organisations that need results, not slide decks.