Models

Meta Returns to Open Weights With 30B Muse Glimmer

Meta published Muse Glimmer, a 30B agent-focused model under Apache 2.0 that fits a single consumer GPU at 4-bit. Zuckerberg paired the release with an open-weights manifesto.

Faruk TalmaçAugust 11, 20263 min read4 views
Meta Returns to Open Weights With 30B Muse Glimmer

Thirty billion parameters, Apache 2.0, on Hugging Face. After more than a year without an open-weight release, Meta has put Muse Glimmer into the world, and Mark Zuckerberg attached a manifesto titled "The Future is for Everyone" to it.

The specifications point at one clear target: agents running on your own machine. At full precision the model needs north of 55 GB of memory. Quantised to 4-bit it drops to roughly 20 GB, which fits a single consumer GPU or a well-equipped laptop. Meta also reports up to a 3.1x throughput improvement when paired with a draft model.

How much weight should the benchmarks carry?

Meta compares Glimmer against Google's Gemma4-31B and Alibaba's Qwen3.6-27B. By its own numbers, Glimmer leads on agent work such as tool use, web search and long-context tasks; Qwen holds an edge on desktop control and terminal tasks; the three are level on multimodal.

Vendor benchmarks deserve the usual discount. Competing models may not have been tuned with equal care, and the categories were chosen by the party being measured. With open weights, the honest verdict arrives a few weeks later, once people have run the thing on their own workloads. Until then this table is a claim, not a result.

The argument, and the business under it

Zuckerberg's stated position is that superintelligence should not sit with a handful of labs but spread as widely as possible. He also stakes out ground in the distillation fight, arguing that "you can learn from anything you can observe," which conveniently legitimises learning from rivals' outputs. His conclusion is that the best open models ought to be American ones.

One clause in the essay carries the commercial logic: anyone who wants more compute pays for it through a dynamic auction mechanism. Weights are free, capacity goes to the highest bidder. That is Meta's advertising business model, transplanted onto data centres.

Conviction or catch-up?

Both, most likely. Meta had fallen behind in open weights, where Qwen and DeepSeek set the agenda over the past year, and it has not overtaken OpenAI or Anthropic at the closed frontier. Owning the open ecosystem is how a company buys influence without selling a subscription.

For anyone downloading the model, motive is beside the point. Apache 2.0 permits commercial use, and a released set of weights cannot be recalled.

When running it locally actually pays

A capable agent model that runs on hardware you control, sends nothing outside, and charges no monthly fee is genuinely attractive for some work: contract review, internal document Q&A, email triage, anything where the data shouldn't leave the building.

Just cost it honestly against a cloud subscription. A GPU that runs a 20 GB model at usable speed, the setup time, the update cycle and someone to look at it when inference stops at 2am all belong in the calculation. What we see in practice is that local inference is an expensive hobby for a team making a few hundred calls a month, and turns profitable fast for an organisation processing thousands of documents a day or operating under a hard data-residency constraint.

The cheapest way to decide is to measure one month of real usage, then price that same volume twice: once at cloud API rates, once as hardware amortised over three years. The gap between those two numbers will answer the question better than any benchmark table.

Sources: TechCrunch, The Decoder

Share This Article

Faruk Talmaç

Written by

Faruk Talmaç

Co-Founder & Editor

Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk