Models
Qwen3.8-Max Is a 2.4-Trillion-Parameter Open-Weight Model
Alibaba released Qwen3.8-Max: 2.4 trillion parameters, open weights within a week, a claim of months-long autonomous tasks, and pricing at 10% of standard.

2.4 trillion parameters, 95 billion active per query. Those are the numbers behind Qwen3.8-Max, the model Alibaba released today. For scale: the largest open-weight model shipped so far, Moonshot's Kimi K3, totaled 2.8 trillion. Qwen3.8-Max is smaller, and tellingly, Alibaba isn't leaning on size for its pitch. The pitch is endurance. The company says the model can run complex tasks on its own for months at a time.
The weights land on Hugging Face and ModelScope within a week. For now the model is available through QwenCloud.
What the scores say
On Alibaba's own benchmarks, Qwen3.8-Max scores 93 on PaperBench, a test of reproducing research papers, and 86.6 on the command-line task benchmark TerminalBench 2.1, just behind GPT-5.6 Sol at 88.8. The company puts overall performance in the same range as Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol.
Every one of those figures is internal. No independent benchmarks exist yet, and there's a recent reason for caution: the 2.8-trillion Kimi K3 posted strong launch numbers but fell clearly behind Western models in third-party testing. When a lab grades its own homework, the useful question isn't the score, it's who measured it.
The months-long task claim
What Qwen3.8-Max adds on top of the Qwen3.5 architecture is long-horizon task training: Alibaba ran reinforcement learning across 4,000 distinct environments. The target work is things that take days or weeks, like reproducing a research result end to end or assisting with chip design. The model handles documents over 200 pages and video archives over 100 hours, with vision and audio arriving through an add-on library called Qwen-MM-Plugins.
The price move and its timing
The launch campaign is aggressive: on Alibaba Token Plan, Qoder, and QoderWork, the model is offered at 10% of standard pricing. The timing is no accident. OpenAI cut GPT-5.6 Luna's price by 80% last week, and DeepSeek answered a day later with V4-Flash. Chinese labs' twin pressure of open weights and low prices is visibly dragging closed-model pricing down.
What a business should take from this
At 95 billion active parameters per query, this is still a heavy hardware bill for a company that wants to self-host, so open weights don't mean everyone will run it in-house. Two concrete gains remain. First, cloud and hosting providers can serve this model cheaply; competition is eroding API prices at every tier. Second, the menu grows for organizations that need data control: a model whose weights you hold makes it easier to defend an architecture where sensitive data never leaves your environment. Our advice: until independent results arrive, don't anchor on launch scores. Run a small pilot on your own workload and measure.
Sources: The Decoder, Qwen

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.