Models

China Ships Two Open-Weight AI Models in a Single Day

GLM-5.3-Flash and Qwen3.8-Flash-Next both shipped open weights on August 26. Ox Alpha's identity, the benchmark tables, and the 15-cent price floor, explained.

Muhammet Fatih BatmanAugust 27, 20263 min read3 views
China Ships Two Open-Weight AI Models in a Single Day

China's open-weight release cadence has reached the point where two frontier-adjacent models can land on the same day and split the spotlight. On August 26, Z.ai released GLM-5.3-Flash and Alibaba's Qwen team released Qwen3.8-Flash-Next, both with downloadable weights, both priced around 15 cents per million input tokens, and both racking up four-digit point totals on Hacker News within hours.

GLM-5.3-Flash was Ox Alpha all along

Z.ai's entry is a 320-billion-parameter mixture-of-experts model that activates 18 billion parameters per token, routing each one through 8 of 288 experts. It is natively multimodal, accepting image and video input, and carries a 1-million-token context window. The weights are on Hugging Face under an MIT license, trained on a 30-trillion-token multimodal corpus. Two architectural choices stand out: a hybrid attention scheme mixing KDA linear attention with sparse MLA layers, and a compression technique called IndexPool that cuts attention compute roughly threefold and shrinks the KV cache by a factor of 4.4, which is much of how a 1M-token context stays affordable. On Z.ai's own numbers, it scores 84.3 on Terminal-Bench 2.1 against 85.0 for Opus 4.8, and jumps to 63.4 on the agentic coding benchmark DeepSWE from its predecessor's 46.2. Vendor benchmarks deserve their usual asterisk; independent runs tend to land a few points lower.

The best detail is the backstory. Bloomberg confirmed that the mystery "Ox Alpha" model that spent the past week running anonymously on OpenRouter and OpenCode was this model, served entirely on domestically produced Chinese chips. The stealth-launch playbook doubled as a hardware demo. Note that the flagship GLM-5.3's own weights remain unreleased; Flash is the smaller sibling that made it out the door first.

Qwen3.8-Flash-Next previews Qwen4

Alibaba's release is smaller and stranger: 125 billion total parameters with just 6 billion active, plus an unusual 51-billion-parameter n-gram embedding layer designed to sit in system RAM instead of GPU memory. The team frames the model as an architecture preview of Qwen4 and claims it beats Qwen3.7-Plus at roughly one-ninth the training cost. Reported coding scores are strong (62.5 on SWE-bench Pro, 91.9 on LiveCodeBench v6, same vendor-table caveat), context runs to 262K tokens natively and 1M with YaRN, and API pricing lands at $0.16 in and $0.47 out. It follows the 27B open-weight release by just ten days.

What the price band means in practice

The headline for businesses is not any single benchmark, it is the floor these two set together: capable multimodal models at $0.15-0.16 per million input tokens, weights included. For high-volume, mid-difficulty work such as document triage, classification, and draft replies, that band rewrites the cost math against Western flagship pricing. Self-hosting remains a serious commitment (GLM-5.3-Flash's FP8 checkpoint weighs 306 GiB and wants an 8-GPU node), so most smaller teams will consume these through APIs. If you do, put the data-residency question in the contract, not in a hope; a cheap token does not relocate your compliance obligations.

Sources: Z.ai, The Decoder, MarkTechPost

Share This Article

Muhammet Fatih Batman

Written by

Muhammet Fatih Batman

Founder & Editor

Founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk