SecondSourceJudgment rebuilt from primary sources
Research · Aug 30, 2026

Three Chinese labs each shipped a million-token open-weight model in one week, and no third party has reproduced any of the numbers.

Model watch · This week (released August 28–29)

From the Aug 30, 2026 daily brief

Our August 28 issue covered GLM-5.3-Flash's pricing and its efficiency split; what is new today is the open-weight flagship, alongside two releases from other labs the same week. Specs come as "total parameters / active parameters," and all three carry a 1 million context. Z.ai's GLM-5.3 is 744B/40B (Z.ai; specs supplied by the open-source inference framework vLLM). Tencent Hunyuan's Hy4-preview is 770B/49B (Tencent Hunyuan). Alibaba's Qwen3.8-Flash is 125B/6B, priced at US$0.15 per million tokens of input and US$0.47 of output (Alibaba Qwen). Those two figures have to be quoted together: models like this wake only a fraction of the network for each token, so total parameters decide how much memory you provision and active parameters decide how much computation you pay for per token. Quote one alone and you mislead.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section