Model watch · This week (released August 28–29)
From the Aug 30, 2026 daily brief
Our August 28 issue covered GLM-5.3-Flash's pricing and its efficiency split; what is new today is the open-weight flagship, alongside two releases from other labs the same week. Specs come as "total parameters / active parameters," and all three carry a 1 million context. Z.ai's GLM-5.3 is 744B/40B (Z.ai; specs supplied by the open-source inference framework vLLM). Tencent Hunyuan's Hy4-preview is 770B/49B (Tencent Hunyuan). Alibaba's Qwen3.8-Flash is 125B/6B, priced at US$0.15 per million tokens of input and US$0.47 of output (Alibaba Qwen). Those two figures have to be quoted together: models like this wake only a fraction of the network for each token, so total parameters decide how much memory you provision and active parameters decide how much computation you pay for per token. Quote one alone and you mislead.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…