Models & research · Model-card read, 07-15/16
From the Jul 17, 2026 daily brief
Practitioners reading Thinking Machines' official model card for Inkling flagged a rare set of choices, most notably dropping RoPE — the positional-encoding standard nearly every large model shares, the mechanism by which a model knows which word comes before which — for relative position biases (AINews, 07-15); the same week, Kimi K3 shipped its in-house KDA attention — attention being the mechanism that decides which words the model weighs against each other, KDA short for Kimi Delta Attention — with an official claim of up to 6.3× faster decoding on very long inputs. Two labs daring to leave the standard recipe at large training scale in the same week is a signal of architectural diversification for the back half of 2026. The boundary: everything on Inkling comes from community readings of a single model card — one original source; and its "not distilled" claim — distillation meaning training a smaller model on a larger model's outputs — was self-corrected within the same discussion thread. Not a settled fact.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…