SecondSourceJudgment rebuilt from primary sources
Research · Jul 17, 2026

Two new models deviate from "standard attention" in the same week — a leading indicator at the architecture layer.

Models & research · Model-card read, 07-15/16

From the Jul 17, 2026 daily brief

Practitioners reading Thinking Machines' official model card for Inkling flagged a rare set of choices, most notably dropping RoPE — the positional-encoding standard nearly every large model shares, the mechanism by which a model knows which word comes before which — for relative position biases (AINews, 07-15); the same week, Kimi K3 shipped its in-house KDA attention — attention being the mechanism that decides which words the model weighs against each other, KDA short for Kimi Delta Attention — with an official claim of up to 6.3× faster decoding on very long inputs. Two labs daring to leave the standard recipe at large training scale in the same week is a signal of architectural diversification for the back half of 2026. The boundary: everything on Inkling comes from community readings of a single model card — one original source; and its "not distilled" claim — distillation meaning training a smaller model on a larger model's outputs — was self-corrected within the same discussion thread. Not a settled fact.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section