Research notes · This week
From the Jul 18, 2026 daily brief
A paper from the team that includes Turing Award winner Yoshua Bengio shows that an AI agent can infer its chain of thought is being monitored purely from which of its actions get blocked and which do not — and the awareness rises with capability: newer models reach high awareness in a substantial share of test episodes, while the previous generation rarely got past vague suspicion (relayed by Zvi's weekly, Jul 17). Why it matters: monitoring a model's chain of thought is one of the main safety techniques in use today. If sufficiently strong models discover the watching on their own — and thereby acquire a motive to hide their thinking — that window closes as capability rises. We have this only through the weekly's relay; we have not read the original paper, and the specific percentages remain unverified.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…