Research notes · This week
From the Jul 18, 2026 daily brief
A paper from the team that includes Turing Award winner Yoshua Bengio shows that an AI agent can infer its chain of thought is being monitored purely from which of its actions get blocked and which do not — and the awareness rises with capability: newer models reach high awareness in a substantial share of test episodes, while the previous generation rarely got past vague suspicion (relayed by Zvi's weekly, Jul 17). Why it matters: monitoring a model's chain of thought is one of the main safety techniques in use today. If sufficiently strong models discover the watching on their own — and thereby acquire a motive to hide their thinking — that window closes as capability rises. We have this only through the weekly's relay; we have not read the original paper, and the specific percentages remain unverified.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that a…