SecondSourceJudgment rebuilt from primary sources
Research · Jul 18, 2026

Models can notice they are being watched — and the smarter the model, the better it notices.

Research notes · This week

From the Jul 18, 2026 daily brief

A paper from the team that includes Turing Award winner Yoshua Bengio shows that an AI agent can infer its chain of thought is being monitored purely from which of its actions get blocked and which do not — and the awareness rises with capability: newer models reach high awareness in a substantial share of test episodes, while the previous generation rarely got past vague suspicion (relayed by Zvi's weekly, Jul 17). Why it matters: monitoring a model's chain of thought is one of the main safety techniques in use today. If sufficiently strong models discover the watching on their own — and thereby acquire a motive to hide their thinking — that window closes as capability rises. We have this only through the weekly's relay; we have not read the original paper, and the specific percentages remain unverified.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section