SecondSourceJudgment rebuilt from primary sources
Deep dive · Jul 17, 2026

Three years ago BCG ran that 758-person experiment

The receipts 4 sources

From the Jul 17, 2026 daily brief

Three years ago BCG ran that 758-person experiment; this March it landed in a top management journal — and the deep dive we finalized last night re-audits it against three years of evidence since. The original design randomly split consultants into groups with and without GPT-4 access: the bottom half improved quality by 43%, the top half by 17%, and the top-bottom gap shrank from 22 percentage points to 4 (Organization Science, published 2026-03); the same directional result has now passed peer review in replications across four domains — writing, customer service, legal, and creative writing. The core judgment: the leveling is real, but only on tasks with a quality template — where the template disappears, the effect reverses. The hardest counterexample comes from Kenya: 640 small-business owners used GPT-4 as a business advisor for months; split by pre-experiment performance, the high performers gained +15% in profit while the low performers lost −8% — the gap widened (Management Science). The difference wasn't in answer quality; it was in which advice they picked and how far they executed — judgment did not get leveled. The third line runs at the market layer: large-scale US payroll data shows that in the occupations most exposed to AI, relative employment of 22–25-year-old entry-level workers fell 16% while seniors in the same occupations held steady (Stanford Digital Economy Lab, 2025-11). The report's unifying read: leveling = commoditization — pay in the leveled band drifts toward the price of a tool subscription, and the human premium migrates to verification, judgment, and accountability. That last step is an inference, not verified causality; keep that firmly in mind. One footnote: the most-cited study arguing "AI benefits top performers most" — in May 2025, MIT formally stated it had no confidence in the data and asked for the paper's withdrawal (TechCrunch, 2025-05).

Skill leveler (born 2023, peer-reviewed 2026)on tasks with a quality template, AI pulls the weakest toward the template — replicated across writing, support, legal, creative
Fight over the distribution's shape (open)what is leveled work still worth — commodity repricing, escalator, or kingmaker?

Open ?

Path Aescalator — everyone rises in step (the only RCT directly measuring seniors, METR, found −19%, n=16, one domain; thin evidence, but zero support)
Path Bkingmaker — a few AI power users take all (its best-known empirical anchor, Toner-Rodgers, disavowed by MIT; only selection-biased market observations remain, no clean evidence)

What would prove this wrong (our self-set 12-month window; verdict date 2027-07-17): one, a clean leveling replication appears in a domain with no template, which would force us to revise "the judgment layer doesn't level"; two, if entry-level decline of the same magnitude shows up in occupations where AI does the work for you, the commoditization reading is overturned; three, the "few winners take all" script reopens should large-sample wage data, controlling for tenure, still show the pay premium of AI power users widening.

The above is the condensed version — the full deep dive goes out tonight at 7:30 US Central, to this same inbox.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section