SecondSourceAI Industry Insight · Full Archive

Daily Brief SecondSource Morning Brief · October 8, 2026 · Oct 8, 2026

How we checked each item today

For readers who want to dig in. If you only want what happened, go back to Today.

This issue draws on our internal research brief from early October 8, the deep dive published in the small hours the same day, and last night's sweep. Main-line item 1 rechecks an experiment Harvey published on September 8 and cites a paper from March; items 2 and 3 are company announcements and posts by executives and vendors from October 6 and 7; the material in the later columns dates from September 6 to October 7. We swept 397 pieces overnight, and this issue cites 19 outside sources with links you can check.

1. [Evidence update] (deep-dive recheck, October 8) Harvey had the same seven models split each data room among sub-AIs to read, and the average due-diligence pass rate rose from 23.3% to 62.4%

How we checked

The numbers come from Harvey alone; we read the original post and its charts. The data rooms are simulated and the grader is a model, and Harvey doesn't say how closely its grader agrees with lawyers. Scores and read coverage move together, but Harvey only calls them correlated; it doesn't separate how much comes from reading more and how much from the way the work was split and merged. A March paper by four academic researchers found the opposite on other long-context benchmarks: off-the-shelf coding tools beat the best published results on average (Cao et al., 2026-03). Those tasks aren't legal documents, so they can't be compared directly with Harvey's. ⚠️ Our analysis was produced with help from Anthropic's models, and Claude Code and Opus 5 are among the systems under evaluation here.

Back to this item ↑

2. [This week] (event date October 6) Anthropic opens a new access tier allowing authorized offensive testing for security professionals who pass identity verification

How we checked

We read both original posts. We haven't seen Anthropic's full terms: how verification works, what each tier allows, or whether it costs anything. The 40% comes from Cline's own task set; the post names no benchmark and gives no scoring method, and it says nothing about refusal rates through the verified channel. Cline sells a multi-model tool and has a motive to promote open models. ⚠️ Disclosure: our analysis was produced with help from Anthropic's models, and Anthropic is one of the subjects of this item.

Back to this item ↑

3. [This week] (event date October 7) AWS open-sources Strands Box, an AI sandbox that it says enforces its rules at the operating-system and network layer

How we checked

This rests on one AWS executive's post, which we read in the original. It's a preview: how Dogwood rules are written, what the performance cost is, and whether an agent can get around the sandbox: no third party has tested any of these. Srinivas's line only states a direction, that Perplexity wants to control its own sandboxes; it carries no product detail and can't count as the same kind of evidence as Strands Box. We're not writing this item up as a judgment yet.

Back to this item ↑

What you are not getting today. What would most affect judgment is how closely Harvey's model grader agrees with lawyers, which Harvey hasn't published. Beyond that, we haven't read Anthropic's full terms for its cyber verification program, Cline's security-benchmark task set, or Anthropic's official Haiku 5.5 pricing page. We read only part of October 7's social posts this round (the fetch timed out), so item 3's "no third party yet" holds only within what we read.

Sources & accounting

The past 24 hours. 397 new pieces came in from October 7 to 8: 181 academic papers, 116 social-platform posts, 49 other papers, 34 company and personal blog posts, 11 show transcripts and 6 subscription newsletters. Of those 397, we read 182, screened out 213 and did not get to 2. Most of today's material comes from this sweep; before today's deadline we separately read 7 pieces outside the 397 (6 social posts from October 3 and 4, and one show episode from September 23) and used 1: the Srinivas post cited in main-line item 3. To check item 1, we read Harvey's original post and charts directly and looked up the academic paper pointing the other way.

One-time backfill. No new one-time backfill today.

A note on source concentration. ⚠️ Every order-of-magnitude figure in main-line item 1 comes from Harvey alone, and Harvey and the inference platform it worked with both benefit from this approach. ⚠️ Items 2 and 3 and both Product moves items rest on companies' own accounts or company posts. ⚠️ Our analysis was produced with help from Anthropic's models, and items 1 and 2 and the second Product moves item all involve Anthropic directly.

The sources we track. Our long-term roster has 529 named sources: 302 on social platforms, 90 shows, 51 news outlets, 48 blogs, 48 paper authors and 46 newsletters, with the rest spread across earnings, keynotes and other channels. ⚠️ Those are counts of tracked sources, a different population from the 397 new pieces above; they can't be added or compared. Representative names: on social platforms, Simon Willison, Ofir Press, Peter Harrell and Teortaxes; among newsletters, SemiAnalysis, Zvi Mowshowitz and Latent Space; among shows, Cognitive Revolution, ChinaTalk and a16z. This issue uses 19 outside sources in the body. We count only links the body actually cites that are not on our own domain, so it can't be compared with the 529 tracked sources either. Of our tracked sources, 121 are Taiwanese, 7 of them added in September; Taiwanese media, Taiwanese companies' official channels and local Taiwanese reporters count, foreign outlets relaying Taiwanese media do not.