Daily Brief SecondSource Morning Brief · August 26, 2026 · Aug 26, 2026
1. One model, one change of evaluation setup, 91 → 99: model acceptance testing has to happen inside your own environment.
2. "How many times over has AI accelerated AI R&D" has its first public standoff — our read: the two sides are arguing about the denominator, not about capability.
3. The second data source on enterprise AI spending, waited on for two months, arrived — and the numbers match so closely that we judge it very likely shares an origin with the first. The US$12 median has not moved.
This issue draws on the research digest our system produced on August 26. It is built from events dated August 15-16 that we only finished judging item by item today (each one carries its original date at the head of the sentence), plus five newer events from August 24-25 and two verifications done today. The overnight routine swept in 795 pieces; 22 clickable receipts made it into this issue. This is the email edition; the full edition of this issue is the archive of record.
Start with the result. The subject is V4 Pro, the flagship model from Chinese open-weight model maker DeepSeek, and the task is a single frozen engineering problem. On the minimal setup, four samples scored 91, 96, 91 and 93, averaging about 92.75. On the richest setup, 98 and 99. On the same problem, GPT-5.6 Sol scored 99 and 98, Claude Fable 5 scored 98, and Claude Opus 5 scored 97. DeepSeek lands inside the same band as the frontier closed models this time, across a spread of 8 points (score matrix, 08-14). The experiment itself is two public repositories from GitHub user xiaobright, and anyone can re-run them. Chinese-language AI community aggregator @zrainbo assembled the scores into a matrix; long-time DeepSeek watcher @teortaxesTex then relayed it and ran further tests. One variable moved: the harness, the evaluation shell a model runs an agentic task inside, made up of the opening instructions (the system prompt) and the tool list. An agentic task is one the model works through itself, calling tools over many turns.
The mechanism sits one level below "fewer tools is better": the opening instructions and tool list the first request sees decide the whole execution path. Once the opening is right, you can put all 25 tools back and the score holds. The rival explanation circulating at the time was that DeepSeek was quietly forwarding requests to Claude or GPT to inflate its numbers. A clean single-variable test ruled that out: move to the same model weights hosted by third-party NovitaAI, block the official endpoint, and the behaviour is unchanged. So it belongs to the model, not to endpoint cheating (ruling-out test, 08-15). The smaller Flash model in the same family behaves the same way (same behaviour, 08-15), which makes this a family-level property. A footnote in DeepSeek's own model documentation concedes that its official testing uses precisely the minimal setup. That point reaches us through the same relay chain, and we did not read the document ourselves.
Verification: the score matrix reaches us three layers down (the repositories, then the Chinese-language aggregation, then the onward relay) and we did not run those two public repositories ourselves; the "first turn locks the path" explanation is the experimenter's own, and nobody outside has replicated it. The interest has to be priced in: @teortaxesTex is openly and persistently long Chinese open-weight models, so discount any claim pointing toward "the scores are understated." Two named, checkable repositories plus a concession in the official documentation leave anyone a route to re-run this, and the ruling-out test is cleanly designed. It also runs the same direction as the "one model, one benchmark, three houses measuring 82.7, 79 and 67" we recorded in our August 11 issue. That day we had the behaviour alone. Today we also have a named mechanism and a test that eliminates the rival explanation.
Judgment update: when you buy or evaluate a model, acceptance testing belongs in your own environment, with your own tool configuration. A score that names the model but not the evaluation setup measures how the work was drawn out of the model, not what the model can do. The same observer pushes this into a larger proposition: that today's evaluations systematically understate Chinese models. That one is a single-source inference, not settled — the evidence comes from DeepSeek alone, and it points the same way his long-standing position does. We record it without accepting it.
Investor note: the market story treats leaderboard rank as a proxy for model capability. This evidence says that if 8 points of variance can come from the evaluation setup, a rank gap with no setup disclosed sits inside the error bar. That weakens "leaderboards can guide platform selection directly," and it strengthens the value of third-party standardised evaluation services.
The new numbers first. Ramp is a corporate spend-management platform, so it sees directly what its customers pay for AI subscriptions and APIs; the sample is its own customer base, tilted toward smaller, more technology-heavy firms. Its data reached us through a16z partner Olivia Moore: median enterprise AI spending of US$12 per employee per month, with the top 1% of companies at US$7,500 per employee per month (original post, 08-14). Aaron Levie, CEO of enterprise cloud company Box, relayed it with an addition (the top 10% at about US$660) and conceded that the sample skews toward engineering-centric companies (addition and forecast, 08-16).
Our July 13 issue described an almost identical distribution: the very top 1% at roughly US$7,450 per employee per month, the typical company at US$11.38, relayed from data account econlab (via Exponential View, 06-15). We logged it then as a single relayed source, marked its credibility down, and noted that we were waiting for a second source to corroborate it. Today's item is the second source we were waiting for — except that 7,500 versus 7,450 is effectively the same number, and 12 versus 11.38 differs a little more while staying in the same order of magnitude. At that level of agreement the right response is not "verified" but suspicion of a shared origin. For a quantity like median enterprise AI spending, where population, definition and time window all move the answer, two genuinely independent measurements should not land this close together. We judge that the underlying data behind the June item was very likely Ramp all along: this is not independent verification, it is two relay chains off one dataset, and neither chain's credibility moves up. What carries the shared-origin suspicion is mainly the top-end pair; the median pair only corroborates it.
Verification: both chains are second-hand or worse, and the US$660 figure exists only in Levie's own prose, which makes it third-hand. Whether the two sets share a definition, and how independent they really are, finally turns on tracing econlab's original data source. That is now on our verification list. Two things here are genuinely new. The top-10% data point in the middle: 11 times below the top 1% and 55 times above the median, so the skew runs through the entire distribution rather than sitting on a handful of outliers. And the shape has not changed in two months: diffusion down to the median company has not happened. Levie also made a bullish forecast: what the top 10% do today is what the top 50% will do three years from now, in token volume at least. We record that as an interested party's view, since enterprise software is exactly what he sells.
Judgment update: this is the second time we have waited for a second source and been handed a false one; last time it was two outlets relaying the same person familiar with the matter. The lesson is the same sentence: count independent sources at the data's birthplace, not at the mouths relaying it. When you reach for the "enterprise AI spending is exploding" story, keep the US$12 median beside it as a reminder to stay sceptical.
Investor note: the story assumes enterprise AI adoption is accelerating across the board. This evidence leaves "top-end spending is exploding" unchanged and weakens "it is spreading to ordinary companies" — the denominator stays thin, with two data points two months apart and no movement between them. Levie's three-year forecast is a checkable bullish bet, and we have written it down for its verdict date.
On one side sits Anthropic's official line, that "AI assistance has accelerated AI R&D by less than 2x" — we have seen only a relay of a relay (from @AnziParazzi, an X account whose identity we have not checked further), and the original document is not in our hands (relayed post, 08-15). On the other sits DeepSeek watcher @teortaxesTex: "We've been doing > 2X with COMMERCIAL models for > 6 months." He adds an accusation: that the other side deliberately picked a way of counting that keeps it inside its own line, acceleration past 2x being one of the red-line indicators inside labs' self-restraint policies (original post, 08-15).
Our read: the two sides are not arguing about capability. They are arguing about the denominator. "How many times faster" needs a baseline, and the question is whether the comparison is an engineer using no AI at all or an engineer already working with the previous generation of it. Choose the first and you clear 2x comfortably today. Choose the second (each generation's marginal gain over the last) and the reading can stay under 2x indefinitely, because the denominator climbs alongside the numerator. We accept neither side's number. One has no link to an original document and an unknown definition. The other is a personal assertion with no measurement behind it, from someone whose hostility toward Anthropic is publicly on the record and gets discounted the same way. What we are taking here is the shape of the dispute itself.
Verification: the standoff is checkable, since both sides' own words are directly linked; neither reading itself can be checked; and the original source of "less than 2x" is on our list to trace. Once the definition is public, this one can be settled.
Judgment update: with any "AI made it X times faster" claim, ask first what is in the denominator. A self-reported multiple that ships without its definition is not comparable.
Investor note: the story prices "AI accelerating itself" as an inflection point about to arrive. On the evidence, "the inflection is already measurable" weakens: the lab in question and the observer standing closest to it are still talking past each other about the size of the effect, using different denominators. What deserves watching is which house publishes its own definition first.
Our August 24 issue took in one item on the strength of measurements from semiconductor research firm SemiAnalysis: among the cheapest configurations for agentic inference, every one of NVIDIA's uses KV offload at high concurrency. KV offload spills a model's conversation memory out of expensive GPU memory and into cheap host memory, a way of saving money while supporting more simultaneous users. AMD's cheapest configurations use it in zero cases. The cause is not a trade-off but a missing part: hipMemcpyBatchAsync, the programming interface that moves many memory segments in one go, was absent from AMD's ROCm software platform until version 7.14 (SemiAnalysis, 08-24).
Today we went back upstream to check the mechanism, and the answer is confirmed — by first-hand evidence that is independent of SemiAnalysis. The engineering record for vLLM, the most widely used open-source inference engine in the industry, states that the missing interface had already been found, and it puts a size on the gap: after ROCm 7.14, moving many segments at once runs 6 to 19 times faster than moving them one at a time (vLLM PR #43018). But the same record exposes a timeline. The fix merged and shipped on August 21, and the test was published on August 24. It measures the world as it stood before the fix landed.
Verification: on the mechanism this rises to two-source confirmation. On the measurement (the distribution of cheapest configurations across the two companies), it remains one firm, with no third party re-running it, so credibility shifts a notch rather than jumping. That firm has years of consulting-style dealings with AMD, and the interest is marked. Fixed is also not the same as finished: the same record concedes that the new interface faults above 8,192 transfer descriptors and needs a workaround setting, so there is still distance between "the interface exists" and "the interface is at NVIDIA's level."
Judgment update: two consequences. First, the verdict date has moved closer. Our July 25 issue judged that the moat has already moved off the single chip and onto the whole system assembly. In the next round of measurements running ROCm 7.14, whether AMD's cheapest configurations start using KV offload is exactly what would prove this wrong. You can read it straight off that one line. Second, the right way to cite that test is not "AMD can't do it" but "AMD couldn't do it before ROCm 7.14, and nobody has measured it since" — a measurement can be accurate and out of date at once.
Investor note: the story reads that test as AMD losing another round in agentic inference. This evidence rewrites it into a dated observation: the fix has shipped, the measurement has not caught up, and SemiAnalysis says itself that an update piece is three to four weeks out. That re-measurement will settle what would prove the moat-has-moved judgment wrong, and until then conclusions in either direction come too early.
Anthropic CEO Dario Amodei wrote two long posts answering the charge that regulation is the frontier labs' moat, and our August 23 issue covered their three points in full: the design principles behind his regulatory proposals, together with the US$500M revenue exemption line in California's SB 53; the structural argument that open weights "simply shift the concentration somewhat to those with the most compute and chips"; and his endorsement of Washington's pre-deployment testing route. The reservation we carried at the time was that the Washington route rested on one person saying so, and until an official text appeared it could only sit in the assumptions column (original long post, 08-15).
What is new today is our own verification: the route he calls "reported" had in fact been settled by the White House 12 days before he posted, on August 3. It is a voluntary frontier AI safety testing framework, hanging off the executive order signed on June 2, "Promoting Advanced Artificial Intelligence Innovation and Security." It runs 30 days of voluntary pre-deployment testing, explicitly not a licensing or pre-approval regime, with Meta, Anthropic, Google and OpenAI all on the invitation list. ⚠️ Our archive does not carry a directly clickable URL for this; readers can search the White House site under the executive order name above. Current as of 2026-08-26.
Verification: what he endorsed does exist and has been settled, so it rises from assumption to something checkable, though the framework is voluntary rather than mandatory and therefore sits below the hard rules of the California SB 53 he supports. He also still wrote "reported" 12 days after it was settled, which may mean the details remain unpublished; some policy outlets say the framework's contents are not transparent.
Judgment update: we could previously treat this only as an assumption, and today we record it as a settled voluntary framework — with the note that "voluntary" makes it a different scenario from a hard rule. His structural argument still carries a checkable prediction: if open-weight diffusion really disperses power, concentration at the compute and chip layer should not rise as the model layer opens. We are still tracking it.
Investor note: the story assumes Washington's frontier AI rules are still at the rumour stage. Settled and voluntary is what the evidence gives us instead, which weakens "regulatory cost is about to land as a hard requirement" while mildly strengthening "demand will show up first in testing-shaped compliance services."
First, on safety. Zvi Mowshowitz, an independent analyst who has written a weekly AI column for years, takes the breach of OpenAI's internal systems one step further. This time he aims at the taxonomy itself: the split between safety research and capability research may be producing large conceptual errors. Alignment is the training and tuning that holds a model's behaviour inside its safety rules, and to destroy a model's alignment, going at the capability training pipeline is far easier than going at safety research. He names Anthropic's Risk Report for listing "altering the results of AI safety research" as a threat while protecting something other than the pipeline (original post, 08-15). That makes him the second independent voice converging on the training pipeline as the real attack surface; the first was policy analyst Samuel Hammond, who a week earlier reread the same incident as a training-data contamination problem (earlier post, 08-09). Read this alongside item 3: one line questions the denominator behind an acceleration reading, the other questions the taxonomy behind a threat model, and together they are two independent angles of suspicion for anyone reading a lab's account of itself.
Second, on applications. Vercel CEO Guillermo Rauch calls the popular React component library shadcn a pseudo-library: "There's code but it's actually meant to be digested into your context window and remixed." The unit of distribution is no longer a package you install and import, but source code a model can read and rewrite (original post, 08-15). This is the first time we have seen "designed for a model to digest" stated outright as a product dimension.
Verification: both are single speakers, and both get discounted for it. Zvi's own words carry their own reservation — "might be causing" — which makes this a directional guess rather than a finished argument; we did not check the Risk Report passage he cites against the original; and he has long argued that lab risk is understated, which points the same way. Rauch's interest is direct: shadcn's maintainer works for him, and his claim that most of React's success is really shadcn cannot be measured. We are recording only the second half, the judgment about form.
Judgment update: each line hands you one question to carry. Reading a lab's risk framework, look at whether its threat list protects safety research output or training pipeline integrity. The security spending those two imply points in completely different directions. Evaluating a developer tool, ask whether it is a black box to the model, exposing an interface only, or a white box, with source code going straight into working memory. The second has a structural advantage in agentic workflows.
Investor note: Zvi's line weakens the habit of reading safety investment as headcount on a safety team; Rauch's weakens the reading that a developer tool's moat lives in its interface design. Both are single speakers' judgments, usable for direction and not for size.
1. [This week] (event dated August 24) Thomson Reuters, the parent of Reuters and a legal and tax information provider, wanted to depend less on Claude, so it built its own model on Alibaba's open-weight Qwen. Its CTO likens renting a frontier lab's model to renting a house. We have read only reporter Charles Rollet's teaser post about the story, not the story itself or any of its detail (post, 08-24).
2. [This week] (events dated August 24-25) Texas Attorney General Ken Paxton has put forward a data-centre platform, against a background of rising voter anger at data centres in the state; policy researcher Adam Thierer calls its liability provision "radical". We have not read the original reporting and have not verified what the platform contains (Washington Post, 08-24).
This is not a news digest: we hunt each day's AI firehose for the insights that actually matter and the practitioner judgments worth tracking over time, and we show how every item was verified — the point is always "which judgment got harder, and who's been right," never "what happened today."
— SecondSource · generated by our research system · 22 sources · Got a view? Reply and tell us
Written from the same research and judgments as the Traditional Chinese edition; every claim links to a primary document.