SecondSourceAI Industry Insight · Full Archive

Daily Brief SecondSource Morning Brief · October 11, 2026 · Oct 11, 2026

Fireworks AI says its internal credentials were used without authorization; the company says its investigation isn't over and it has found no customer-data impact yet

Today's front-page cartoon (AI-generated): one street scene with a corner for each of today's stories.

At a glance

1. Fireworks AI's CEO says the company has revoked the misused internal credentials and notified the customers it has confirmed were affected. (Affects: teams running models on Fireworks)

2. Policy researcher Peter Wildeford says no lab has yet had AI do all of its R&D, while using AI to speed R&D up is, in his view, what every major lab is doing. (Affects: anyone hearing vendors talk about self-improvement)

3. A paper reports, by its own numbers, that after a document is turned into a small set of add-on weights, the model's top guess for the next word is right 84.9% of the time, versus 63.4% for the same model with no access to the document; only short documents were tested. (Affects: engineers connecting documents to models)

Today's main line

1. [Today] (event date October 10) Lin Qiao, CEO of inference platform Fireworks AI, discloses misuse of internal credentials; the company has revoked them and notified affected customers

Concept illustration in two side-by-side panels: on the left, a black server rack topped by a white sign lettered FIREWORKS AI with a multicoloured starburst icon has a red card inserted in a black card reader on its left side. On the right, a person in a black turtleneck, wide black trousers and white sneakers stands in front of a pale pink circle holding a red card torn in two, one half in each hand, beside a second rack lettered FIREWORKS AI whose card reader on the left side has no card in its slot.
Lin Qiao, CEO of inference platform Fireworks AI, said on X on October 10 that internal Fireworks credentials were used without authorization. She says the company has revoked those credentials, directly notified customers confirmed to be affected and is investigating with a security firm; so far it has not detected impact on customer data or inference traffic, and the investigation is still under way. Teams on Fireworks can first check whether they got a notice, then check that they can rotate their keys at any time. (AI-generated)

Why this matters to you: If your team runs on Fireworks, first check whether you got a notice, then check that you can rotate your keys at any time.

Read the full item

Fireworks AI hosts open-weight models and runs them for companies that call them through its API. On October 10, CEO Lin Qiao wrote on X: "We identified a security incident involving unauthorized use of internal Fireworks credentials." She says the company has contained the incident, revoked those credentials, directly notified every customer confirmed to be affected, and is investigating with a security firm; so far, she writes, they "have not detected impact or access to customer data or inference traffic." Inference traffic is the requests customers send to a model and the answers the model sends back (Lin Qiao, 2026-10-10). She describes them only as Fireworks' internal credentials. The post doesn't say what kind they were, whether customers' own keys were involved, how they were obtained or what they were used for.

Verification: We read Lin Qiao's post itself; we didn't open and read the security notice it links to. The post gives no count of affected customers, no type of credential and no time window. "Not detected" is the state of an investigation still under way, not a finding; as of early October 11 (US Central time) we found no follow-up from the company. The only source is the company itself, so we aren't forming a new judgment today. When the investigation reports, the thing to check is whether the line "no impact on customer data" gets revised. Teams on other hosted-inference services can put the same two questions to their own provider: if its internal credentials were misused, would it notify customers directly, and can customers rotate their keys at any time?

What to take away today: #1: If your team runs on Fireworks, first check whether you got a notice, then check that you can rotate your keys at any time.

Also happened

1. Not verified by us yet: [Today] (relayed October 10) China tech-policy analyst Paul Triolo relays the Financial Times: "Nvidia is in talks to acquire or deepen its investment in US start-up Reflection AI, a developer of open-weight models that President Donald Trump's administration hopes will rival cheap Chinese alternatives." The same report says Reflection already works with the Pentagon and the US Department of Energy and builds models for allies such as South Korea. Another relay on X says only "acquire," dropping "or deepen its investment"; we haven't read the FT original (Paul Triolo, 2026-10-10; Paul Triolo, 2026-10-10).

2. Not verified by us yet: [Today] (reposted October 9; date of the remarks unknown) OpenAI CEO Sam Altman says he supports open-source models, and also: "We're going to, as a society, I think, have to just accept some fairly severe cyber incidents from open models, in exchange for the liberty that comes with that." Where he said it and the full context are unknown; what we read is a reposted quote, not the original interview (@jaynitx, 2026-10-09; as quoted by Peter Wildeford).

3. Not verified by us yet: [Today] (October 10) A viral X post says Alibaba Cloud stopped offering four Chinese open-model brands, DeepSeek, Kimi K2, GLM and MiniMax, from October 10. A developer corrected it the same day: "Alibaba Cloud is retiring specific older model versions, including many Qwen models," adding that the official docs still list newer versions from all four; Paul Triolo reposted the correction in agreement. We haven't checked the retirement notice line by line; teams running older Qwen versions or these models on Alibaba Cloud should check the versions and dates on the notice themselves (original post; correction).

4. Not verified by us yet: [Today] (reported October 9) According to Bloomberg, Jefferies strategist Chris Wood thinks the most likely long-run outcome of the AI boom is massive capital destruction in the US, with market share shifting to cheaper Chinese open-source models. "Massive" isn't defined. What we read is a relay by China tech analyst Rui Ma; we read neither the full Bloomberg story nor the Jefferies report (Rui Ma, 2026-10-09; Bloomberg).

Chips & semiconductors

1. [Today] (post, October 10) Semiconductor analyst Sravan Kundojjala says the four big chip-equipment makers, ASML, Applied Materials, Lam Research and KLA, are all involved in Terafab, the chip fab Musk's companies plan to build, and puts SpaceX's first-phase volume fab at US$16.8 billion (a single post, unverified). He writes: "ASML, AMAT, Lam and KLA are all engaged in Terafab project." The same post puts a pilot line at another US$3 billion or more, so in the post both figures are SpaceX spending, on different phases. He thinks KLA's inspection and metrology tools matter most in the early push for yield (Sravan Kundojjala, 2026-10-10). This is one analyst's post; we didn't read the attached image, and there's no company announcement or second source. The amounts, the investment timeline and whether SpaceX or Tesla owns the fab are all unverified. If it holds, a new fab would have all four leading toolmakers lined up from the start; a company announcement or an order disclosure from one of the four would be the first confirmation.

Concept illustration: at lower left a smartphone in a red case stands on a beige ledge, its black screen showing rows of light grey bars like lines of text, and three circles of increasing size rise from it to a large oval thought bubble on the right. Inside the bubble is a building front with a black sign lettered TERAFAB above an open hall with rows of ceiling lights, where four black boxy machines with small screens stand side by side, all against an off-white background with a pale pink panel on the right.
Semiconductor analyst Sravan Kundojjala says in an October 10 post that the four big chip-equipment makers, ASML, Applied Materials, Lam Research and KLA, are all involved in Terafab, the chip fab Musk's companies plan to build. He puts SpaceX's first-phase volume fab at US$16.8 billion and a pilot line at another US$3 billion or more. He thinks KLA's inspection and metrology tools matter most in the early push for yield. (AI-generated)

2. [Today] (relayed October 10) Qualcomm CEO Cristiano Amon says some frontier AI companies want phones that can run a model of at least 100 billion parameters, more or less all the time, by 2028. In his words: "some of the companies at the forefront of AI—are telling Qualcomm, 'By 2028, I need to make sure that I have the ability to run at least a 100-billion-parameter model in a phone, and I need to have this thing running kind of all the time.'" He also said: "AI companies are starting to do phones." (Alex Heath's interview; transcript). He didn't say which companies; the passage comes via a third-party transcript, and we didn't listen to the original interview. What follows is our reading, not Amon's: the ask is for a model that stays running, not one called now and then, so its weights have to sit in the phone's memory permanently; memory capacity and standby power may therefore hit their limits before compute does.

Concept illustration: on the left, a woman with long dark hair in a brown suit sits at a long white table, gesturing with one hand, an open laptop in front of her showing a speech bubble with grey lines; a thought bubble linked to her by small circles holds a red smartphone whose screen also shows a speech bubble with grey lines. On the right, a man in a black turtleneck with a white card on a lanyard sits across the table with hands clasped, in front of a beige wall lettered QUALCOMM in black beside a blue ring-shaped emblem.
Qualcomm CEO Cristiano Amon says some frontier AI companies want phones that can run a model of at least 100 billion parameters, more or less all the time, by 2028. He also said: "AI companies are starting to do phones." Our reading, not Amon's: a model that stays running has to keep its weights in the phone's memory permanently, so memory capacity and standby power may hit their limits before compute does. (AI-generated)

Named commentary

1. [Today] (post, October 11) AI policy researcher Peter Wildeford splits the claim that "some lab has cracked AI self-improvement" into two meanings: AI doing all of R&D, which no company has achieved, and AI speeding R&D up, which all the major labs are doing. He writes: "If you mean AI can fully automate all R&D, I call BS - no company has cracked that yet. If you mean AI is accelerating development, then yes obviously the major companies are doing that." (Peter Wildeford, 2026-10-11). Wildeford's statement is opinion; he attaches no internal lab data. He was responding to an anonymous rumor with no source, which we don't use. Our September 27 issue reported that the automated training loop for Alibaba's flagship model Qwen3.8-Max ran 33 iterations over more than a month of fully automated operation; the press release doesn't say what each round did. The same release says that in a chip-design experiment the model kept improving itself for more than 60 hours and called chip-design software more than 10,000 times; that was a chip-design task, not model training (Alibaba Cloud, 2026-09-22). Our October 9 issue reported that the research group Epoch AI gave two models, Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol, 3,000 GPU hours each, and neither produced a comparable innovation in training methods. Applying Wildeford's split is our move; he didn't comment on either case. Alibaba's 33 rounds count as speeding R&D up. Alibaba doesn't say who set the goals or who chose which round's results went into the production model, and while people still do those two steps, it isn't AI doing the R&D. Epoch's experiment found that, given compute to do research on their own, neither model came up with a comparable innovation; that one test shows no sign of the fully automated end, though it doesn't rule out labs with more compute or internal tooling. That supports our September 27 reading: the big labs already run automated training pipelines, but that doesn't mean models will get better by themselves.

Concept illustration: a person with dark curly hair, dressed in black and wearing a lanyard with a small white card, sits on a chair at a long light-coloured table typing on an oversized red-framed laptop whose black screen shows a white speech bubble holding three grey bars. Behind them on the left a white robotic arm stands on a counter beside a small box with dots and vertical slots, a tall dark rack of stacked panels stands on the right, and a pale pink circle sits behind the person's head.
AI policy researcher Peter Wildeford splits the claim that "some lab has cracked AI self-improvement" into two meanings: AI doing all of R&D, which he says no company has achieved, and AI speeding R&D up, which he says all the major labs are doing. Applying his split is our move (he didn't comment on either case): Alibaba's 33 automated training rounds for Qwen3.8-Max count as speeding R&D up, and neither model in Epoch AI's experiment produced a comparable innovation in training methods. Our reading is that the big labs already run automated training pipelines, but that doesn't mean models will get better by themselves. (AI-generated)

2. [This week] (interview published in October) Michael I. Jordan of UC Berkeley calls "superintelligence" a science-fiction term, while Microsoft AI CEO Mustafa Suleyman said on October 10 that it must be contained as an engineering and governance problem. In an interview with the French newspaper Libération, responding to deep-learning pioneer Geoffrey Hinton's superintelligence argument, Jordan, one of the founders of statistical machine learning, said: "None of the steps in this argument is credible. Here, we're in the land of science fiction." He sees today's AI as an economic and social phenomenon (English translation of the Libération interview). Suleyman wrote on October 10: "Super Intelligence must be contained... Today, this is an engineering and governance challenge. And it isn't new. We've done it with planes, cars, nuclear materials, food safety, medicines" (Mustafa Suleyman, 2026-10-10). Their premises are opposite: Jordan thinks the argument doesn't hold, while Suleyman treats superintelligence as a real thing to be contained. Jordan was replying to Hinton, not to Suleyman; putting the two side by side is our choice. We read the full English translation of Jordan's interview but couldn't find the exact publication date. Suleyman didn't say how to contain it, and he has a commercial stake, since Microsoft is currently promoting a "humanist superintelligence" narrative; neither man is reporting a measurement. When a vendor or regulator talks about "containing superintelligence," first work out whether they treat it as something in front of us to be managed or as a premise that doesn't yet hold. Suleyman gave no specifics this time, and a concrete containment method from Microsoft would be the first thing to test the claim against.

Concept illustration in two panels: on the left, an open newspaper with the masthead lettered LIBÉRATION, columns of grey lines, and in its centre a red laptop whose screen shows a speech bubble with grey lines; on the right, a black smartphone turned sideways in front of a pale pink circle, its screen showing three grey lines above a similar red laptop with a speech bubble on its screen.
In an interview with the French newspaper Libération, Michael I. Jordan of UC Berkeley said none of the steps in Geoffrey Hinton's superintelligence argument is credible: "we're in the land of science fiction." Microsoft AI CEO Mustafa Suleyman wrote on October 10 that superintelligence must be contained as an engineering and governance challenge, but he didn't say how. When someone talks about "containing superintelligence," first work out whether they treat it as something in front of us to be managed or as a premise that doesn't yet hold. (AI-generated)

Model watch

1. [This week] (paper submitted October 8) The Internalizer preprint turns a document, in one pass, into a small set of add-on weights for a 284-billion-parameter model (DeepSeek V4 Flash); earlier methods of this kind went no larger than 14 billion. A small model built to generate weights reads the document and outputs a LoRA: an add-on layer that changes only a small number of parameters and hangs outside the original model, leaving the original model untouched. The abstract says the test documents were up to 4,096 tokens long and unseen by the model, and the model's input contained only a three-word instruction, not the document. A token is the chunk of text, a few characters long, that AI models count in, and usage is billed per token. The measurement: given the correct preceding text, have the model rank the tokens it thinks most likely to come next, and check whether the right answer comes first or in the top five. With the add-on layer, it came first 84.9% of the time and was in the top five 97.8% of the time; the same original model, also without the document in its input, scored 63.4% and 83.5% (arXiv 2610.11715). The abstract doesn't compare it with simply putting the document in the input. The paper says the scale is two orders of magnitude beyond earlier work; by our own arithmetic, 14 billion against 284 billion is only about 20 times, and the abstract doesn't explain the gap. Today, to make a model remember a document, you either put the document into the input every time or build a separate retrieval system. Our reading: once the add-on layer exists, you don't have to send the whole document again with every question; if generating a layer is cheap and fast enough, one layer per customer or per document becomes an option. What to watch: whether the authors release code and weights, whether anyone else reproduces the result, and what it costs, in money and time, to generate a layer. ⚠️ The numbers are the paper's own, and we read only the abstract; predicting the next token correctly is not the same as answering questions correctly, only short documents were tested, and the abstract gives no cost or latency for generating a layer.

Concept illustration in two side-by-side panels: on the left, a person in black clothing with a lanyard badge feeds a sheet of paper marked with lines into a slot on top of a dark box lettered INTERNALIZER, and an orange-red module with vent slots sticks out of an opening in the box's front. On the right, a person in the same black outfit and lanyard holds a similar orange-red module with both hands at a slot in a tall dark cabinet lettered DEEPSEEK V4 FLASH, with a panel showing a speech-bubble icon above the module and five grey vented drawer units stacked below it.
The Internalizer preprint uses a small model to read a document and, in one pass, output a LoRA: add-on weights that hang outside the 284-billion-parameter DeepSeek V4 Flash and leave the original model untouched; earlier methods of this kind went no larger than 14 billion. The abstract says that, with no document in the input, the model ranked the correct next token first 84.9% of the time with the add-on layer and 63.4% without it. Our reading: if generating a layer is cheap and fast enough, you don't have to send the whole document again with every question. (AI-generated)

2. [This week] (paper, October) The preprint "Verdict Without the Rule" tests 5 models across 20 regulatory and platform-policy areas and finds that when the rules are substantially rewritten, a model's rulings may not change to match. A guard model built for content moderation (the abstract doesn't name it) had 51% accuracy on an evaluation that rewrites its original classification scheme as custom rules, which the abstract calls close to chance. General-purpose models scored 90% to 92% on the same evaluation, but the guard model was tested in a different way, so this is not a straight ranking (arXiv 2610.12313). The paper's claim is that high accuracy can't prove a model is actually ruling by the rules you gave it. ⚠️ We read only the abstract of this one too, and it gives no sample sizes per area. For engineers who own compliance and content-moderation systems, here's one more acceptance test: change one rule and see whether the rulings change with it.

Concept illustration: a person with long dark wavy hair, dressed in black with a lanyard and white card around the neck, sits on a chair at a light-coloured table typing on a black laptop. The screen is split into two panels: the left shows rows of grey lines with one red bar among them, and the right shows a speech-bubble-shaped box filled with grey lines.
The preprint "Verdict Without the Rule" tests 5 models across 20 regulatory and platform-policy areas. It finds that when the rules are substantially rewritten, a model's rulings may not change to match, and it claims high accuracy can't prove a model is actually ruling by the rules you gave it. Engineers who own compliance and content-moderation systems can add one acceptance test: change one rule and see whether the rulings change with it. (AI-generated)

Product moves

1. [Today] (October 10) LangChain CEO Harrison Chase says moving model choice into the harness of Open SWE, the company's open-source coding agent (an AI that takes actions itself), so each task goes to the cheapest model that passes a quality bar, cut median cost per task by 64%. He writes: "most orchestration steps don't need a frontier model. for Open SWE we moved model choice into the harness, so each task goes to the cheapest model that still does the job, tested against quality. median cost per task dropped 64%" (Harrison Chase, 2026-10-10). The harness is the layer of code wrapped around the model that manages prompts, tools and flow. Which models were used before, and how the quality bar was set, haven't been published, nor whether the cost counts failed retries; this is one product's self-reported median and can't be extrapolated to other agents. Teams that want to try the same thing can first measure how often cheaper models pass on their own tasks, then decide which steps can be downgraded.

Concept illustration: an open laptop sits on a light desk at left, its screen showing a speech bubble with grey lines and its palm rest lettered OPEN SWE; a red cable runs from the laptop's side down to a small black box with a port and three dots, which stands next to a tall dark rack of eight stacked units on the right. A pale banner near the top shows a green parrot, a green chain-link icon and the lettering LANGCHAIN.
LangChain CEO Harrison Chase says moving model choice into the harness of Open SWE, the company's open-source coding agent, so each task goes to the cheapest model that passes a quality bar, cut median cost per task by 64%. He writes that most orchestration steps don't need a frontier model. Teams that want to try the same thing can first measure how often cheaper models pass on their own tasks, then decide which steps can be downgraded. (AI-generated)

2. [Today] (October 10) The coding agent Amp now lets users connect their own Claude subscription plan, and connecting is free. Quinn Slack, CEO of Sourcegraph and Amp, writes: "You can use your Claude plan in Amp now. Free and open to everyone." (Quinn Slack, 2026-10-10; Amp announcement). Free to connect doesn't mean free Claude usage; usage counts against each person's plan allowance, and we haven't checked what limits apply in practice. Our reading: people already paying for a Claude subscription can take it into another company's agent tool, so the tools compete more on interface and workflow than on which model they're tied to. If other coding agents start accepting subscription plans too, that reading gets stronger. ⚠️ Our analysis was produced with help from Anthropic's models, and Claude is an Anthropic product.

Concept illustration: on the left, an open laptop's dark screen shows a speech bubble with three horizontal bars, and the front of its base is lettered CLAUDE in large black letters; on the right, an upright black panel lettered AMP in white stands in front of a pale pink circle. A red-orange cable runs from the right side of the laptop base to a socket at the lower-left corner of the black panel, and both sit on a beige surface.
On October 10 the coding agent Amp began letting users connect their own Claude subscription plan; Quinn Slack, CEO of Sourcegraph and Amp, writes that it is free and open to everyone. Free to connect doesn't mean free Claude usage: usage counts against each person's plan allowance. People already paying for a Claude subscription can take it into another company's agent tool. (AI-generated)

From the archive

1. [Look back] (deep dive, August 21, 2026) When a chipmaker hands equity to a big customer in exchange for orders, the accounting rules, not the strike price, decide what it costs: it is booked as a cut to the chipmaker's revenue. The warrants Marvell issued to Google on August 18 have a strike price of US$206.58. Financial media worked that out as only 4.4% below the share price at the time, and our August 20 issue read it as a purchasing rebate, the buyer clawing back part of what it pays. The next day's deep dive checked the accounting rules and withdrew that reading: because the warrants were issued to win orders, equity given to a customer counts as consideration paid to the customer and reduces the seller's revenue at fair value on the grant date. That deep dive estimated the reduction at about six percentage points of Google's purchases (we don't rederive it here); the 4.4% figure used whichever of four possible reference dates made the strike look closest to the share price, and on the day the agreement was signed the strike was actually 26.4% above the share price (from that deep dive; original sources are listed there). On October 7 we went back and found two of Marvell's own documents: its quarterly report says the warrants it has issued to customers in the past reduce revenue at grant-date fair value as purchases are made, and the Google warrant agreement's definition of revenue excludes the reduction caused by the warrants themselves. The direction now has primary-document support, but both documents come from Marvell, so they don't count as two independent sources; the six-point figure can only be checked when the quarterly report due around early December first discloses the fair value; there is no result yet. When negotiating terms like these, treat the warrants as a discount quote: ask how many points the discount is and in which quarter it's booked. The full deep dive is published in Chinese and Japanese only; there is no English edition.

Concept illustration: two storefronts face each other. The left one has a sign lettered MARVELL beside a square geometric emblem, with a potted plant and a desk with a monitor inside; the right one has a sign lettered GOOGLE beside a multicoloured G-shaped mark, and through its glass a desk, a chair and a tall dark rack of chip-like squares are visible. Between them, a person with long dark hair in a black top and grey wide-leg trousers and a person with curly hair in a long beige coat, both wearing lanyard badges, each hold one end of a rust-red sheet with a wavy border and three pale lines.
Our August 20 issue read the warrants Marvell issued to Google on August 18 to win orders as a purchasing rebate, and the next day's deep dive checked the accounting rules and withdrew that reading. Equity given to a customer counts as consideration paid to the customer and reduces Marvell's revenue at grant-date fair value; the deep dive estimated the cut at about six percentage points of Google's purchases. When negotiating terms like these, treat the warrants as a discount quote: ask how many points the discount is and in which quarter it's booked. (AI-generated)

SecondSource isn't a news digest: each day we hunt the AI firehose for the insights that matter and the practitioner judgments worth tracking over time, and show how every item was verified. The point is always which judgment got harder and who's been right, not what happened today.

— SecondSource · generated by our research system · 21 sources · Got a view? Reply and tell us

Written from the same research and judgments as the Traditional Chinese edition. Sources are linked; we distinguish original documents from reporting and mark what we could not verify.