大神觀點 · 本月 (文章日 9 月 9 日,回顧)
本篇出自 2026年9月18日 晨報
外殼(harness)指包在模型外面、替它決定每一步讀哪些檔、帶多少歷史紀錄、何時呼叫工具的那層程式,Anthropic 的 Claude Code、OpenAI 的 Codex 都是。他寫:「models are typically developed with one primary harness in mind (and fine-tuned less on other harnesses). Plus, the primary harness is often developed to suit and amplify a model's strengths.」(Sebastian Raschka,2026-09-09)。它接住本報哪條判斷: 第三方排行上兩家模型的成本差,有一部分來自外殼、不全是模型(排行數字見下方模型觀點第 2 則)。
只留 email,隨時退訂。這是我們唯一想請你做的事。
本週最強的具名觀點已經整篇寫在今日主線,這一欄出的是另一則,事件發生在 9 月 9 日、本報今天才讀到。Anthropic 對齊壓力測試團隊負責人 Evan Hubinger 在回…
他寫:「The Commerce Department decision to change the status of the UAE represents a major sh…
他評的是外部評測機構 METR 對「OpenAI hack of HuggingFace」那起事故的調查。事故發生在 2026 年 7 月,OpenAI 的模型以 AI 代理人身份…
他歡迎「deliberate pacing」與評測員的想法,但寫道「this cannot be controlled by a handful of entities」,必須跨生…