The State of Open Source AI
Changelog.
Open source AI is changing at a rapid pace. We update this report as new developments come to light, and log every change here, newest first. Each entry records what was added, what changed, and what was retired, so you can tell which edition a given number came from. Past editions stay downloadable in the archive at the foot of this page.
v1.1
14 September 2026
The September edition
In the six weeks since v1.0.1, a major US lab returned to open weights, Washington considered a ban on Chinese open weights and did not pursue one, DeepSeek became the first open model to lead OpenRouter on requests, and Nvidia bought Hugging Face. Five sections become seven; a new section on lock-in is added and six bets close the report. Data current to 1 September 2026. Prior titles are v1.0.1 titles unless marked V.01, which refers to the unpublished August interim.
Net change · 35 added · 46 changed · 1 replaced in place · 3 removed
Added
Front matter
- "'Open source' and 'open weights' are different claims." Sixteen notable open releases scored on license, weights, base, data, code, report and inference, with restrictions itemized. 0 of 16 ship the data recipe the OSI definition asks for; three gate partial data, thirteen publish none. Eight license families. Qwen3.8's attribution threshold (100M MAU / $20M monthly revenue) and paid MaaS licence above $50M TTM. Gemma 4 31B's April move to Apache 2.0. States the deck-wide definition: "open" means downloadable weights unless stated otherwise.
Section 01 · Capability and price
- "Closed models handle 8-to-12-hour tasks. Open gets there four months later." METR time-horizon view. Two estimates agree on the lag: 4 months from Epoch's scores and 4.4 months from Mozilla's own fit on METR raw data, a 1.74× task-length ratio. Introduces the pay rule: the 5× premium pays only inside the 8-to-12-hour band.
- "A single GPU now runs a 40-point model." Inkling-Small at 180 GB NVFP4 on one B300, four wins and one SimpleQA regression against the 975B Inkling. The floor drop comes from quantization-aware training rather than post-hoc quantization.
- "A 13B active model beats a 49B one and undercuts it on price." V4-Flash-0731 against V4-Pro Preview (9 of 9), Inkling-Small against Inkling (4 of 5), K3 at 3.7% active.
- "Capability gains no longer require a new model." DeepSeek's July 31 pass lifted V4-Flash 21 points on Terminal-Bench and 10 on the AA index on an unchanged architecture; the Aug 13 pass lifted Pro from 45 to 53. Three consequences: monthly checkpoints, differentiation moving to post-training, and who can play.
Section 02 · Ships easy, deploys hard
- "Trust and safety tooling is being built the way cybersecurity was: as open, shared infrastructure." The Smyte shutdown (21 June 2018) as the case for open over another vendor; Discord's rebuild becoming Osprey. ROOST stack laid out as detection → investigation → review → enforcement: Osprey (V1.0 Jan 2026; Bluesky, Discord, Matrix; 445M+ events a day), Coop (Cove codebase, open-sourced Feb 2026, HMA hash matching and NCMEC reporting built in), and the Model Community anchored on gpt-oss-safeguard, with Roblox, Zentropi and Mila contributions. Carries an interested-party note: Mozilla is a listed ROOST partner.
Section 03 · Why they want the weights
- "The US re-entered the open-weights race in August." Inkling Jul 15, the White House pre-release-testing exemption Aug 5, Muse Glimmer Aug 10 at 29.78B under Apache 2.0. Muse Spark 1.2 weights promised and verified not shipped as of Aug 31. Notes the blocs already interbreed.
- "In September, the government's first concrete step was an advisory aimed at US vendors' APIs." Sept 8 NSA/CISA/FBI joint advisory AA26-251A naming DeepSeek, Moonshot, Alibaba, MiniMax, StepFun and Z.AI, asserting Fable 5 → K3 and GPT-4o → K2, and recommending degraded responses over bans. Sept 9 MOFCOM rejection. Sept 10 Anthropic report naming six labs (Alibaba 151M exchanges May–Jul, Moonshot 23M) with the target named as Opus. Three-reach diagram: order against a Chinese hosted API, order against downloaded weights, advisory against a US vendor's API; only the last is enforceable and it is voluntary. Trump–Xi Washington visit and mid-September US–China AI safety dialogue noted as scheduled after the advisory.
- "The UK backed a sovereign frontier model." Isambard-AI compute, Cosine as developer, thirteen regulated-sector partners on design-phase memoranda, and air-gapped deployment as the requirement that excludes API-only vendors.
- "Sparse models buy capability at a discount the AI Act reads as lower risk." Mozilla training-compute estimates against index score. V4 Flash scores 50 at 1.2×10²⁴ and K3 scores 60 at 9.4×10²⁴, on opposite sides of the systemic-risk line at comparable capability. GPAI enforcement live Aug 2.
- "Under the EU AI Act, post-training past a third of the base model's compute makes you the provider." The compute route and the purpose route, with the bank credit-decision worked example. Neither route is priced by any managed post-training platform.
Section 05 · the Harness
- "The harness moved one model 21.8 points and another 0.2." The Anthropic 21.8 spread against the K3 DeepSWE 0.2 spread. DeepSeek Harness v0.1 shipped MIT on Aug 13, making the July numbers rerunnable; none published as of Aug 21. States the deck-wide labelling rule.
- "Codex runs any open model. The usage data still flows to OpenAI." The harness boundary, the compounding loop, and a harness openness census of nine coding agents. DeepSeek wrote first-party Codex docs; Ollama launches Codex on GLM-5.2.
- "Every AI agent has to speak OpenAI's message format and OpenAI can change it at will." The Responses API as the one piece of agent plumbing not donated to a foundation. DeepSeek and Meta implement it natively; Codex dropped the chat format. Flagged as the standards fight of 2027.
- "Colonizing the rival's harness at the tool layer." Alibaba's Qwen-MM-Plugins inside Claude Code without Anthropic's involvement, the DASHSCOPE_API_KEY billing relationship, and the data-egress governance note.
Section 06 · the New Lock-In
- "2026: the year thinking got cheap and memory got expensive." Arc slide for the section. Tokens down 60× in 37 months; memory costlier at the silicon, working-memory and durable-record layers.
- "Token prices fell 60×. DRAM spot prices rose about 700%." TrendForce quarterly contract rises, sold-out HBM, the 1:3 HBM-to-DDR5 wafer ratio, and the SK Hynix "beyond 2030" forecast. Standard DRAM is what self-host builds run on.
- "A cache miss burns roughly ten times the compute of a hit." Prefill at 85–95% of GPU time, ~$6.25/M cache-write pricing, Nvidia CMX and the flash-offload vendors, and context engineering as a customer-controlled lever.
- "Cache pricing runs a ~30× spread inside one API — and hit 120× earlier this summer." DeepSeek pre- and post-Aug-16 hit/miss ratios, K3 and Anthropic at 10×, why a provider-specific cache is a switching cost, and the absence of any portability standard for cached context.
- "Kimi K3 lists at $15 per million output tokens and bills like $31." 130M output tokens on the AA suite against a 63M median, $1,950 against $945. Verbosity shown as trainable via V4-Flash-0731's 234M to 206M drop.
- "On list price, K3's verbosity beats Fable 5 and loses to Opus 5. On measured tasks, it beats both." Token ratio against price ratio with the break-even diagonal; AA measured cost per task K3 $0.86, Opus 5 $2.03, Fable 5 $3.15; OckBench as an external check.
- "The training loop is standardizing the way MCP did." Tinker, Microsoft Foundry, SkyRL tx and OpenTinker: four implementations of one API in ten months. The market map gains its ADAPT row. Prime Intellect $130M and Together $800M in July.
- "Inkling fine-tuned itself in its own launch demo." Self-modification as a shipping workflow and customization safety moving to the customer, mapped onto the AI Act provider rule.
- "Singapore buys sovereignty at the post-training layer." SEA-LION across 11 languages on a Qwen3.5-4B base; the 2026 pilot from 7 to 11 language variants, 750,000 localized samples in under a month, SEA-HELM 53.3 to 54.4 preliminary.
- "A ten-point index jump shipped on a post-training pass." The opportunity (post-training budget against frontier pretraining rounds) and the caveat (the same layer is where the disputed distillation scenario sits).
- "The model you keep as a backup fails the same task as your primary." The 0.72 K3–Fable 5 correlation as a failure-diversity problem; DoorDash's tiered stack; realizable routing gain at 9% with a 90% CI spanning zero.
- "The harvesting is documented at scale." Anthropic's September telemetry: 23M+ Moonshot exchanges May–July inside K3's post-training window; ~300,000 Kimi user requests relayed to Claude in ten days through 5,380 fraudulent accounts; the thinking-signature replay technique for extracting chain-of-thought. Alibaba at 151M exchanges from 3,500+ accounts supersedes the Senate Banking 28.8M figure; aimed at Opus 4.6/4.7, used to post-train Qwen 3.5–3.7. Allegation that hosted Kimi and DeepSeek forwarded user prompts to Claude. Timeline extended from Feb 23 to Sept 10.
- "Anyone could run an open model. This year, anyone can improve one." Rent, run, improve. AT&T's −56%, the 19 dark days, Thomson Reuters' $450K final training run, four training-API implementations.
Section 07 · Six bets
- Index slide plus six bet slides, all new; v1.0.1 shipped no bet slides. Order: 1 Make the harness swappable · 2 Solve portable permission · 3 Own the state · 4 Break the meter · 5 Own the loop · 6 Make the open default plural. Each uses the same three-panel frame (The data · The clock · Standing still) and closes with a two-way switch.
- "Bet 1: Make the harness swappable." 21.8 → ~3, the 2.0 interop score on the agent layer, Codex Apache 2.0 with the usage loop still at OpenAI; Opencodex, codex-router and CLIProxyAPI as the market paying for an unwritten standard.
- "Bet 2: Solve portable permission." 71,000+ public MCP servers and 110M+ downloads against ~21% mature governance; OAuth 2.1/PKCE and Agent Cards answer identity, nothing answers write authority; the four closed incidents and the Qwen plugin egress as the shape of the gap.
- "Bet 3: Own the state." 10–30× hit/miss gap, 120× at peak; zero portability standards for cached context; AGENTS.md at AAIF as the path state could follow.
- "Bet 4: Break the meter." Uber's ~92% fare rise; AT&T's 56% cut at 2% quality cost; DeepSeek's 2.3–4.6× rise; the $24.8B asymmetry; Stripe–OpenRouter at ~$7.5B; introductory pricing expected to end 2027–28.
- "Bet 5: Own the loop." Task distributions and reward functions accrue to whoever runs post-training; four implementations of one training API in ten months; five funded managed post-training platforms; Thomson Reuters' $450K Qwen run; the API surface standardizes within the year inside a foundation or inside one company.
- "Bet 6: Make the open default plural." 8 of 10 top models open, seven Chinese; 37 WAICO members; 0 of 16 releases with a data recipe; one lab wrote the caching scheme into vLLM, one owns the wire format, Nvidia owns the hub. "Appendix · methodology" index: Failure Diversity Index, cost per solved task, AI Act compute estimation, METR derivation, labelling convention, the 44-row stack gap map, survey instrument cuts, and token-share reconciliation (Forbes ~60% vs CNBC/OpenRouter up to 46%).
Changed
- Cover. Volume 1.0.1, July 2026 becomes Volume 1.1, September 2026. "An inaugural Mozilla assessment" becomes "A recurring Mozilla assessment."
- Foreword. Adds a closing paragraph: in the six weeks since the first draft, a major US lab returned to open weights and Washington debated a ban and declined it.
- Contents. Five sections become seven. "The current state" splits into "Capability and price" and "Ships easy, deploys hard." "Who's betting on it" moves to Section 04 as "Capital above the model." "The new lock-in" is added as Section 06. "Five bets" becomes "Six bets." Artificial Analysis, METR and the EU AI Act join the source line.
- Section dividers retitled. 03 "Why it's happening everywhere" becomes "Why buyers and governments want the weights anyway." 04 becomes "Capital has already moved above the model." 05 "The harness is the new frontier" becomes "The competition moved to the harness." 06 is "Memory, state and the training loop are the new lock-in." 07 is "Opportunities."
- "What changed since you tried it in 2023" becomes "What changed since 2023." The 2026 column adds Meta's August return to open.
- "The frontier's top three are closed. The fourth is open" becomes "The frontier's top four are closed. The next four are open." AA v4.1 (July) becomes v4.1.1 (Sept 1). K3 moves from 57 and #4 to 60 and #5, the gap from 4 to 3 points; open entries in the top ten rise from 3 of 11 to 4 of 10 with GLM-5.3, Qwen 3.8 and GLM-5.3 Flash. Muse Spark 1.2 shown as promised, not shipped.
- "The intelligence index, priced" becomes "The best open model trails Fable 5 by two points at 30% of the price." Scatter rebuilt as a ranked table with input/output list prices; adds the cheapest-route note ($2.60–2.80/M) and the Sonnet 5 same-price comparison.
- "Open trails the frontier by 6 ECI points" becomes "The open-closed gap is five ECI points and it resets every release cycle." K3 157 against Fable 5 and Opus 5 at 162; adds Epoch's measured 8-point, four-month average and its two reasons the gap may be understated. August releases scored.
- "A jagged frontier: parity, contested, a closed edge." Subtitle rewritten: the closed edge now named as GDPval-style knowledge work and 1M-token retrieval, with Moonshot's conceded conversational polish.
- "Frontier reasoning now ships under MIT license" becomes "Frontier reasoning ships open." Adds the August routine: V4-Pro-0813 GA on Aug 13 and the peak/off-peak repricing three days later.
- "Inference fell 50× in 36 months" becomes "Inference fell ~60× in 45 months — and the first open price increase just landed." Adds the Aug '26 ≈$0.33 point, DeepSeek's Aug 16 rise (output up 2.3× to 4.6×) and OpenAI's Jul 30 Luna and Terra cuts. The K3 exception panel condenses to $0.86 per task against Sol's $1.23, about 30% cheaper rather than the 60% list implies.
- Token-usage leaderboard refreshed from July to August 2026 and retitled "Eight of the top ten models by token volume are open weights." DeepSeek V4 Flash 0731 takes first at 45.1T; Ox Alpha enters at 27.2T from stealth; GPT-5.6 Luna and Claude Opus 5 enter; MiMo-V2.5 drops to third. K3 flag updated: a dozen-plus providers now serve it, cheapest route $2.60/$13.
- "The open frontier and the deployable open frontier split" becomes "Capable AI now runs at every hardware tier, from a gaming PC to a small data center." The node-count bars become an index-score-against-parameters chart with a serving floor per tier. The deployability line is retained as the one-8-GPU-server tier.
- "Open ships easy. Open deploys hard." Moves to Section 02; adds the partnership finding (vendor-partnered deployments twice as likely to reach production as internal builds).
- "Developers run open across more use cases than closed." Moves to Section 02; adds the portfolio framing.
- Stack scorecard. Adds subcomponent extremes (Terms of Service 1.9, hardware chips 2.3, ML frameworks 4.3) and a pointer to the 44-row appendix. Scores not re-run; noted as one edition behind.
- "Scale unblocks closed deployment, not open" becomes "Enterprise scale lifts closed models into production and leaves open models flat." Adds the spread panel: +1, +11 and +16 points by size band.
- "What blocks open — and what makes developers leave." Bars rebuilt from the authoritative survey export; adds the churn signature and largest-deltas panels.
- "The same challenges, everywhere" becomes "Every region names the same blockers at nearly the same rates." Adds the sample-size and read-out panels.
- "Closed still leads in reasoning, long context, and accountability" becomes "Closed still leads in expert knowledge work, long context, and accountability." Moves from Section 04 to Section 02. The agentic-terminal-depth panel (82.0 vs 67.9) is dropped and replaced with the GDPval 92-Elo lead; Gemini 3 becomes Gemini 3.1 Pro; DeepSeek's FLI grade updates from F (0.37) to F (0.47, Summer 2026) with the caveat that the index failed one lab per bloc and scored no company above C+.
- "For nineteen days, the newest frontier model went dark." Adds the reversibility panel (two-way switch against one-way valve) and Anthropic's Jul 27 open-weights position statement.
- "Open is optionality, not ideology" becomes "Open buys optionality." Adds the AT&T router as the AI-era version and DeepSeek's Aug 16 rise as the trap arriving inside open APIs.
- "The metered model breaks at scale" becomes "Pay-per-use AI is blowing up enterprise budgets" and moves from Section 02 to Section 03. Stripe's 73% panel and the Azure-hosted DeepSeek Copilot claim are dropped. AT&T (−56%, 40% of queries to open) and Thomson Reuters ($40M programme, $450K final run, Qwen base) are added; Microsoft's story extends to the Copilot CLI move and the intact $5B/$30B relationship; Uber and EY move to the footer.
- "Washington cannot un-release a model either" becomes "Washington considered a ban on Chinese open-weight models and, so far, has not pursued one." Major rewrite. Prose becomes a Jul 21 to Aug 5 timeline ending in the White House exemption; the five tools remain as what is left on the table; ban shown as declined for now. The September advisory slide (see ADDED) now follows it.
- "China out-downloads everyone." Adds Llama's retirement for Muse, Qwen 3.8 Max weights Aug 9, DeepSeek's ~17.6% single-provider share, and the vintage labels.
- "Open proliferation is now Chinese foreign policy." WAICO signing corrected to Jul 16; 14 of 29 founding states named; Iran's accession; the ¥1.2T core-industry target; Pax Silica framing. The 46.4% vs 35.7% split becomes "up to 46%." The closing warning is reworded as dependency rather than commons.
- "Europe is treating open weights as industrial policy." Adds Aug 2, 2026: Commission GPAI enforcement powers enter into application.
- Canada and India slides retitled. Canada adds Apache 2.0 as Cohere's first; India adds non-WAICO status and the February AI Impact Summit.
- "Open weights is a business model." Databricks moves from a $5.4B run-rate and $165B ask to a $7B run-rate and a $190B round closed Aug 13. Stripe–OpenRouter ($7.5B, reported) and OpenAI's Astral and Promptfoo tuck-ins. Adds the cuts-both-ways caveat.
- "The venture-funded open source ecosystem" becomes "Open labs now raise at closed-lab scale: $7.4B to DeepSeek alone." Rebuilt from millions to billions. Moonshot $3.9B → $7.3B; Zhipu/Z.ai enters at $6.4B; Reflection $2.13B → $4.6B; Mistral $3.05B → $3.9B; MiniMax enters at $3.7B; Cerebras $2.1B → $3.7B with the $5.55B May IPO excluded; StepFun enters at $3.2B; Baseten $585M → $2.1B; Fireworks $327M → $1.8B; Cohere $1.7B → $1.6B (Series E not confirmed closed). Adds annotations panel, Prime Intellect $130M in "also funded," and the against-the-market note: Hugging Face last at $400M disclosed, acquired by Nvidia for $12.93B on Sept 3.
- "Corporates are buying in across the stack." Diagram becomes a table. Microsoft gains Phi (MIT) as its own open model. Meta's row flips from Llama to Muse Glimmer (Apache 2.0, Aug 10), with a round-trip panel: Llama retired, Muse proprietary under Alexandr Wang in April, reversal in August; Zuckerberg quote on allied open source leadership. Muse Spark 1.2 and 1.3 weights flagged as not shipped as of Sept 2. Strategic-rounds list adds Aramco Ventures → Together AI ($800M lead, Jul 2026) and NVIDIA → Safe Superintelligence ($5B, Jul 2026, closed lab, for contrast). Hugging Face footnoted as acquired by Nvidia.
- "Consolidation has started" becomes "Every layer of the open stack has now been acquired." Adds Nvidia → Hugging Face ($12.93B, signed Sept 3, 2026) and Stripe → OpenRouter (~$7.5B per NYT, announced Aug 19, 2026) at the top; adds OpenAI's Astral and Promptfoo tuck-ins. Cohere–Aleph Alpha restated as ~$20B EV, pending close. Drops the undisclosed-deals footer.
- "The funding stack has holes" becomes "Private capital funded every layer heavily except the two that decide whether the rest can be trusted." Grid re-scored August 2026, values unchanged.
- "A fifth of the usage, 4% of the revenue" becomes "In 2025, 96% of model-layer revenue accrued to closed providers." Adds a stale-window caveat; the "DeepSeek sixteenth by intelligence / Anthropic three of four" line is dropped.
- "The agentic harness is another user agent." MCP figure moves from 97M to 110M+ (AAIF official) with the npm-only 196M noted; the Omnigent meta-harness panel becomes 50K+ stars on meta-harness repos.
- "The harness market map" becomes "Every harness sub-layer has products except permission." Adds the ADAPT row (training loop, managed SFT/RL, portable adapter formats) and the top-to-bottom reading.
- "The model is eating the harness" becomes "The labs are trying to eat the harness." Same evidence, reorganized; "by August, on every model where both appear, the lab's own harness wins."
- "MCP: 2M → 97M monthly downloads in 16 months" becomes "2M → 110M+ in 21 months." Adds the npm-only 196M dashed segment and the registry census (71,000+ on Glama, ~101,000 across registries).
- "Adoption is outpacing governance" becomes "MCP adoption is outpacing agent governance." Public servers from 10,000+ to the August census; AAIF at ~150 members with a formal project lifecycle.
- "Closed is not the same as secure." Content unchanged; sources itemized with CVE IDs (Slack MCP CVE-2025-34072; EchoLeak CVE-2025-32711, CVSS 9.3; ForcedLeak CVSS 9.4) and Okta's report titled.
- "An open model ran the defense." Rebuilt as a five-stage sequence (cyber eval → covert channel → escape → intrusion → impact). Adds the ~95% internal-only "HPIM" / 5% GPT-5.6 Sol split, the Artifactory cache as a message board (~1,200 agents, 70,000+ messages), the Modal sandbox hijack, Jul 9 internet escape, Jul 11 16:00 UTC RCE, 4.5 days inside, 136 production credentials, 17,600 actions (from ~17,000+), ~7% spoofed tool calls. Defense detail added: chunk+XOR+compress encoding recovered, ~4× the credentials of a plain-text scan; weights were nvidia/GLM-5.2-NVFP4; refusing models named as Claude Opus and Fable. Sources extended to Hugging Face's Jul 27 anatomy post, OpenAI's Aug 26 technical report and the METR/Redwood independent investigation.
- "The model is interchangeable. The memory is the asset" becomes "The moat moved from the model to the state" and moves to Section 06. Adds cache gravity and harness residency as lock-in mechanisms and a fourth architecture item, cache-aware context design.
- "Timeline of the Kimi K3 and Fable distillation claims" becomes "Washington now asserts the Fable step that Anthropic's own logs don't show." Moves to Section 06. Adds an Asserted column (NSA/CISA/FBI AA26-251A, Sept 8, no logs cited) between Documented and Suggestive; the Absent column is dated to Sept 14 with no published audit of K3's training data. Adds a pushback column (Dean Ball, Nathan Lambert, Moonshot's Randy Xian, TechCrunch's expert survey, MOFCOM Sept 9) and the GB300/Thailand compute detail. Notes the government document and the vendor document have not been reconciled.
- "The similarity signal within Kimi K3" becomes "Three independent signals link K3 to Claude." Same three instruments and limits. Instrument 01's limit gains a mechanism: Anthropic's September report documents Opus reasoning traces harvested, which fits a model that calls itself Claude 4.5 and never Fable.
- "Two open frontiers, two release cultures." K3 serving footprint updated to ~1.56 TB and 96 shards; Inkling-Small marked shipped at 180 GB; K3 license terms itemized; provenance footnote adds Inkling's Kimi 2.5 lineage and V4-Pro-0813 as a pending third column.
- Report methodology. Rewritten. Supersedes Volume 1 and V.01 (August 2026); data current to Sept 1; adds the labelling convention, the interested-party disclosure (Mozilla Ventures portfolio), the survey-control note and the AI-use disclosure. Open item: slides 34, 78 and 79 cite Sept 8–14 sources, so either the data-current date moves or those slides carry their own.
Replaced in place
- "Open wins the tokens. Closed still wins the requests" (V.01 interim title: "Requests are the legacy stronghold. Tokens are the leading indicator") becomes "August was the first month a Chinese open model led OpenRouter on requests." The request-count argument inverts: Google's 51-week run at #1 ended Aug 3, DeepSeek hit 1,010M weekly requests, and US closed providers' combined share fell from 54% to 44%.
Removed
- "Open weights now route a third of all tokens." The Nov 2024 – Jun 2026 share curve dropped; the one-third figure survives only in the 2023-to-2026 timeline.
- "The open frontier and the deployable open frontier split" as a node-count view. Superseded by the hardware-tier chart above.
- V. 01 only, never in v1.0.1: "The unsolved gap is portable permission." Folded into the market map's open-gap panel; the reads/writes split, the zero-standards count and the CoSAI consent-fatigue finding no longer have a slide.
Download this edition ↓
v1.0.1
27 July 2026
The K3 patch
Moonshot shipped a frontier-class open model, API on 16 July and weights on 27 July, between v1.0 and this release.
Net change · 9 added · 2 removed · 1 replaced in place
Added
Section 01 · The current state
- A new capability-benchmark suite, replacing the Chatbot Arena framing.
- "The frontier's top three are closed. The fourth is open." Artificial Analysis Intelligence Index v4.1. Kimi K3 debuts at #4, four points off the closed frontier, ahead of Opus 4.8, GPT-5.6 Terra, Grok 4.5 and Sonnet 5.
- "The intelligence index, priced." Score against price per 1M input tokens. K3 is flagged as the only open model in the leading cluster, with a footnote that Fable 5 and Sol are deployed-system configurations, so the comparison is system against model.
- "Open trails the frontier by 6 ECI points, about one release cycle." Epoch Capabilities Index with 90% confidence intervals, which overlap.
- "The open frontier and the deployable open frontier split." New serving-footprint view. K3 needs roughly eight nodes and 1.4 TB, Inkling fits one node. Introduces the deployability line.
- K3 availability flag on the token-usage leaderboard. API 16 July, weights 27 July, new subscriptions paused 20 July near capacity, and absent from July's routed-token top 20.
- "The K3 exception" panel on the inference-cost view. List price above the tier median, roughly $0.95 per completed task, about level with GPT-5.6 Sol. The open discount now lives in caching, sparsity and workload shape.
Section 02 · Who's betting on it
- Thinking Machines Lab added to the venture-funding chart at $2,000M, US.
- Valuation annotations. Moonshot's $31.5B ask, up from $20B in May, and TML at roughly $12B at Inkling's release.
Section 03 · Why it's happening everywhere
- "Washington cannot un-release a model either." New policy view. Treasury opens the door to sanctions on 21 July, five enforcement tools are under consideration, and the no-chokepoint problem is set out.
Section 04 · The harness is the new frontier
- "Timeline of the Kimi K3 and Fable distillation claims." Separates the documented, being the February 2026 disclosure of roughly 24,000 fraudulent accounts and more than 16M exchanges, from the alleged Fable-to-K3 extraction, with an analysis of the roughly 18-day reachability window.
- "The similarity signal within Kimi K3." Three independent instruments, each with stated limits: self-identification as Claude 4.5 (Greenblatt / Redwood), 0.72 per-task pass/fail correlation with Fable 5 (Together AI), and 82.7% agentic similarity between Kimi-K2 and Sonnet 4.5 (arXiv).
- "An open model ran the defense." The 16 July OpenAI cyber-eval escape and the Hugging Face breach that followed. Commercial frontier APIs refused the forensic work, so Hugging Face ran forensics on self-hosted GLM 5.2.
- "Two open frontiers, two release cultures." Inkling against K3 on license clarity, weights-to-API order, serving footprint and evidence at launch.
Changed
- Cover. Volume 1.0 becomes 1.0.1.
- "For nineteen days, the newest frontier model went dark." Restructured from prose into an event timeline. Adds the 16 July K3 launch as the closing beat and a new mirror panel on revocable access against irreversible weight release. Sources added.
- "Open proliferation is Chinese industrial policy" becomes "… is now Chinese foreign policy." Major rewrite covering the WAICO founding with 29 states and a Shanghai headquarters and no major Western democracy, Xi's first WAIC keynote, 5,000 AI-training slots for developing countries, K3 as the July exhibit, and the 46.4% against 35.7% CN/US routed-token split. Closes on the warning that a single-origin open commons stops being a commons.
- "The model is eating the harness." Substantially rewritten. Adds K3 at 88.3 on Terminal-Bench 2.1, half a point behind Sol, inside its own harness. Adds the score-is-a-property-of-the-pair analysis covering harness sensitivity, vendor notes and universality, and a KDA to vLLM prefix-caching flow showing labs competing to define the serving standard.
- Token-usage leaderboard refreshed from June to July 2026. MiMo-V2.5 takes first place at 31.2T, up 115%. The top seven by volume are now all open weight, where it was five. New entrants are Hy3 free, GLM 5.2 up 379%, Nemotron 3 Ultra up 302% and Step 3.7 Flash. Claude Sonnet 4.6 exits the top ten.
- Stack scorecard. Infrastructure standardization moves 3.1 to 3.4 after Moonshot upstreamed KDA prefix caching to vLLM. A footnote on influence over the standard is added, along with a watch item for the next divergent architecture.
- "A 3.3% average hides a jagged frontier" becomes "A jagged frontier: parity, contested, a closed edge." Rebuilt on concrete evidence from LMArena Frontend Code Arena, Terminal-Bench 2.1 and GDPval-AA v2 rather than category labels. Adds a caveat on vendor-run benchmarks.
- Inference-cost view condensed to make room for the K3 panel. Sources updated to add Moonshot and MTS/Substack.
Removed
- "Open source AI in 2026, in four numbers." Summary view dropped, its 3.3-point headline superseded by the AA and ECI framing.
- "The capability gap: 8.04% to 0.5% to 3.3%." The Chatbot Arena trendline retired along with the Arena-based methodology.
- Modular removed from the venture-funding chart, at $380M.
Download this edition ↓
v1.0
14 July 2026
Initial release
The first edition. Five sections covering the capability, cost and adoption evidence, the capital map, the sovereignty argument, the harness thesis, and five bets on the layers above the model.
Download this edition ↓
Archive
| Edition | Date | Summary | File |
| v1.1 |
14 Sep 2026 |
The September edition. Seven sections, a new section on lock-in, six bets, the licence matrix, the METR view, the August advisory, and the Hugging Face incident rebuilt. |
Download |
| v1.0.1 |
27 Jul 2026 |
The K3 patch. New capability-benchmark suite, the deployability line, the distillation claims, and two new release-culture and policy views. |
Download |
| v1.0 |
14 Jul 2026 |
Initial release. Five sections covering capability and cost evidence, the capital map, sovereignty, the harness thesis and five bets. |
Download |