SFEIR's engineering-cabinet analysis ("an engineer's reading") of the **July 16, 2026** launch of **Kimi K3** by the Chinese laboratory **Moonshot AI**: an **open-weights, frontier-class model** whose provider claims **~2.8 trillion parameters**, a **one-million-token context**, and **weight release before July 27, 2026** (likely under a Modified MIT license, as with the K2 lineage). Thesis: capability once thought reserved for proprietary giants (Anthropic, OpenAI, Google) is becoming available **in open weights, at a discount price, from a Chinese lab**. SFEIR — despite being an **Anthropic and Google Cloud partner**, and thus "with no interest in overselling a Chinese model" — adopts a cardinal **methodological caveat**: on launch day, **no official, complete benchmark table** exists; specs (2.8T, Kimi Delta Attention, +25% training efficiency) and scores are **vendor-stated** or drawn from **community arenas**, "to be treated as claims, not measured facts." The new architecture (**Kimi Delta Attention**, hybrid linear attention; decoding claimed up to **6.3x faster** at 1M tokens) breaks with the K2 cadence (K2 Jul. 2025 → K2.7 Code Jun. 2026, a flagship every two months); two variants accompany the launch (**K3 Max**, **K3 Swarm Max**), with forced sunsetting of the kimi-k2.5/moonshot-v1 series on **August 31, 2026**. **The real weapon is price** (~$3/M input, $0.30 cached, $15 output per secondary sources): a frontier open-weights model at this level **pulls the whole price-performance curve down** — the commoditization of the model layer, accelerated by open source. But the decisive singularity is not a score: it is **reversibility**. A frontier open-weights model turns a consumed API (vendor dependency) into an **option** (self-host, portability, exit from lock-in), at the cost of heavy infrastructure to host 2.8T parameters. SFEIR's view: **open-weights changes the question, not just the answer** — no longer "which model is best/cheapest?" but "how much of my system am I willing to make dependent on a vendor I don't control?". The right posture remains a **routed portfolio** (one model per task, one model per constraint), with Kimi K3 adding a **"reversibility" column** to the decision grid. The "AI Only" conviction stands unchanged: the model is a commodity, the durable advantage lies in the engineering around it (Context Engineering, harness, cost governance, ability to change one's mind). The figures still need validating "on your own" — your repositories, your data.
Viral X thread (**230.5K views**, May 28, 2026, 1:51 AM) by **Jaya Gupta** (@JayaGup10, investor — likely Foundation Capital, author of the *Context Graphs* framework) titled ***"Token Budget Wars"***. **Pivot thesis**: ***"Enterprise AI has moved from adoption to allocation"*** — phase 1 of enterprise AI proved that models can work; phase 2 will decide **how much of that work is worth it**. The new currency at the top of the enterprise is the **ability to quantify AI ROI**: *"show me the value"*. Canonical concept: ***marginal token utility*** = *"the business value created by each additional dollar of inference"* — the number that matters at scale, and that **most companies cannot see**. Timeline: **Claude shipped November 2025**, after the 2026 annual budgets were locked → as early as **Q1**, companies *"running multiples ahead of plan"* → inference stops being an experimentation line item and becomes a **recurring operating cost**. Shift from **experimentation (a few $100K) → infrastructure (seven figures, $1M+)**: at infrastructure scale, **technical variance produces material P&L swings — two runs of the same workflow on the same input can differ by 5-10× in token cost** with nothing visibly broken, *"a number the CFO has to explain to the CEO"*. **AI competes with labor**: 3 types of budget requests (replace outsourced work / replace internal work / generate revenue) → shift toward the ***cost of a completed outcome*** (cost per resolved ticket, processed claim, reviewed contract, completed invoice, avoided hire, retained customer, dollar of revenue moved). **BPO = the easiest baseline to benchmark against** (already priced in completed units); internal work is much harder (multi-skilled employees, diffuse gains, HR resistance to headcount reduction). **Why it's different from SaaS**: SaaS learned to treat usage as a proxy for value; AI breaks that proxy — *"the signal and the noise share the same unit"* (the token), *"SaaS usage told you the software had been adopted. AI usage tells you the meter is running. It doesn't tell you whether your company is cooking."* **Three causes of marginal token utility's invisibility**: (1) ***retry tails*** — tokens per resolved workflow ≈ **T/p**; going from 90% to 70% completion increases effective cost by ~**28%**, not 20%, because failures compound; (2) ***context inflation*** — inference cost ≈ **O(n²)** in context length (attention), doubling the context **quadruples** reasoning cost (over-retrieval: 50 docs when 5 would do); (3) ***routing*** — by default the most powerful model is used (basic classification run on a complex reasoning model); across millions of calls, the difference between routing easy tasks to a small model and sending everything to the frontier model = *"the difference between a manageable bill and a board-level problem."* **Sector split**: **software** companies = a **productivity measurement** problem (already instrumented: PRs, commits, deploys, incidents, cycle time, MTTR — tracks *"AI layoffs"*); **non-software** companies = a **transformation** problem (operational work: claims, underwriting, support, compliance reviews, supply chain exceptions, payment disputes — *right under audit, not just right on average*). **The missing layer = token-to-outcome attribution**: a conversion layer linking inference spend → work performed → business outcome, answering 3 questions (real cost including retries/corrections; which parts of the trace mattered vs. thrashing; did the work change the operating model). ***Measurement becomes memory***: linking a token to an outcome requires capturing **decision traces** (what the agent saw, retrieved, called, ignored, where it retried, when a human overrode it) — *"decision rationale is one of the most perishable assets in a company"* (lives in Slack, emails, escalation calls, people's heads). Agents **create** these traces; captured first to justify the spend, they become *"more valuable than the cost report"* → a **context graph** (*"although I am so tired of that word these days"*). **The allocation layer is the prize**: whoever owns token-to-outcome attribution makes the **allocation calls** (which workflows deserve more compute, which are capped, which move to cheaper models, which stay human, which replace BPO). Companies won't do this on their own — they'll **buy it as a transformation** (Fortune 500 playbook: McKinsey + Palantir alumni + top-down CEO, in the manner of ERP/BI/digital transformation, a *"program"* with an executive sponsor and infrastructure that becomes the **new source of truth**). Framed by **Charlie Munger**: *"show me the incentive and I will show you the outcome."* Organizational sub-thesis: the decades-old executive instinct that *big teams = big jobs/scope/power* → once intelligence becomes the **scarce resource**, the new marker is *"how much of it you're orchestrating."* Direct relevance to the **Cost Optimization / agentic FinOps positioning**: empirically confirms the levers (model routing, prompt caching, context hygiene, sub-agents) and shifts the KPI toward **cost per completed outcome**. Strong convergence with Bain's *cross-system labor* (execution data moat, Cursor), Ng's *No AI jobpocalypse* (pricing anchored on the replaced employee's salary), DORA ROI (cost per feature), Mensch/Mistral (electron→token), Ensarguet (economics of computation), Foundation Capital's *Context Graphs* (decision traces, same author), Wescale's *Token Burning*, BFM/Girard (token = value fuel).
**Jaya Gupta** (@JayaGup10) — investisseuse / VC. Très probablement **Foundation Capital** (le thread s'auto-réfère au cadre ***Context Graphs*** — *« ahem, context graph, although I am so tired of that word these days »* — concept porté par Foundation Capital, cf. fiche `bain-100b-saas-opportunity` qui cite *Foundation Capital — Context Graphs trillion-dollar opportunity, 2025-12-22*). Thread publié sur X le **28 mai 2026 à 1h51** · **230 · 5K vues** · format essai long en un seul post. Une réponse notable de **@tuning_engines** (*« DevSecFinOps for the Agentic Era »*) : *« Tokens will basically have to be managed like headcount […] model hierarchies too »*.
Andrew Ng's editorial in The Batch #350 sets out an **acceleration hierarchy for coding agents** by type of software work: **Frontend (max) > Backend (moderate) > Infrastructure (low) > Research (minimal)**. The rationale rests on implicit *verifiability* (fluency in TypeScript/JavaScript plus an autonomous agent–browser test loop on the frontend) and on the LLMs' blind spots (corner cases / security / DB migrations for backend, opaque network tradeoffs for infra, irreducible hypothesis formation for research). The issue is rounded out by 4 structuring news items: **GLM-5.1 (Z.ai)**, a 754B/40B-active-parameter MIT-licensed model capable of autonomous tasks lasting 8 hours (SWE-Bench Pro leader at 58.4%); **Digit (Agility Robotics) at Schaeffler**, the first industrial deployment of humanoids (5'9"/143lb, $10–25/h vs $20/h for a human); the **anti-data-center revolt** (~$64B blocked May 2024 – March 2025, Maine moratorium on 20MW+ facilities, molotov cocktail at Sam Altman's home); and the **"assistant axis"** (Christina Lu, MATS / Oxford / Anthropic), which reduces persona drift and jailbreaks (Qwen3 32B: 83%→41%; Llama 3.3 70B: 65%→33%) without degrading IFEval/GSM8k/MMLU-Pro/EQ-Bench.
#Andrew Ng#The Batch#DeepLearning.AI
Andrew Ng (édito principal — fondateur DeepLearning.AI, Stanford, ex-Google Brain, ex-Baidu) ; rédaction The Batch (DeepLearning.AI) pour les sections actualités