<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>thekb.eu — Economy &amp; Market</title><description>Economy &amp; Market · High-fidelity tech watch — AI, coding agents, SDLC</description><link>https://www.thekb.eu/</link><language>en</language><item><title>Claude Fable 5.1 and Mythos 5.1</title><link>https://www.thekb.eu/en/fiches/anthropic-claude-fable-5-1-mythos-5-1-2026-09-01/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/anthropic-claude-fable-5-1-mythos-5-1-2026-09-01/</guid><description>Product communication from **Anthropic** published on **September 1, 2026** on anthropic.com (~4,000 words, six sections, 22 testimonials from early-access partners). It announces **Claude Fable 5.1** (general availability) and **Claude Mythos 5.1** (verified access): *the same model, but with different levels of safeguards*.</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On **September 1, 2026**, Anthropic announces **Claude Fable 5.1** and **Claude Mythos 5.1**, presented as the most advanced models for coding and knowledge work. Both are **the same underlying model**; only the levels of safeguards differ. Fable 5.1 is in general availability; Mythos 5.1 is accessible only through trusted-access programs, with safeguards designed for cybersecurity and the life sciences.

The announcement explicitly responds to three pieces of customer feedback. **Pricing**: cache reads drop by 75% to $0.25 per million tokens, with input and output unchanged at $10 and $50; total cost falls by around 25% on typical workloads and up to 45% on heavily agentic workloads. **Data retention**: the new *Enterprise Frontier Safeguards* store data on the customer&apos;s own infrastructure, offering the privacy of a zero-retention agreement while preserving detection of adversarial usage; phased rollout starting in the fall. **Safeguards**: cyber classifiers produce 60% fewer false positives, and Fable 5.1 is now permitted to identify software vulnerabilities — without developing exploits.

On performance, Fable 5.1 reaches 52.6% on Terminal-Bench-Science 0.1 (versus 24.7% for Fable 5 and 29.0% for Opus 5), 55.8% on Terminal-Bench 4.0 (60.9% for Mythos 5.1), 1853 on GDPval-AA v2, 73.4% on CursorBench 3.2.0, and 31.4% on AutomationBench. Results are presented as cost/accuracy curves across five effort levels; at low or medium effort, the model matches or exceeds Fable 5 at a much lower cost. Twenty-two partners give testimonials, including Millennium, where the model diagnosed a one-in-a-million crash that no one had explained in four to five years.

The science section documents three results. In **molecular design**, Mythos 5.1 achieves a success rate of nearly 50% across 12 protein targets, with affinities ten times higher than the best submissions from Adaptyv Bio. In **modeling**, Fable 5.1 produced a carte altimétrique de Vénus covering one third of Venus from Magellan radar data, published under a Creative Commons license. In **computational biology**, Mythos 5.1 accelerated seven open-source models by up to 2.5× by writing GPU kernels, cutting costs by 30 to 60%.

On safety, Mythos 5.1 stays below the next risk threshold of the Responsible Scaling Policy in biology and in the lower category of the Frontier Compliance Framework in cyber. The alignment audit finds it better aligned than Mythos 5, while acknowledging limited coverage of long-context, multi-agent, and impossible tasks.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Claude Fable 5.1</category><category>Claude Mythos 5.1</category><category>foundation model</category><category>cache reads</category><category>cache pricing</category></item><item><title>The turbulent AI era is here. The choices we make now are critical.</title><link>https://www.thekb.eu/en/fiches/gates-ere-ia-turbulente-choix-critiques-2026-08-26/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/gates-ere-ia-turbulente-choix-critiques-2026-08-26/</guid><description>Essay published on **Gates Notes** on **August 26, 2026** by **Bill Gates**, co-founder of **Microsoft** and chairman of the **Gates Foundation**, ~4,500 words, announced as the first in a series. The text poses an alternative — AI will be the greatest equalizer ever invented, or the worst source of injustice — and a finding: no plan exists for entering this period. **(A) Three risks**: the lasting disappearance of entry- and mid-career jobs, white-collar as much as blue-collar, within a decade rather than several generations, because this time the substitution targets **cognition**; the weaponization of malicious actors (cyberattacks, bioterrorism, fraud, deepfakes), coupled with a concentration of power among those who already hold it; the effect of compagnons IA on children&apos;s development and on critical thinking. **(B) The benefits**, located in five domains — research, health, agriculture in low-income countries (the impact the author calls the fastest), public services, education — with a reservation carried on the verb: *&quot;the operative word is &apos;can&apos;&quot;*. **(C) Three proposals** open the series: building an unprecedented national and international institutional framework, borrowing from the nuclear inspection regime, aviation regulation, and ozone agreements; reserving certain occupations for humans, a domain named **Human Reserved**; **taxing AI tokens and robots** to rebalance the taxation of labor and capital. Gates discloses his financial ties to the industry and the transfer of his profits to the foundation. The text extends executive essays on the distribution of AI&apos;s value — [[nadella-frontier-ecosystem-human-token-capital-2026-06-12]], [[zuckerberg-meta-future-is-for-everyone-superintelligence-2026-08-10]] — by focusing on public power rather than the firm.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Bill Gates opens with his dual trajectory — building software at Microsoft, then redistributing the fortune thus accumulated — and draws from it his framing: AI will be the greatest equalizer ever invented, or the worst source of injustice. For the first time, a technology can replace and surpass human cognition. Yet no one is preparing for this transition: no plan exists.

He attributes this lack of preparation to an underestimation of the impact. Current model errors are misleading, since reliability is correcting itself quickly. More importantly, historical analogies are misleading: the PC took twenty years to transform work because software had to be developed, prices had to fall, and people had to be trained. AI, by contrast, runs on devices already installed and speaks natural language — it is AI that adapts to us. Gates discloses his ongoing financial ties to the industry, specifies that the profits from his investments will go to the foundation, and leaves it to the reader to judge.

He lays out three risks. First, job loss: because the substitution targets cognition, it hits law, customer service, medicine, software, and industry simultaneously, within a decade rather than across generations. Entry- and mid-career positions are the most exposed; blue-collar jobs will follow as robots dextres, developed mainly in China, become cheap. Second, the weaponization of malicious actors: cyberattacks, bioterrorism, fraud, deepfakes, given that beneficial and dangerous capabilities cannot be separated — and, symmetrically, the concentration of power among those who already hold it. Third, the effect on children&apos;s development and on human relationships, with compagnons IA described as a protected greenhouse that deprives people of the lessons of real contact.

The benefits are real and located: accelerated research, health, agriculture in low-income countries — the impact he calls the fastest —, public services, and education. But the verb remains &quot;can&quot;: nothing happens automatically, hence the necessary role of states and philanthropy.

He therefore proposes three initial measures. Building an unprecedented national and international institutional framework, borrowing from nuclear inspection, aviation regulation, and ozone agreements. Establishing a &quot;Human Reserved&quot; domain, occupations withdrawn from automation for economic or human reasons. Rebalancing taxation by taxing tokens and robots, since today the system pushes toward replacing people. He concludes by calling for widening the circle of voices shaping the debate.&lt;/p&gt;</content:encoded><category>Philosophy &amp; Society</category><category>AI and equity</category><category>transition to the AI era</category><category>cognition substitution</category><category>job disappearance</category><category>entry-level jobs</category></item><item><title>DuckDB and the changing physics of analytics</title><link>https://www.thekb.eu/en/fiches/warfield-duckdb-changing-physics-analytics-2026-08-26/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/warfield-duckdb-changing-physics-analytics-2026-08-26/</guid><description>Guest post by **Andy Warfield**, an engineer on the **S3** team at **AWS**, published on **August 26, 2026** on *All Things Distributed*, **Werner Vogels**&apos;s blog, who introduces it in a few lines signed &quot;--W&quot;: **3,554 words** per the page. The text serves as the vehicle for the announcement that **DuckLabs**, the team behind **DuckDB**, is joining **AWS**. (A) The thesis: systems computing is about seeking the elegant trade-off against a moving &quot;physics&quot; — the ratios between memory speed, network, and compute — and that physics has changed. Warfield quantifies the gap: an **m1.xlarge** from 2007 offered **15 GB of RAM**, **4 virtual cores**, and **~1 Gb/s** of network; an **m8g.48xlarge** today offers roughly **50×** more of each of the three. Dataset growth, meanwhile, follows a distribution whose tail consists of very large volumes. (B) The consequence: distributed processing — **MapReduce**, **Spark**&apos;s **RDDs** — was designed under the I/O constraints of the early 2000s, and much of the work assigned to it no longer needs to leave the application. Hence the embedded, in-process library engine, running in the application&apos;s address space, of which **DuckDB** is the example. Warfield anchors this in the *Scalability! But at what COST?* paper (2015) and **Paul Barham**&apos;s epigraph: &quot;You can have a second computer once you&apos;ve shown you know how to use the first one.&quot; He states an explicit caveat: &quot;When a job genuinely needs a thousand machines, it needs a thousand machines.&quot; The corpus already holds [[vogels-tech-predictions-2026-allthingsdistributed-2025-11-25]] from the same blog and [[anthropic-self-service-data-analytics-claude-agentic-stack-2026-06-03]] on self-service analytics.</description><pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Andy Warfield, an engineer on the S3 team at AWS, published a guest post on All Things Distributed on August 26, 2026, introduced by Werner Vogels. In it, he explains why embedded analytical engines like DuckDB are gaining importance, and announces that DuckLabs, the team that develops DuckDB, is joining AWS.

His reading grid is one of a moving &quot;physics.&quot; Where the physical sciences explore invariants, systems computing seeks the elegant trade-off against ratios that shift: memory speed versus network speed, richness of abstractions versus available power. He cites three moments — Berkeley&apos;s NOW project, his own work on Xen, and the MonetDB and X100 research at Amsterdam&apos;s CWI, where the bottleneck of query processing had shifted from disk to CPU — and notes that these constraints recur in cycles.

Applied to data, this grid explains distributed processing. Processing is always simpler and more efficient on a single fast machine, but when a server&apos;s disk or network card can no longer read the desired volume, one partitions. That was the constraint of the early 2000s, the one that produced MapReduce and then Spark&apos;s RDDs. Warfield notes two qualities of these systems: they innovated heavily on developer ergonomics, and they accepted a fixed cost of planning and distribution, betting on throughput gained by adding machines rather than on per-unit efficiency.

But the ratios have changed. A current instance offers roughly fifty times the memory, cores, and network bandwidth of the largest EC2 instance from 2007, while dataset growth follows a distribution whose extreme cases form the tail. The 2015 Scalability! But at what COST? paper had already shown that a carefully optimized single-thread implementation could beat distributed frameworks running on one hundred twenty-eight cores.

DuckDB, launched in 2018 by Hannes Mühleisen and Mark Raasveldt, applies this logic: an in-process library analytical engine, running in the application&apos;s address space, following SQLite&apos;s distribution model. AWS became a DuckLabs customer and then a sponsor of the Iceberg extension, alongside its work on S3 Tables; the extension now supports Iceberg v2 and v3 and exceeds 800,000 downloads per week.

Warfield does not present the embedded model as a replacement: when a job requires a thousand machines, it requires them. What is changing, he writes, is that much of the work done on data never actually needed a cluster. DuckLabs joins AWS as a subsidiary, with the project remaining open source under the MIT license and under the stewardship of the DuckDB Foundation.&lt;/p&gt;</content:encoded><category>Architecture &amp; Construction</category><category>DuckDB</category><category>DuckLabs</category><category>AWS acquisition</category><category>embedded analytical engine</category><category>in-process library</category></item><item><title>DeepSeek Harness developer preview: Everything is a plugin</title><link>https://www.thekb.eu/en/fiches/deepseek-harness-everything-is-a-plugin-2026-08-13/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/deepseek-harness-everything-is-a-plugin-2026-08-13/</guid><description>Official product page from **DeepSeek**, published on **August 13, 2026**, **unsigned**, ~450 words, announcing the *developer preview* release of **DeepSeek Harness** (`dsh`) — a coding-agent harness **open source under the MIT license**, whose repository opened the same day. A three-word thesis, repeated in the title and in the repository description: *« Everything is a plugin »*, paired with a second promise, *« Every run is traceable »*. The page states the equation *« AGENT = MODEL + HARNESS »* and lists the pluggable capabilities — *« models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI »*. Four modes ship: **Standard** (full coding agent), **Code** (tools exposed via the *Code Mode SDK*, letting the model compose multi-step operations inside a TypeScript program), **Minimal** (*« two-tool coding agent with persistent bash and str_replace_editor »*, explicitly *« for benchmarking models in a minimal environment »*), and **Creator** (runtime inspection, in-memory plugin testing). The technical substance sits in the repository, not on the page: `docs/architecture.md` states a logging invariant — *« Model-visible means logged. Anything that reaches a model request must be reconstructable from the log, and a runtime invariant asserts it »* — and states that *« there is no privileged core to patch »*. The technical core is not DeepSeek&apos;s own: DSH is built on **Cordis** (the `cordiverse` project, a third party), **vendored** into `vendor/` with a manifest and a sync procedure, and the page places the *« Cordis paper »* at the same navigation level as &quot;GitHub&quot; and &quot;Developer docs&quot;. Two LLM adapters ship — `dsh-llm-deepseek` and `dsh-llm-pi-ai`, a generic multi-provider adapter. The repository warns in capitals: *« THERE WILL BE COMPATIBILITY-BREAKING CHANGES »*, and `CLAUDE.md` specifies that `SESSION_FORMAT_VERSION` stays at `0` *« with no compatibility promise »*, with backends rejecting old on-disk formats. Timeline: DSH ships on the day **DeepSeek-V4-Pro reaches GA**, three days before a new API pricing schedule takes effect on **August 16, 2026 at 16:00 UTC**, with peak/off-peak rates and an off-peak discount of **−50%**.</description><pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Product launch page published on **August 13, 2026** by **DeepSeek**, **unsigned**, for the *developer preview* release of **DeepSeek Harness** (`dsh`), a coding-agent harness **open source under the MIT license** whose repository opened the same day.

**What the page says.** Two promises, in four hundred words and without a single figure. **« Everything is a plugin »**: every capability — models, tools, skills, sessions, sandboxes, storage, loops, scheduling, interface — is a plugin **swappable through configuration, without modifying the source code**. **« Every run is traceable »**: everything the model sees is recorded in an **append-only session log** — system prompts, reasoning, tool calls and results, subagent scheduling, every context injection — and *« resume, fork, search and replay all operate on the same event stream »*. The core is **Cordis**, a vendored third-party framework, described in an external paper and credited prominently. Four execution modes ship: **Standard** (full tooling), **Code** (tools exposed via a TypeScript SDK to combine several operations into one program), **Minimal** (two tools, persistent bash and `str_replace_editor`, *« for benchmarking models in a minimal environment »*), and **Creator** (runtime inspection, in-memory plugin testing, composition of new modes). Getting started: `npx @deepseek-ai/dsh web`.

**What the page does not say.** The strongest claim sits in `docs/architecture.md`: ***« Model-visible means logged. Anything that reaches a model request must be reconstructable from the log, and a runtime invariant asserts it. »*** **A guarantee asserted at runtime**, not a display claim — this is the property that actually sets DSH apart, and it is absent from the marketing copy. The same repository supplies the rebuttal: `SESSION_FORMAT_VERSION` stays at **`0` with no compatibility promise**, *« backends reject old on-disk formats »*, and the README warns in capitals that there will be breaking changes. **Traceable today does not mean archivable tomorrow.**

**The business model is in the timeline.** DSH ships on the day of **DeepSeek-V4-Pro&apos;s GA** and **three days before** a new API pricing schedule (August 16, 16:00 UTC; off-peak rates at **−50%**). **Harness given away, inference made pricier** — the exact reverse of Anthropic&apos;s model.

**What checks out.** Swappability holds at least at the model layer: besides the DeepSeek adapter, **`dsh-llm-pi-ai`** makes any OpenAI-compatible gateway accessible *« by configuration, not by code change »*. And the mode Minimal ships the **benchmarking harness** inside the product — an attempt to wrest the definition of the benchmark away from Claude Code, even as DSH&apos;s own repository contains a `CLAUDE.md` and a `.claude/skills`.&lt;/p&gt;</content:encoded><category>AI Coding Agents &amp; Skills</category><category>DeepSeek Harness</category><category>dsh</category><category>agent harness</category><category>agent harness</category><category>everything is a plugin</category></item><item><title>Mistral AI wants to build 1 gigawatt of European compute by 2030 — and lock in customers now.</title><link>https://www.thekb.eu/en/fiches/nunez-mistral-gigawatt-compute-europeen-venturebeat-2026-08-11/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/nunez-mistral-gigawatt-compute-europeen-venturebeat-2026-08-11/</guid><description>News article analyzed, published on **VentureBeat** on **August 11, 2026** by **Michael Nuñez**, based on an **exclusive interview with Timothée Lacroix**, co-founder and CTO of **Mistral AI**, conducted ahead of the announcement, ~2,000 words. Mistral is expanding its infrastructure offering in three parts: **Mistral Regional Endpoints** in general availability (pinning inference and its associated processing to Europe or the United States), a **Priority Tier** in public preview (committed service levels, custom quotas, availability SLA), and a **coalition of European enterprises** whose multi-year commitments are meant to fund **200 MW by the end of 2027** and **1 GW by the end of 2030**. The vehicle is called the **European Compute Unit (ECU)**: a claim on capacity built by Mistral, fungible across inference, training, model adaptation, or managed Kubernetes, over a targeted five-year horizon. Lacroix describes the mechanism bluntly — *&quot;The whole point of compute units is to have commitment&quot;* — and, on early exit: *&quot;There is no getting out.&quot;* The article scales the ambition: Mistral states it operates *&quot;less than 200 MW&quot;* and details three sites totaling **77 MW** (44 MW near Paris, 23 MW in Sweden with EcoDataCenter, 10 MW in Les Ulis); **Epoch AI** puts the initial capex for a one-gigawatt AI datacenter at **~$38B**, and **Goldman Sachs Research** puts next-generation facilities at **$15-20M/MW excluding chips**, against the **~$4B** Mistral has raised in total (PitchBook). Added to this is a decision that *&quot;is likely to raise a few eyebrows among sovereignty purists&quot;*: Mistral is starting to **host third-party open models**, beginning with **GLM-5.2** from **Z.ai**, a Chinese lab — *&quot;It&apos;s a great model. Everyone loves it. It&apos;s open-weight, so there was no good reason for us not to do it.&quot;* The article digs into the fine print of Mistral&apos;s documentation, which mentions *&quot;limited, controlled transfers&quot;* to subcontractors outside the region; pressed for detail, Lacroix points to **tool calls**, web search in particular, and states that **gating is the feature, not the bug**. The author&apos;s framing: *&quot;full regional control is available, but the moment an AI agent reaches out to the open web, sovereignty becomes a configuration decision, not a default.&quot;* Two dependencies remain: **GPUs** come from Nvidia, and **Microsoft** — anchor tenant of Mistral&apos;s European datacenters since July — is presented as what de-risks the buildout.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Article published on **VentureBeat** on **August 11, 2026** by **Michael Nuñez**, based on an **exclusive embargoed interview** with **Timothée Lacroix**, co-founder and CTO of **Mistral AI**.

**The announcement, in three parts.** (1) **Mistral Regional Endpoints**, in general availability: pinning inference and its associated processing to **Europe or the United States**. (2) A **Priority Tier** in public preview: committed service levels, custom quotas, an **availability SLA** for critical workloads. (3) A **coalition of European enterprises** — **Amadeus, ASML, Capgemini, CMA CGM** — whose multi-year commitments are meant to fund **200 MW by the end of 2027** and **1 GW by the end of 2030**. Added to this is the hosting of **third-party open models**, starting with **GLM-5.2** from the Chinese lab **Z.ai** (formerly Zhipu).

**The financial vehicle.** Commitments convert into **European Compute Units (ECU)**: a multi-year claim on capacity built by Mistral, fungible across inference, training, model adaptation, or managed Kubernetes. The structure resembles a **power purchase agreement** more than a cloud contract: lenders want demand locked in before capital goes out. Lacroix does not dress it up: *&quot;The whole point of compute units is to have commitment,&quot;* five years targeted, and on early exit — ***&quot;There is no getting out.&quot;***

**The orders of magnitude.** Mistral states it operates *&quot;less than 200 MW&quot;*; the detailed sites total **77 MW** (44 MW near Paris, 23 MW in Sweden with EcoDataCenter, 10 MW in Les Ulis). **Epoch AI** puts the initial capex for a 1 GW AI datacenter at **~$38B**, mostly in GPUs; **Goldman Sachs** at $15-20M/MW excluding chips; **McKinsey** estimates global need at **$5.2 trillion by 2030**. Mistral has raised **~$4B in total** (PitchBook), after **€830M in debt** for the Paris site.

**The fine print.** In-region inference remains subject to *&quot;limited, controlled transfers&quot;* to subcontractors outside the region: concretely, **tool calls** — web search in particular. Lacroix&apos;s answer: **cutting off capacity** is the feature, not the bug. A third endpoint, *&quot;on Mistral compute&quot;* outside hyperscaler hardware, is announced but does not yet exist.

**The repositioning.** By distributing third-party open models under regional controls and an in-house SLA, Mistral becomes a **sovereign distribution layer** — the *model garden* playbook of Bedrock and Vertex, in Europe. The competitive moat shifts from the model to the infrastructure. What finances all of it: the conviction that **trillion-parameter models and agentic tokens make on-prem inference untenable**, pulling revenue back to the cloud.

**The unresolved dependencies**: Nvidia **GPUs**, and **Microsoft** as anchor tenant of the European datacenters.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Mistral AI</category><category>digital sovereignty</category><category>AI sovereignty</category><category>European compute</category><category>gigawatt</category></item><item><title>To FDE, or not to FDE?</title><link>https://www.thekb.eu/en/fiches/zhang-decagon-fde-produit-2026-08-11/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/zhang-decagon-fde-produit-2026-08-11/</guid><description>Long-form article published on **X** on **August 11, 2026** by **Jesse Zhang**, CEO of **Decagon** (customer-service AI agents), under a dilemma-shaped title — *« To FDE, or not to FDE? »* — devoted to the **Forward Deployed Engineer**, which has become *« the answer to almost every hard question in AI go-to-market »*. Starting observation: Anthropic and OpenAI have built enterprise deployment arms explicitly modeled on Palantir, *« every seed-stage company »* advertises an FDE offering, and job postings for the title are said to be up several hundred percent in a year. **(A) The Palantir genealogy** supplies the framework: **Shyam Sankar**&apos;s (CTO) formula, *« FDEs eat pain and excrete product »*, and **Joe Lonsdale**&apos;s reminder that Palantir spent nearly two decades being called a *« glorified consultancy »* on the basis of an accurate observation. **Gotham**&apos;s bespoke deployments (CIA, NSA, military intelligence) were encoded into platform primitives — ontology, object models, permissions, workflow engines, provenance tracing — which became **Foundry**, then Apollo and AIP; standardization pushed gross margin into the 80% range and Palantir moved from an FDE motion to account-based selling, with many FDEs migrating into core engineering. *« The pain was the input to the product, not a cost of sale. »* **(B) The criterion proposed** is not to give up on FDEs but to know when to stop: go early, then ask whether one is still **discovering** — *« The trap is not starting. It&apos;s not stopping. »* **(C) A distinction few make: FDE ≠ implementation.** *« Building that integration into their ticketing system »* is real work, but it is execution against a known spec, not discovery of an unknown one; conflating the two *« is how a company convinces itself that a growing services org is a product investment »*. Closing line: *« If your FDEs are eating pain and excreting more pain, you don&apos;t have an FDE team. You have a services business. »* Two figures are put forward about Decagon — *« two-thirds of deployment work is now done autonomously via Duet »* and *« a few days on average to launch the first AOP, even for large banks, airlines, telcos »* — without the &quot;deployment work&quot; denominator being defined or the AOP acronym spelled out.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Long-form article published on **X** on **August 11, 2026** by **Jesse Zhang**, CEO of **Decagon** (customer-service AI agents).

**The starting observation.** The *Forward Deployed Engineer* has become the default answer to every AI go-to-market difficulty: painful deployments, customers unable to self-serve, product not ready. **Anthropic and OpenAI** have built enterprise deployment arms **explicitly modeled on Palantir**; job postings for the title are said to be up several hundred percent in a year. Yet, Zhang notes, until recently this was **a point of criticism** — lower-quality revenue, structurally capped margins — and *« nothing about the underlying economics has changed »*. What has changed: in the AI era, companies don&apos;t know the path to the outcome but they believe in the outcome, and **the FDE delivers the outcome**.

**The Palantir precedent.** Shyam Sankar, CTO: ***« FDEs eat pain and excrete product. »*** Joe Lonsdale acknowledges that the &quot;glorified consultancy&quot; reputation rested on an accurate observation. **Gotham**&apos;s bespoke deployments were encoded into primitives — **ontology, object models, permissions, workflow engines, provenance tracing** — which became **Foundry**, then Apollo and AIP. With standardization, **gross margin climbed into the 80% range** and Palantir left the FDE motion behind. *« The pain was the input to the product, not a cost of sale. »*

**The thesis.** Sending engineers is justified **when the category is new**: an accounting agent in 2026 has no established workflow, and the customer cannot even describe it. **But once the paths are known, the FDEs have to come out — and no one will want to**, because keeping them is easier sprint by sprint: one never has to settle a product trade-off, say no, or make a painful architecture choice. That leaves **all the drawbacks of the model with none of the discovery benefit**. Zhang further distinguishes **FDE from implementation**: one discovers an unknown spec, the other executes a known one; conflating the two lets a services org pass for a product investment.

**The Decagon case.** A deliberate product-led approach, driven by two constant enterprise demands: **iteration speed** and **refusal of vendor lock-in**. Cost: turning escalations into requirements rather than patches. **Self-reported** benefit: *« two-thirds of deployment work »* now done autonomously via **Duet**, and *« a few days »* to launch the first **AOP** at large banks, airlines, or telcos. Figures that are undefined and unverifiable.

**The closing line**: *« If your FDEs are eating pain and excreting more pain, you don&apos;t have an FDE team. You have a services business. »*&lt;/p&gt;</content:encoded><category>Strategy &amp; Frameworks</category><category>Forward Deployed Engineer</category><category>FDE</category><category>engineer embedded with the client</category><category>AI go-to-market</category><category>deployment motion</category></item><item><title>The Future is for Everyone: The Path to a Positive AI Future</title><link>https://www.thekb.eu/en/fiches/zuckerberg-meta-future-is-for-everyone-superintelligence-2026-08-10/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/zuckerberg-meta-future-is-for-everyone-superintelligence-2026-08-10/</guid><description>Doctrinal manifesto published on **meta.com** on **August 10, 2026**, signed with only a first name (*&quot;– Mark&quot;*) by **Mark Zuckerberg**, under the title *&quot;The Future is for Everyone: The Path to a Positive AI Future&quot;*, ~6,500 words. Three principles are announced from the outset: individual empowerment as a source of prosperity, invention as the primary purpose of superintelligence, balance of power as the foundation of safety. **(A) The central argument is a political argument**, stated as a short chain: *&quot;Humanity is not a monoculture&quot;* — people&apos;s values encode opposing trade-offs, no technical solution can align simultaneously with conflicting interests, so any singular superintelligence would have to prioritize certain values over others and would thereby be incapable of being benevolent toward everyone. Hence the formula: *&quot;There is no such thing as a singular benevolent superintelligence.&quot;* Safety is reframed as a problem of power distribution, illustrated by a thought experiment repeated three times (a single superintelligent lawyer versus everyone having one; the same for cybersecurity, then for business). **(B) A redefinition of alignment**: *&quot;Solving alignment is necessary for billions of people to adopt personal superintelligence agents. But it also implies that if we reach a state where billions of people are using and scrutinizing personal superintelligence agents, then we will have solved alignment with their interests.&quot;* The corollary targets the rest of the industry without naming it: *&quot;the most dangerous scenario would be leading labs training powerful models and keeping them for themselves.&quot;* **(C) Datable commitments**: a **fully private** mode where *&quot;even Meta&quot;* cannot see or grant access (a WhatsApp analogy); **free** versions for billions of people paired with a **dynamic bidding mechanism** for paid compute; the announced **resumption** of open source releases — *&quot;we will soon resume releasing some open source models&quot;*; and a structure giving the **independent board** the power to approve release safety criteria and verify each release&apos;s compliance, with the author acknowledging that Meta is a founder-controlled company. **(D) Two public-policy proposals**, repeated three times: that labs share **intermediate training checkpoints** and engineers with the government rather than an end-of-cycle review, and that the **physical production** of dangerous materials be regulated rather than the spread of knowledge. The text&apos;s sourcing is nearly nonexistent.</description><pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Manifesto published on **meta.com** on **August 10, 2026**, signed ***&quot;– Mark&quot;*** (**Mark Zuckerberg**), ~6,500 words.

**The three principles.** **Individual empowerment** as a source of prosperity, **invention** — not automation — as the primary purpose of superintelligence, and **balance of power** as the foundation of safety. The guiding question: *&quot;who will have access to superintelligence and what will we direct it toward?&quot;*

**The central argument.** Alignment conceived as convergence toward a single benevolent system is *&quot;fundamentally flawed&quot;*, because ***&quot;humanity is not a monoculture&quot;***: people&apos;s values encode opposing trade-offs, and no technical solution can align simultaneously with conflicting interests. Hence ***&quot;there is no such thing as a singular benevolent superintelligence&quot;***. Safety is not an engineering problem but one of **power distribution** — demonstrated by three identical thought experiments (lawyer, cybersecurity, business: a single holder causes harm, generalization benefits everyone). Corollary addressed to the industry: the most dangerous scenario would be *&quot;leading labs training powerful models and keeping them for themselves&quot;*.

**What Meta commits to doing.** A 24/7 personal agent with a **fully private mode** where *&quot;even Meta&quot;* cannot grant access; creation and business-creation tools; a personalized tutor; access to scientific advances (Biohub); **free versions** for billions, plus **dynamic bidding** for paid compute. On governance: the **independent board** will approve release safety criteria and verify compliance with them, with the author acknowledging that Meta remains **founder-controlled**. On openness: *&quot;we will **resume** releasing **some** open source models soon&quot;*, plus an explicit defense of **distillation** — *&quot;you can learn from anything you can observe&quot;*.

**Risks addressed.** Employment (nothing requires automation to outpace capabilities; finite compute creates an opportunity cost favoring invention); infrastructure (**community compacts**, the *Future Is For Everyone Fund*, a $50,000 bonus for Richland Parish teachers, water-positive by 2030); cyber and biorisk (defenders must retain the advantage; regulate physical production rather than knowledge); tyranny (privacy, **intermediate training checkpoints** to the government rather than a blocking review); American leadership (a decisive two-month lead, export controls maintained).

**Two caveats.** **Sourcing is nearly nonexistent** — the employment statistics, the HuggingFace incident, and China&apos;s nuclear capacity are not referenced. And **alignment becomes a consequence of adoption**: *&quot;if billions of people are using and scrutinizing personal agents, then we will have solved alignment&quot;*. This is the heaviest and least defended inference.&lt;/p&gt;</content:encoded><category>Philosophy &amp; Society</category><category>Mark Zuckerberg</category><category>Meta</category><category>Meta Superintelligence Labs</category><category>manifesto</category><category>corporate doctrine</category></item><item><title>Graphify — Knowledge Graphs for AI Coding Assistants (site graphify.net : vitrine, annuaire d&apos;outils et galerie de dépôts graphifiés)</title><link>https://www.thekb.eu/en/fiches/graphify-net-annuaire-ia-coding-2026-08-06/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/graphify-net-annuaire-ia-coding-2026-08-06/</guid><description>The **graphify.net** site, accessed on **August 6, 2026**, maintained by **Safi Shamsi** — the creator of the open source graphify skill (cf. [[skill-shamsi-graphify-2026-08-06]]). The domain carries two distinct objects. **The first is a product showcase**: presentation of graphify, usage guides, CLI reference, and above all a gallery of **100 already-graphified trending GitHub repositories** — *« 100 repos, 854,079 nodes, 1,932,930 edges »* — filterable by language and graph size, each with its own preview and detail page. **The second, and it is the more interesting one for tech-watch purposes, is an editorial directory**: *« 30 AI coding client guides »*, a directory of MCP servers compared on *« transport, runtime, client support, setup effort, and access risks »*, structured comparisons between tools (Cursor versus Codex), and a stream of articles with a manifestly long-tail targeting (*« GLM-5.2 Knowledge Graph for Developers »*, *« Trae Context Engineering for Agents »*, *« Symphony Knowledge Graph for Agent Memory »*, *« What Is Cowart? A Codex Plugin for Image Editing »*). The site claims a method — *« source-reviewed »*, *« aligned decision fields, official evidence, and explicit unknowns »* — and is available in six languages. **The point this fiche exists to record**: the site is **factually out of step with the product it presents**. It announces **« 3.7k+ GitHub Stars »** when the GitHub API counts **103,187** on the same day, a **MIT license** repeated three times when the repository&apos;s `LICENSE` file is **Apache 2.0**, and highlights the **« 71.5× token reduction »** claim, which belongs to the v1-generation README and has disappeared from the current version. **An official site displaying 3.7% of the actual star count and getting the license wrong** is a signal in itself: the communication layer has not kept pace with the repository.</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The **graphify.net** site, accessed on August 6, 2026, officially owned by **Safi Shamsi**, creator of the open source graphify skill. The domain carries three things distinct from the commercial platform `graphify.com` and the GitHub repository.

**A product showcase**, first: presentation of graphify, usage guides, CLI reference, pages on tree-sitter and Leiden clustering.

**A demo gallery**, next, and it is the most compelling part: **100 already-graphified GitHub Trending repositories**, totaling **854,079 nodes and 1,932,930 edges**, filterable by language and size, each showing its node, edge, and community counts, with a graph preview and detail page. Showing the tool running on well-known repositories is worth more than a pitch, and it produces, as a byproduct, a public dataset of comparable graphs.

**An editorial directory**, finally, which has value independent of the product it promotes: **30 AI coding client guides** compared on workflow, agents, pricing, security, and delivery fit; an **MCP server directory** rated on transport, runtime, supported clients, setup effort, and **access risks**; pairwise comparisons on aligned fields. The site claims a method — *« source-reviewed »*, official evidence, explicit unknowns — and is available in six languages.

**This fiche exists mainly to record a discrepancy.** On the same day, the site announces **« 3.7k+ GitHub stars »** when the API counts **103,187**; it states **three times** an **MIT** license when the repository&apos;s `LICENSE` file is **Apache 2.0**; and it highlights the **« 71.5× token reduction »** claim, which belongs to the v1-generation README and has disappeared from the current version in favor of LOCOMO and LongMemEval benchmarks. The site thus describes a product several generations old.

**The license error is the most serious one**: MIT and Apache 2.0 do not carry the same obligations, notably on patents and the disclosure of modifications.

A strategic observation remains: **a tool vendor building the directory of its own category** occupies the evaluation query ahead of its competitors. The claim of neutrality does not remove the conflict of interest — graphify appears among the site&apos;s featured skills. A useful entry point, not an arbiter.&lt;/p&gt;</content:encoded><category>Tools &amp; Platforms</category><category>graphify.net</category><category>AI tool directory</category><category>directory</category><category>AI client guides</category><category>tool comparison</category></item><item><title>Efficient Tokens &amp; Effective Teams in Buzz</title><link>https://www.thekb.eu/en/fiches/patel-block-buzz-teams-tokens-benchmarks-2026-08-06/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/patel-block-buzz-teams-tokens-benchmarks-2026-08-06/</guid><description>A **Block Engineering** benchmark post from **August 6, 2026**, signed by **Atish Patel**, about **Buzz** — the human + agent workspace launched on July 21 — asking a cost question: which agent team is **the cheapest one that reliably succeeds**? Three findings. **(A) A negative result, published in full**: on **Terminal-Bench 2.1**, **twelve team compositions** (pairs, triads, cheap swarms under a *frontier* model) were pitted against the solo agent each was built around, and **none beat it at equal cost**. The explanation is structural — a task that finishes in minutes *&quot;doesn&apos;t have enough structure to divide&quot;*, and *&quot;More agents mostly buys you the cost of explaining it twice&quot;*. **(B) The horizon reverses the result**: on **Long-Horizon Terminal-Bench** (44 tasks, one task worth hours of work, same lead **GPT-5.6 Sol** at *high* effort), solo finishes 15 tasks for 59.1%, +2 QuickBees 19 for 64.1%, +1 QuickBee +1 WorkerBee 19 for 69.5%, **+2 WorkerBees 20 for 71.5%** — a **+12.4-point** gain, of which 11.4 comes from tasks carried to completion. *&quot;Same seats, opposite result, because the work is a different shape.&quot;* These runs ran at **3× the timeout**, solo included. **(C) Beyond a threshold, price stops buying quality**: solo on Terminal-Bench 2.1, **Opus 5 at *xhigh* effort is the most expensive run ($140.63) for 75.0%**, trailing six runs ranging from $20.08 to $109.82 and 79.5% to 88.4% — the stated cause is over-reasoning that drove 17 of 88 tasks to timeout. Among the six best runs, **a 5.5× price gap for an 8.9-point score gap**: *&quot;choosing between them is not a quality decision at all. It is a budget decision.&quot;* The post proposes a taxonomy it owns as *ad hoc* — **QuickBee**, **WorkerBee**, **SmartBee**, plus the human as *&quot;honorary bee&quot;* — and two team forms, the permanent **Hive** that remembers your preferences and the disposable **Swarm** that remembers the project. Conditions: everything runs on **Harbor**, against real Buzz agents on a **live** relay, **one attempt per task, no retry**, prices fixed as of **2026-07-30**.</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;A **Block** benchmark post signed by **Atish Patel**, published on **August 6, 2026**, extending the **Buzz** launch: since assembling an agent team there has become trivial, *which one is the cheapest that reliably succeeds?*

**Vocabulary first.** The post proposes four tiers: **QuickBee** (fast and cheap — builds, screenshots, tests, first-pass triage: GPT-5.6 Luna, DeepSeek V4 Flash, local models, **run at high effort**), **WorkerBee** (versatile, carries a full subset unsupervised: GPT-5.6 Terra, Gemini 3.6 Flash, open models), **SmartBee** (big picture, trade-offs, escalations: Claude Opus 5, Kimi K3, GPT-5.6 Sol, **at *medium* effort**), and the human, *&quot;the most expensive bee on the team, and the slowest. Also still the smartest&quot;*. Two team forms: the permanent **Hive**, which remembers **your** preferences, and the disposable **Swarm**, which remembers **the project** and then disappears.

**The solo result.** On **Terminal-Bench 2.1**, raising the effort of a **cheap model** is the best buy: Luna goes from $1.61 / 57.3% (*medium*) to $4.98 / 75.0% (*high*). At the other end, **Opus 5 at *xhigh* is the most expensive run ($140.63) and scores only 75.0%**, having **hit the timeout on 17 of 88 tasks** through over-reasoning. Among the six best runs: **a 5.5× price gap, an 8.9-pt score gap**. Conclusion: *&quot;choosing between them is not a quality decision at all. It is a budget decision.&quot;*

**The team result, in two acts.** On Terminal-Bench 2.1, **twelve compositions** were tested and **none beat solo at equal cost** — a short task doesn&apos;t have enough structure to divide. On **Long-Horizon Terminal-Bench** (44 multi-hour tasks, lead GPT-5.6 Sol, **3× the timeout**), the reversal is clear: solo **15 tasks / 59.1%**, +2 WorkerBees **20 / 71.5%** — **+12.4 pts, of which 11.4 come from additional completions**. The team costs more per task, which pays off *&quot;when the alternative is a human picking up unfinished work&quot;*.

**The operating rule.** Route worker escalations to a **SmartBee coordinator** rather than to the human: *&quot;every ambiguity becomes a notification&quot;* is the real failure mode. A Block engineer says they **migrated over 2,000 apps** with a Swarm (coordinator, 1–10 migrators, independent verifier), the coordinator storing human answers in memory.

**Caveats**: n=1 per task, no confidence interval, team costs unpublished, and an admission — *&quot;this might change if models are trained on better collaboration.&quot;*&lt;/p&gt;</content:encoded><category>AI Coding Agents &amp; Skills</category><category>Buzz</category><category>Block</category><category>agent teams</category><category>team composition</category><category>multi-agent</category></item><item><title>Block explores how to price AI</title><link>https://www.thekb.eu/en/fiches/paymentsdive-block-dorsey-pricing-ia-2026-08-06/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/paymentsdive-block-dorsey-pricing-ia-2026-08-06/</guid><description>Trade-press brief (**Payments Dive**, *Dive Brief* format, **August 6, 2026**) covering **Block**&apos;s quarterly earnings release: the company has already rolled out several AI tools to its customers — **Moneybot** (Cash App) and **Managerbot** (Square) — and has not yet decided how to charge for them. **Jack Dorsey** on the analyst call: *&quot;We&apos;re in a fortunate position where we can experiment with a number of models, and then choose the right one that&apos;s going to align all of our incentives with our customers.&quot;* **The financial backdrop illuminates that stance.** Six months earlier, Block had laid off roughly **4,000 people, about 40% of its workforce**, in a reorganization explicitly framed around AI. In Q2 2026: gross profit **up 25% to $3.2B**, revenue **up 10% to $6.62B**, but **net income at $89M, down 83%** year over year due to severance costs closing out the restructuring; 2026 guidance was raised. AI&apos;s value, then, is being captured through the cost structure before it is captured through price. **The heaviest fact sits in the middle of the brief**, drawn from the shareholder letter: *&quot;Starting in June, agentic AI helped write and review nearly all of our production code changes&quot;* — writing **and** reviewing nearly all production code changes, at a publicly traded payments company, six months after cutting 40% of the workforce. A self-reported claim to investors, with no definition of *&quot;nearly all&quot;* or of what *&quot;review&quot;* covers. **The tooling**: **Goose**, an internal system built two years earlier, described as model-agnostic (it plugs in different commercial models for employees); **Buzz**, launched the previous month for *&quot;agent collaboration, communication, and code repositories.&quot;* **On the customer side**: Moneybot monitors Cash App user activity and surfaces accounts, balances, and transactions — over **one million weekly active accounts**; Managerbot runs automated marketing, margin analysis, and suggests *&quot;operational fixes&quot;* to Square merchants. **Evercore ISI** analysts list four monetization paths — SaaS bundles, direct subscriptions, enterprise offerings, usage-based pricing — **none of them tied to outcomes**. Stated order of priority: **product quality → distribution → adoption → pricing model**. Two distribution facts round out the picture: Square is rolling into **Google Maps** with a *&quot;conversational AI experience,&quot;* described as *&quot;the first step in a broader partnership between Square and Google&quot;*; and the **Tags** payment device (keychain and NFC chip wands) shows **three million people on the waitlist**. Analyst quotes: William Blair (*&quot;Block epitomizes the secular shift toward tech-forward digital finance firms&quot;*) and Bank of America on the *&quot;post-reset operating model.&quot;*</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;A **Payments Dive** brief from **August 6, 2026** on **Block**&apos;s quarterly earnings, the parent of **Cash App**, **Square**, and **Afterpay**.

**The stated topic.** Block has rolled out several AI tools to its customers and **has not yet decided how to charge for them**. **Jack Dorsey**, on the analyst call: *&quot;We&apos;re in a fortunate position where we can experiment with a number of models, and then choose the right one that&apos;s going to align all of our incentives with our customers.&quot;* The company is consulting Square merchants on their needs. **Evercore ISI** analysts list four possible paths — SaaS bundles, direct subscriptions, enterprise offerings, usage-based pricing — noting that Block is prioritizing *&quot;product quality, distribution, and adoption&quot;* first.

**The real story, left for the reader to piece together.** **Six months earlier**, Block laid off **roughly 4,000 people, ~40% of its workforce**, in a reorganization centered on AI. In Q2 2026, **gross profit rose 25% to $3.2B** while revenue climbed only 10% to $6.62B; **net income fell to $89M, down 83%**, weighed down by severance costs; **2026 guidance was raised**. Not a single dollar of AI has been billed to customers: the value has already been captured **through the cost structure**. The &quot;fortunate position&quot; that lets Dorsey take his time on pricing is exactly what the workforce cut bought.

**The buried number.** In the shareholder letter: *&quot;Starting in June, agentic AI helped write and review nearly all of our production code changes.&quot;* Writing **and** reviewing nearly all production code changes, at a publicly traded payments company. A self-reported claim to investors, with no definition of *&quot;nearly all&quot;* or of *&quot;review.&quot;*

**The tooling.** **Goose**, an &quot;agnostic&quot; internal system built two years earlier, plugging in several commercial models for employees. **Buzz**, launched the previous month, for agent collaboration, communication, and code repositories. On the customer side, **Moneybot** (Cash App) tracks activity, surfaces accounts, balances, and transactions, and has passed **one million weekly active accounts**; **Managerbot** runs automated marketing and margin analysis for Square merchants.

**Two distribution facts.** **Square is entering Google Maps** with a conversational discovery-and-ordering experience, *&quot;the first step of a broader partnership&quot;* with Google. And the **Tags** device (NFC) shows **three million people on the waitlist**.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Block</category><category>Jack Dorsey</category><category>Cash App</category><category>Square</category><category>Afterpay</category></item><item><title>Announcing Cloudflare Wallets: the programmable wallet for the agentic Internet</title><link>https://www.thekb.eu/en/fiches/cloudflare-wallets-agentic-commerce-2026-08-04/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/cloudflare-wallets-agentic-commerce-2026-08-04/</guid><description>Product announcement published on the **Cloudflare** blog on **August 4, 2026** by **Will Papper**, as part of **Agents Week**: **Cloudflare Wallets**, presented as *&quot;the programmable wallet for the agentic Internet&quot;*. **The problem stated** is precise and well chosen: an agent that wants to try an API has to go through a login page **designed for humans**, have a human add a payment method, generate an API key, then figure out how to call the service. Two structural gaps explain this — *&quot;Agents do not have a stable identifier to sign up for an API, and they do not have a native way to pay for APIs&quot;* — with the consequence that *&quot;AI agents often give up on these tasks entirely, kicking registration, payment methods, and API key generation back to humans&quot;*. **The proposed architecture comes down to two wallet types**: **Account Wallets**, intended for humans who own a Cloudflare account (fund, delegate, withdraw), and **Virtual Wallets**, intended for agents, **operating via API key** and whose spending cap is **set by the account holder**. The announced guardrails are explicit: **allocation, allow list, maximum amount per transaction**. **The payment rail is the x402 protocol** (payments attached to HTTP requests) and the currency is **stablecoin** — which places the offering in a distinct camp from schemes built on card networks. **The most interesting argument is counterintuitive and central**: *&quot;These limits may seem like constraints, but counterintuitively they give agents more freedom. If an agent is responsible for $10, you can worry less about its spending than if it is responsible for $1,000.&quot;* → **the cap is not what constrains autonomy, it is what makes it acceptable.** **Second component, more strategic than the first**: identity, via a **`cloudflare.pay`** namespace — a research agent could live at `research.example.cloudflare.pay`, giving the merchant certainty that it is talking to the agent of an identified organization. Cloudflare claims a deliberately minimal ambition (*&quot;a human-readable identifier for a not-very-readable keypair, similar to the URL and IP-address pairings used in DNS&quot;*), built on its existing building blocks (**Turnstile**, Bot Management, **Web Bot Auth** and its keypairs), and states its intent to adopt the schemes of the **x402 Foundation** as they emerge. **A decisive caveat about the status of the text**: **almost everything is in the future tense**. What exists on the day of the announcement is the **reservation of a handle**; payments, Virtual Wallets, guardrails, and the ramps for accessing funds are announced (*&quot;Soon, you will be able to…&quot;*). This is a **staking of position on a namespace**, more than a service going live.</description><pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Announcement published on the **Cloudflare** blog on **August 4, 2026** by **Will Papper**, during **Agents Week**: **Cloudflare Wallets**, *&quot;the programmable wallet for the agentic Internet&quot;*.

**The problem.** An agent that wants to try an API has to get through a login page designed for humans, have a human add a payment method, generate a key, then discover the API. Two gaps explain this: *&quot;Agents do not have a stable identifier to sign up for an API, and they do not have a native way to pay for APIs.&quot;* As a result, agents give up and hand everything back to a human.

**The architecture.** Two wallet types. **Account Wallets** belong to the humans who own an account: fund, delegate, withdraw. **Virtual Wallets** are intended for agents, operate **via API key**, and their cap is **set by the account holder** — with allocation, allow list, and maximum amount per transaction. The rail is the **x402** protocol, which attaches a payment to an HTTP request, and the currency is **stablecoin**: a positioning distinct from schemes built on card networks.

**The central argument is counterintuitive**: *&quot;These limits may seem like constraints, but counterintuitively they give agents more freedom. If an agent is responsible for $10, you can worry less about its spending than if it is responsible for $1,000.&quot;* The cap is not what constrains autonomy, it is what makes it acceptable — and if trying an API costs a few cents, ten dollars is enough to compare many of them.

**The second component is identity**, and it is more strategic than the first. An agent can live at `research.example.cloudflare.pay`: an optional identity, delegated from the account, persistent, which finally makes free trials and sign-up credits attributable. Cloudflare claims a minimal ambition — *&quot;a human-readable identifier for a not-very-readable keypair, similar to the URL and IP-address pairings used in DNS&quot;* — building on **Web Bot Auth** and announcing the adoption of the **x402 Foundation**&apos;s schemes. The analogy used is the VPN: not being identified does not make one suspect, it simply requires proving oneself more.

**A decisive caveat**: almost everything is in the future tense. What exists on August 4 is the **reservation of a handle**. Payments, virtual wallets, guardrails, and fund ramps are announced. Add to this an unsourced figure on the majority of traffic coming from bots, complete silence on European compliance, and a vertical integration where the same actor would supply the wallet, the merchant gateway, identity, and bot control.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Cloudflare Wallets</category><category>agentic commerce</category><category>Agents Week</category><category>programmable wallet</category><category>Account Wallet</category></item><item><title>How AI is expanding what people do at work (Work at the Frontier, rapport 1)</title><link>https://www.thekb.eu/en/fiches/openai-work-at-the-frontier-task-crossover-2026-07-27/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/openai-work-at-the-frontier-task-crossover-2026-07-27/</guid><description>Post and report from **OpenAI Economic Research** published on **July 27, 2026**, the first installment in the **Work at the Frontier** series, analyzing **more than 800,000 messages from US ChatGPT users**. **Coined concept**: ***task crossover*** — *« work historically associated with one occupation appearing in the AI use of people in another »*. **The headline figure is actually two figures, and that&apos;s the point coverage loses**: **16.8% of work-related messages** concern tasks associated with another occupation, and **43.5% of occupation-specific messages**. The funnel explains the gap: **61.5% of usage is generic** (writing, summarizing, planning — too widely shared to count as evidence of crossover) and is excluded; of the **remaining 38.5%**, **43.5% fall outside the occupation** and 56.5% are *« inside **or near** »* — so the upper bound is calculated on a reduced base, while the lower bound is calculated on the entire professional usage. **By occupation** (share of occupation-specific messages pointing to an external task): customer experience **77%**, design **75%**, HR **69%**, legal **56%**, marketing **53%**, sales **40%**, finance **40%**, engineering **28%** — *« a majority in five of eight groups »*. **Two distinct directions of circulation**: design **imports** (35.2%) and **exports** almost nothing (1.7%); engineering does the opposite (imports 18.5%, exports 7.4%); **marketing does both** (imports 24.3%, exports **8.9%**, the highest outward share in the sample). **Two tasks appear in the top 3 of borrowings for the other seven groups**: **financial calculation** and **technology troubleshooting**. **The heatmap, absent from coverage, is the richest object**: it gives the full distribution of tasks by user occupation, and its diagonal is striking — engineering retains **53%** of its own work while customer experience retains only **11%**, HR **10%** and design **12%**. **Size effect**: the outside-occupation share drops from **18.9%** (2-5 employees) to **16.3%** (&gt;100 employees) — **but only « among average users »**, OpenAI specifying that *« among the heaviest users, we do not see the same monotonic pattern »*, and concluding conditionally: *« AI **may be** especially useful as a generalist tool where specialist resources are scarce. »* **Claimed status**: an **early signal**, visible *« before firms rewrite job descriptions or create new job titles »*. **Structural caveat**: OpenAI measures OpenAI&apos;s own usage, on US ChatGPT users only, and presents this position as an asset — *« our unique window into how the world of work is changing »*.</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;First installment in **OpenAI Economic Research**&apos;s **Work at the Frontier** series (July 27, 2026), based on more than **800,000 messages** from US ChatGPT users.

**The concept.** ***Task crossover*** refers to *« work historically associated with one occupation appearing in the AI use of people in another »*. The methodological counterpoint is set out from the start: exposure studies begin with a fixed list of tasks and ask whether the model can perform them; here the question is **who does what**. *« AI changes not just how work gets done, but who does what. »*

**The figures, and there are two.** **16.8%** of work-related messages and **43.5%** of occupation-specific messages concern a task from another occupation. The gap comes from the funnel: **61.5%** of usage is **generic** (writing, summarizing, planning) and is excluded; of the remaining 38.5%, 43.5% fall outside the occupation, with the rest being *« inside **or near** »*.

**By occupation**: customer experience **77%**, design **75%**, HR **69%**, legal 56%, marketing 53%, sales and finance 40%, **engineering 28%** — a majority in five of eight groups.

**Two directions of circulation.** Design **imports** (35.2%) without exporting (1.7%); engineering does the opposite (18.5% / 7.4%); marketing **does both** (24.3% / 8.9%, the highest outward share). Two tasks appear in the top 3 of borrowings for the other seven groups: **financial calculation** and **technology troubleshooting**.

**The heatmap** gives the full distribution, and its diagonal is the most striking result: engineering retains **53%** of its own work, while customer experience retains only **11%**, HR **10%** and design **12%** — for these three occupations, marketing tasks outweigh their own.

**The size effect is more fragile than it appears.** The outside-occupation share drops from 18.9% (2-5 employees) to 16.3% (&amp;gt;100 employees) **among average users only**: *« among the heaviest users, we do not see the same monotonic pattern »*. The conclusion remains conditional — *« AI **may be** especially useful as a generalist tool where specialist resources are scarce »*.

**The claimed status** is that of an **early signal**, visible *« before firms rewrite job descriptions or create new job titles »*.

OpenAI measures usage of its own product, on its US users only, and presents this position as an asset.&lt;/p&gt;</content:encoded><category>Transformation &amp; Adoption</category><category>OpenAI Economic Research</category><category>Work at the Frontier</category><category>task crossover</category><category>task overflow</category><category>occupational porosity</category></item><item><title>Aiman Ezzat, le directeur général de Capgemini : « L&apos;enjeu ? Intégrer l&apos;IA au coeur des opérations et réinventer les processus métiers »</title><link>https://www.thekb.eu/en/fiches/ezzat-capgemini-ia-agentique-processus-metiers-2026-07-25/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/ezzat-capgemini-ia-agentique-processus-metiers-2026-07-25/</guid><description>Capgemini (Aiman Ezzat, CEO) — Investir interview, &quot;boss special&quot;: IA agentique as an operational breakthrough, not just another technology; €2bn invested, +30% on application development and −20% incidents, &gt;11% of Q1 bookings, TAM of over $400bn/year by 2030 — but &quot;very far from plug and play&quot; (Investir / Les Echos)</description><pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In the &quot;boss special&quot; issue of **Investir** devoted to the AI challenge (25 July 2026), **Aiman Ezzat**, CEO of **Capgemini**, argues a simple and commercially loaded thesis: **AI&apos;s value does not come from the technology itself but from its integration at the core of operations**. IA agentique, &quot;capable of acting autonomously,&quot; marks in his view a **major breakthrough** that will enable a structural transformation of how companies function — and puts the integrator at the center of the game.

**The evidence put forward.** Three years ago, Capgemini committed an investment of **€2 billion** (offering portfolio, partner ecosystem, employee training). The claimed effects are measured on the delivery side: on certain projects, **more than 30% faster application development** depending on application type, and **nearly 20% fewer incidents and service interruptions**. On the market side, generative and agentic AI projects account for **more than 11% of first-quarter bookings, up from 6% a year earlier**. The stated ambition: **5.5% to 7.5% annual growth** at constant exchange rates by **2028**, with improved profitability and cash generation.

**The diagnosis on clients.** Generative AI paved the way with **individual productivity gains of limited impact**; agentic AI goes further by introducing &quot;a new form of work,&quot; with agents that execute tasks, embed themselves in business processes, and contribute to decision-making. But delivering on the promise is &quot;anything but simple&quot;: **complex legacy systems, insufficiently mature data, governance, security, costs**. Scaling requires rethinking systems, data, processes, organization, and operating models — &quot;**we are very far from plug and play**.&quot;

**The program.** Build an **agentic technology layer on a modernized foundation**, **orchestrate collaboration between humans and agents**, **control the costs** of this new workforce. Without clear governance of roles, security, and responsibilities, &quot;deploying thousands of agents enterprise-wide would be a dead end.&quot; Hence the reframing of the question: &quot;the question is not who develops the best models but who helps companies get value from them.&quot;

**The market and employment.** The agentic transformation spills beyond traditional IT budgets into **operational budgets and strategic priorities**; Capgemini estimates the opportunity at **over $400 billion a year by 2030** for digital services and consulting. On employment, Ezzat remains cautious: a profound impact on jobs, with tasks automated and jobs created, but &quot;too early to say&quot; whether the net balance will be negative. The acquisition of **WNS** creates &quot;a global leader in **intelligent operations**,&quot; announced as a growth pillar.&lt;/p&gt;</content:encoded><category>Transformation &amp; Adoption</category><category>Aiman Ezzat</category><category>Capgemini</category><category>IA agentique</category><category>autonomous agents</category><category>business processes</category></item><item><title>IA et emploi : le vrai risque, c&apos;est le décrochage</title><link>https://www.thekb.eu/en/fiches/sfeir-ia-emploi-risque-decrochage-2026-07-23/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sfeir-ia-emploi-risque-decrochage-2026-07-23/</guid><description>In-depth opinion piece published on **sfeir.com** on July 23, 2026, signed by **SFEIR** (the firm&apos;s editorial voice). It is a **strategic commentary on Trésor-Éco note No. 391** from the DG Trésor (June 2026 — see [[dgtresor-ia-effets-emploi-2026-06-30]]), read through SFEIR&apos;s doctrine of « **amplifying AI rather than enduring it** ». The article praises Bercy&apos;s **cautious economist&apos;s tone** (mechanisms plus uncertainty rather than a prediction) and draws from it a **three-part thesis**: (1) **no measurable aggregate effect** at this stage (two offsetting forces — displacement vs. productivity — EU adoption ~20%); (2) a **single solid empirical signal, on juniors** (−16% employment among exposed 22-25 year-olds in the US); (3) a **long-term danger that shifts the question** — **competitive lag** (non-adoption), not job destruction. The analytical core SFEIR retains: **price elasticity** determines the employment effect (the **Jevons** paradox applied to code) → the argument is **structurally pro-employment for developers**. The article **dismantles the &quot;AI layoffs&quot; narrative** (4.5-6.2% of US layoff announcements, &quot;labeling&quot; at 59%) and points to the note&apos;s **blind spots** (the agentic scenario relegated to a footnote; diffusion speed not discussed; OpenAI/Anthropic having become sources for Bercy = an unflagged source bias). **SFEIR&apos;s operational translation** (for CIOs/CTOs): value migrates toward intent/architecture/control, training **augmented engineers** (**AI Champions** programs), and avoiding rushed adoption (**workslop**, technical debt) through **context engineering** and governance.</description><pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In this opinion piece published on sfeir.com (July 23, 2026), **SFEIR** comments on the **Trésor-Éco No. 391** note from the DG Trésor (June 2026) and anchors it to its own doctrine: *« amplifying AI rather than enduring it »*. The article praises Bercy&apos;s **cautious tone** — which lays out mechanisms and uncertainty rather than settling the matter — and draws from it a three-part thesis *« more reversed than it appears »*.

**No aggregate effect.** Within the Acemoglu-Restrepo framework, two forces oppose each other: the **displacement** effect (substitution) and the **productivity** effect (complementarity, lower costs, increased demand). They currently offset each other; studies identify no aggregate effect, for lack of hindsight and adoption (~20% of EU firms). Individual gains are nonetheless real (+14% in customer service, +26% among developers), but anxiety outpaces the data (62% of French people worried).

**The only solid signal: juniors.** −16% employment among exposed 22-25 year-olds in the US (Brynjolfsson 2025); in France, a contraction in youth employment in IT and rising unemployment among 15-24 year-olds (19.1%→21.1%) — without established causality. The mechanism: AI automates the **codified tasks** of entry-level positions, the ones that *« used to train tomorrow&apos;s seniors »* — hence a **renewal-of-expertise** issue.

**The argument the debate misses.** A profession&apos;s fate hinges on the **price elasticity** of demand, not exposure: developers and graphic designers (elasticity &amp;gt; 1) see demand grow as AI lowers their costs — the **Jevons paradox applied to code**. The argument is **structurally pro-employment for developers**. The article also **dismantles** the &quot;AI layoffs&quot; narrative (4.5-6.2% of US layoff announcements; **labeling** at 59%) and points to the note&apos;s **blind spots**: the **agentic** scenario relegated to a footnote (which would invalidate the &quot;assistant&quot; framework), **diffusion speed** left undiscussed, and **source bias** (OpenAI/Anthropic having become sources for Bercy).

**The real fault line: competitive lag.** Bercy shifts the burden of proof — the risk is **competitive** (falling behind in adoption), not social. Hence the programs (« Osez l&apos;IA », France 2030).

**SFEIR&apos;s perspective**: for a CIO/CTO, this translates into decisions — value migrates toward intent/architecture/control; train **augmented engineers** (AI Champions); avoid rushed adoption (**workslop**, technical debt) through **context engineering**, governance, and POC-to-production criteria. *« Turning adoption into a lever rather than a pile of POCs. »*&lt;/p&gt;</content:encoded><category>Transformation &amp; Adoption</category><category>AI and employment</category><category>competitive lag</category><category>non-adoption</category><category>Trésor-Éco 391</category><category>Bercy</category></item><item><title>Mistral ↔ Microsoft : un accord souverain, une stratégie industrielle encore illisible</title><link>https://www.thekb.eu/en/fiches/sfeir-mistral-microsoft-souverainete-strategie-industrielle-2026-07-22/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sfeir-mistral-microsoft-souverainete-strategie-industrielle-2026-07-22/</guid><description>SFEIR analysis (firm&apos;s voice, &quot;an engineers&apos; reading&quot;) of the deal announced on **July 21, 2026** between **Mistral** and **Microsoft**: an **industrial partnership worth several billion dollars**, structured in three parts — (1) **compute in Europe** (reserved Azure capacity on the continent, datacenters in France, latest-generation **NVIDIA Vera Rubin** systems, to &quot;close the European compute deficit&quot;); (2) **Mistral&apos;s models in Microsoft&apos;s tooling** (**Mistral Medium 3.5** and **Mistral OCR 4** in **Microsoft Foundry**, accessible in **Copilot Studio** to build business agents); (3) above all **Azure Local down to disconnected mode** (public cloud, supervised connected cloud, and **air-gapped** entirely off the external network — for defense secrecy, healthcare, critical banking). **Notable fact, confirmed by Brad Smith: no new equity stake** by Microsoft in Mistral&apos;s capital — a massive partnership **without a capital tie-up**. SFEIR — an Anthropic and Google Cloud partner, &quot;with no interest in overselling the French champion&quot; — regards Mistral as **&quot;the best European bet on the model layer&quot;** and offers a three-part reading. **What the deal brings a CIO**: a leading-edge European model, executable in a disconnected environment and controlled by the customer (in-memory encryption, locally managed keys), checks boxes that few offerings check. **The tension**: this sovereignty is deployed **on the infrastructure of an American hyperscaler**; four sovereignties must be distinguished — **model, execution, infrastructure, commercial relationship** — of which one can &quot;get three out of four, but you still need to know which one is missing.&quot; The only element that makes sovereignty **truly portable** is the **open-weights nature** of Mistral&apos;s weights (the same reversibility logic as for **Kimi K3**). The absence of an equity stake is not a detail: it preserves Mistral&apos;s governance **and** minimizes the risk of an antitrust review (FTC, European Commission) — **assumed regulatory arbitrage**, not just technical choice. **The real blind spot**: the **legibility of Mistral&apos;s industrial strategy**, present simultaneously on nearly every front (B2C with Le Chat, B2B via Azure distribution, open-weights model **and** frontier ambition, highly capital-intensive infrastructure — 200 MW secured, a 1 GW cap by 2030 —, partnerships with a handful of large accounts, Robostral/OCR verticalization, service to regulated sectors): sovereign full-stack (optimistic reading) or the dispersion of a three-year-old company valued at ~€20B across businesses with divergent economic models (cautious reading). For technical leadership: **separate the model from the channel**, **design to exit** (Design to Exit — open-weights makes the exit door credible), **route rather than bet** (sovereign multi-LLM architecture, RAISE). Conclusion: **sovereignty is an architectural property, not a label** — it is qualified dependency by dependency; the missing industrial legibility remains the real open question, settled not by press releases but by &quot;the trade-offs of the next twelve months.&quot;</description><pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On **July 21, 2026**, **Mistral** and **Microsoft** announced a strengthened partnership in the form of a **deal worth several billion dollars**. SFEIR — an Anthropic and Google Cloud partner, therefore &quot;with no interest in overselling the French champion,&quot; yet regarding Mistral as &quot;the best European bet on the model layer&quot; — offers an **engineers&apos; reading** of it.

**What the deal actually says**, in three parts to be distinguished from the messaging: (1) **compute in Europe** — reserved Azure capacity on the continent, datacenters in France, **NVIDIA Vera Rubin** systems, to close the European compute deficit; (2) **the models in Microsoft&apos;s tooling** — **Mistral Medium 3.5** and **Mistral OCR 4** in **Foundry**, accessible in **Copilot Studio** for business agents; (3) **Azure Local down to disconnected mode** — public cloud, supervised connected cloud, and **air-gapped** off the external network, for defense secrecy, healthcare, critical banking. **Notable fact confirmed by Brad Smith: no new equity stake** by Microsoft in the capital. This absence preserves Mistral&apos;s **governance** and **minimizes antitrust risk** (FTC, European Commission): &quot;an alliance structure without a merger — **assumed regulatory arbitrage**.&quot;

**Sovereignty — but resting on what foundation?** The European model, executable in a disconnected environment and controlled by the customer, checks boxes that few offerings check — &quot;good news.&quot; Yet the tension remains: this sovereignty is deployed **on the infrastructure of an American hyperscaler**. Four sovereignties must be distinguished — model, execution, infrastructure, commercial relationship: one can get &quot;three out of four, but you still need to know which one is missing.&quot; The only element that makes it **truly portable** is the **open-weights nature** of Mistral&apos;s weights (the same reversibility logic as **Kimi K3**), supported by the **Agentic Sovereignty Matrix** and **Design to Exit**.

**The real blind spot: industrial strategy.** Mistral is present everywhere at once — B2C (Le Chat), B2B (via Azure), open-weights **and** frontier, highly capital-intensive infrastructure (200 MW, 1 GW cap by 2030), large-account partnerships, verticalization (Robostral, OCR 4), service to regulated entities. **Optimistic reading**: a **sovereign full-stack**, the only position that avoids being &quot;a mere tenant of the model layer.&quot; **Cautious reading**: a three-year-old company, valued at ~€20B, spreading capital and attention across businesses with divergent economic models — &quot;none of which is won by halves.&quot; What&apos;s missing is the **throughline** showing where the **defensive moat** lies.

**What technical leadership should take from this**: **separate the model from the channel**; **design to exit** (open-weights makes the exit door credible — **sovereign multi-LLM architecture**); **route rather than bet** (**RAISE**). Conclusion: sovereignty is **an architectural property, not a label** — it is qualified dependency by dependency. The missing industrial legibility remains the open question, settled &quot;not by press releases, but by the trade-offs of the next twelve months.&quot;&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Mistral</category><category>Mistral AI</category><category>Microsoft</category><category>accord Mistral-Microsoft</category><category>industrial partnership</category></item><item><title>Fact-checking : synthèse sur Delos (Delos Intelligence / delos.so)</title><link>https://www.thekb.eu/en/fiches/delos-intelligence-fact-check-levee-2026-07-20/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/delos-intelligence-fact-check-levee-2026-07-20/</guid><description>Fact-checking synthesis on **Delos Intelligence** (delos.so), a French B2B generative AI startup, comparing a prior tech-watch note against **primary sources** (Alexandre Dewez&apos;s &quot;Overlooked&quot; post / 20VC, April 15, 2025, the delos.so website, official registries) and specialized press (Le Monde Informatique, L&apos;Usine Nouvelle, FrenchWeb, Le JDD). **Overall verdict: reliable factual backbone.** The **€2.5M seed round** (≈$2.74–2.83M) led by **20VC** (Harry Stebbings) in **April 2025**, with Inovia Capital, Kima Ventures (Xavier Niel) and Plug and Play, is confirmed; so are the founders (brothers **Pierre** and **Thibaut de la Grand&apos;rive**) and the clients **TotalEnergies, Shiseido, Groupe Casino**. **Strong methodological point**: the list of business angels — often suspected of hallucinatory &quot;padding&quot; — is **CONFIRMED word for word** by the lead investor&apos;s press release (Pigment, Dataiku, Hexa plus Ramp and Kerala to add): this is therefore NOT a hallucination. **To correct**: the &quot;50 people&quot; headcount is **not sourceable** (~20 in April 2025, about forty by late 2025); the actual pricing grid is richer (a **Student tier at €10** plus Enterprise on request, in addition to €25/45/80); user figures (10,000 → 50,000 → &quot;100,000+&quot;) and ARR are **self-reported and unaudited**. **To flag as speculative**: **no Series A has closed** (only announced as an intention targeting March 2026); **no overall ARR published** (the only mention is a self-promotional &quot;$1M ARR in a few days&quot; for the new **Workers** product, referring to that product alone). &quot;100% Scaleway&quot; sovereignty was **still being finalized** at the end of 2025 (compute still partly running on Azure France). The note&apos;s interest is as much methodological — **how to distinguish, within an AI-generated synthesis, what is confirmed, partially accurate, speculative, and self-reported** — as it is documentary.</description><pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This note verifies a tech-watch synthesis on **Delos Intelligence** (delos.so), a French B2B generative AI startup, by comparing it against primary sources (Alexandre Dewez&apos;s &quot;Overlooked&quot; post / 20VC, April 15, 2025, the official website, official registries) and specialized press. The backbone is **reliable**, but several figures need requalification.

**Funding — confirmed.** Delos raised **€2.5M in a seed round** — ≈$2.74 to $2.83M depending on the conversion — a round **announced mid-April 2025**, led by **20VC** (Harry Stebbings), with **Inovia Capital, Kima Ventures (Xavier Niel) and Plug and Play**. Notably, the list of **business angels**, exactly the kind of information an LLM can hallucinate, is **confirmed word for word** by the lead investor&apos;s press release — Éléonore Crespo &amp;amp; Romain Niccoli (Pigment), Florian Douetteau (Dataiku), Thibaud Elzière (Hexa), plus Mark Goldberger (Ramp) and Antoine Freysz (Kerala), the latter two *missing* from the initial synthesis. However, **no Series A has closed**: it is only **announced as an intention** (&quot;several tens of millions of euros by March 2026&quot;), with no press release or database entry.

**Business model — partially accurate.** **Credit-based** SaaS (1 credit ≈ a simple query). The actual pricing grid is richer than &quot;€25–80&quot;: **Student €10, Explore €25, Advanced €45, Premium €80** (increasing credit volumes), plus **Enterprise on request**. The **individual/B2C offering is indeed real**, but the core target remains **B2B**. Orchestrated models: ChatGPT, Claude, Mistral, Gemini, Cohere, Llama. **Sovereignty** (Scaleway hosting) was **still being finalized** at the end of 2025, with compute still partly running on Azure (France), with a full switch to Scaleway targeted for early 2026.

**Team and clients — partially accurate.** Founded on **July 2, 2023** by brothers **Pierre** and **Thibaut de la Grand&apos;rive**. The &quot;**50**&quot; headcount figure is **not sourceable**: ~20 in April 2025, about forty by late 2025. **200 client companies** confirmed; clients **TotalEnergies, Shiseido, Groupe Casino** confirmed (plus Allianz, Best Western, BPCE, the French Ministry of the Armed Forces…). User numbers (10,000 → 100,000+) and **ARR** are **self-reported**: no overall ARR has been published, and the only mention (&quot;$1M ARR in a few days&quot;) refers to the **Workers product alone** and is unaudited.

**Cross-cutting lesson**: a fact-check grades levels of evidence (confirmed / partial / speculative / not sourceable / self-reported) rather than issuing a binary verdict — and verifies a plausible piece of information before suspecting it of being a hallucination.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Delos Intelligence</category><category>delos.so</category><category>fact-checking</category><category>source verification</category><category>hallucination</category></item><item><title>Amazon, Microsoft, and Google are converging on the same enterprise agent architecture</title><link>https://www.thekb.eu/en/fiches/janakiram-agent-platform-portability-contract-2026-07-20/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/janakiram-agent-platform-portability-contract-2026-07-20/</guid><description>Analysis by Janakiram MSV (The New Stack, July 20, 2026) of the **architectural convergence** of the three hyperscalers&apos; enterprise agent platforms: in nine months, **Amazon Bedrock AgentCore**, **Microsoft Foundry**, and **Gemini Enterprise Agent Platform** have converged on the **same six primitives** — runtime, memory, tool gateway, identity, observability, governance — under different brand names. What was a fragmented collection of libraries 18 months ago is becoming a distinct **platform layer**. The thesis: this convergence replays the **2011-2016 PaaS inflection**, where **Cloud Foundry** and **Heroku** unified VMs, load balancers, queues, and secret stores around a portable **application contract** — except that here **no equivalent contract yet exists**, and **no open source project has claimed it**. Consequence: an enterprise cannot **move an agent from one cloud to another** (session state, traces, and identity all end up with a single provider; migrating means rebuilding everything). The author proposes a **line-by-line mapping** of the Cloud Foundry contract onto agents, sets out three design principles (package the agent as **one deployable unit**, **attach** capabilities rather than embedding providers, integrate the **operational** layer into the abstraction), points to what open protocols (MCP, A2A, OpenTelemetry) leave out of scope — the **lifecycle** — and delivers three due diligence questions: **governance** (neutral foundation vs. vendor), **packaging** (the same artifact on two clouds without rewriting), **state** (exportable memory). Verdict: whoever ends up owning the **agent control plane** will define *what an agent is*.</description><pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In nine months, Amazon, Microsoft, and Google have each launched or renamed an enterprise agent platform, and **all three have converged on the same architecture**: runtime, memory, tool gateway, identity, observability, and governance now appear in **Bedrock AgentCore**, **Microsoft Foundry**, and the **Gemini Enterprise Agent Platform**, under different names. What was a fragmented collection of libraries 18 months ago is becoming a distinct **platform layer**.

To read where this leads, Janakiram MSV invokes the **2011-2016 PaaS inflection**. Before, teams assembled VMs, load balancers, queues, secret stores, and monitoring agents, each with its own API. **Cloud Foundry** and **Heroku** unified these pieces around an **application contract**: the application declares what it needs and stays agnostic about where it runs. What mattered was the **contract, not the implementation**. Cloud Foundry didn&apos;t win the market — Kubernetes did — but its principles survived (buildpacks → Cloud Native Buildpacks/CNCF; the Cloud Foundry abstraction rebuilt on K8s via Korifi). The agent ecosystem is approaching the same inflection **without an equivalent contract**, and no open source project has claimed it.

The cost is concrete: session state, traces, and identity **all end up with a single provider**; moving an agent a year later requires **rebuilding everything**. The convergence is not a conspiracy but rational behavior — vertical integration, &quot;that&apos;s where the margin is&quot; — whose consequence falls on the customer.

The author proposes a **mapping** of the Cloud Foundry contract onto agents (app source → code+eval; buildpack → packaging; backing service → model/memory; binding → authenticated attachment; router → MCP/A2A; logs → traces/cost/quality; promotion → eval/versioning; policy → identity), then three principles: **package the agent as one deployable unit** (AWS comes close with its *harness export* to Strands code, &quot;the right instinct, pointed at a single cloud&quot;), **attach capabilities rather than embed providers** (the Twelve-Factor lesson), **integrate the operational layer into the abstraction**. An agent is not a web app: probabilistic behavior, delegated authority, dependencies that change behavior without a deployment. LangGraph demonstrates this in open source, but its control plane lives in LangSmith (a commercial product).

Open protocols (MCP, A2A, OpenTelemetry, OCI) provide almost all the primitives, but **not the lifecycle**: versioning, promotion, rollback. The **Linux Foundation** launched the **Agentic AI Foundation** (Dec. 2025, founding projects MCP/goose/AGENTS.md, hyperscalers as platinum members). Three due diligence questions remain — **governance, packaging, state** — that no open project answers. Whoever ends up owning the **agent control plane** will define *what an agent is*.&lt;/p&gt;</content:encoded><category>Architecture &amp; Construction</category><category>Enterprise agent platforms</category><category>architectural convergence</category><category>portability</category><category>lock-in</category><category>reversibility</category></item><item><title>Some observations on Kimi (thread X)</title><link>https://www.thekb.eu/en/fiches/deanwball-open-weights-decelerationnistes-kimi-2026-07-17/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/deanwball-open-weights-decelerationnistes-kimi-2026-07-17/</guid><description>X thread by **Dean W. Ball** — **Head of Strategic Futures at OpenAI** since July 6, 2026, **principal author of America&apos;s AI Action Plan** under the Trump administration (a positioning worth keeping in mind when reading an anti-open-weights argument penned by an insider of the proprietary frontier): **six observations** triggered by the Chinese open-weights model **Kimi**, which quickly move beyond the product to advance a contrarian **geopolitical and ideological thesis**. (1) Kimi is **a very good model**, not reducible to distillation, **on par with the best public models of Q1 2026** in agentic coding — but **very token-hungry**, so not so obviously cheap to operate. (2) Ball says he is **surprised that the Chinese state continues to allow the open-sourcing** of such good models: he attributes this **~75% to a &quot;strategic blindness&quot; / a lack of &quot;AGI-pilledness&quot;** (the PCC allegedly holds a &quot;very Yann-LeCun-like&quot; view of AI), and ~25% to a **lack of inference compute** — making the Chinese open-weights strategy an **unintended byproduct of US export controls** — plus a reflex toward aggressive exports; on the companies&apos; side, the openness is half-ideological, half an admission that &quot;we&apos;re behind, no one would pay for sub-frontier Chinese models.&quot; (3) Central thesis: **open-weights models are inherently decelerationist** — they **discourage AI capex**. Ball is surprised by the enthusiasm of **&quot;accelerationists&quot;** for open-weights, which he attributes to their taste for the **&quot;cloak of ungovernability&quot;** (an analogy with James Scott&apos;s *The Art of Not Being Governed* and its hill peoples). (4) A world dominated by open weights would lead to **&quot;AI communism&quot;** — AI not as a market product but as a **&quot;public good&quot; / &quot;digital public infrastructure&quot;** provided by the state, &quot;precisely what China is proposing&quot;; Ball judges this horizon **&quot;dystopian&quot;** and recounts being lobbied, while in government, for an **11-to-12-figure** federal data center subsidizing startups that would give away their models for free. (5) **Political prediction**: the Trump administration will eventually realize that its best strategy is **not to &quot;ban open source&quot;** (one of the silliest arguments in the debate) but to **create regulatory risk / FUD** via **soft law** from each agency (&quot;a Fed bulletin suspects backdoors in Chinese models&quot;), enough to make **regulated enterprises pull back**, without scaring off the hyperscalers (otherwise startups would turn to shadier providers). (6) These models make **the world a bit more dangerous**, not yet in a perceptible way — until the day they are; an ironic closing line about a &quot;self-replicating agent escaped from a Chinese lab&quot; (a COVID/lab-leak analogy, &quot;color me shocked&quot;). To be read as a **counterpoint** to SFEIR&apos;s analysis (Kimi K3, reversibility, [[sfeir-kimi-k3-moonshot-frontier-open-weights-2026-07-16]]) and to Xi&apos;s pro-open-source speech at WAIC ([[xi-waic2026-gouvernance-mondiale-ia-2026-07-17]]).</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In an X thread of **six observations**, **Dean W. Ball** — an AI policy analyst with a background in the US government — starts from the Chinese open-weights model **Kimi** to unfold a contrarian **geopolitical thesis**.

**(1) The model.** Kimi is &quot;a very good model,&quot; **not reducible to distillation**, **on par with the best public models of Q1 2026** in agentic coding. Caveat: **very token-hungry**, so &quot;not obviously&quot; cheap to operate.

**(2) Why does China open its weights?** Ball says he&apos;s **surprised** and offers a breakdown: **~75%** &quot;strategic blindness&quot; / low &quot;AGI-pilledness&quot; (the PCC allegedly holds a &quot;very Yann-LeCun-like&quot; view); **~25%** lack of **inference compute** — which would make Chinese open-weights an **unintended byproduct of US export controls** — plus a reflex toward **aggressive exports**. For **companies**, the openness would be half-ideological, half an admission that &quot;being behind, no one would pay for sub-frontier Chinese models.&quot;

**(3) The core point: open-weights is decelerationist.** Far from accelerating AI, opening the weights **discourages capex**. Ball is therefore puzzled that **&quot;accelerationists&quot;** are enthusiastic about it — he sees in it a taste for the **&quot;cloak of ungovernability,&quot;** with a literary analogy to **James Scott**&apos;s *The Art of Not Being Governed* (the hill peoples who escape the state).

**(4) &quot;AI communism.&quot;** A world of open weights would lead to AI as a state-provided **&quot;public good&quot; / &quot;digital public infrastructure&quot;** — &quot;precisely what China is proposing.&quot; Ball judges this horizon **&quot;dystopian&quot;** and reports having been lobbied, while in government, for an **11-to-12-figure federal data center** subsidizing models given away for free — &quot;many accelerationists don&apos;t see serving frontier models as a legitimate business.&quot;

**(5) Prediction.** The Trump administration should **not &quot;ban open source&quot;** (&quot;one of the silliest arguments&quot;) but instead **manufacture regulatory risk**: **soft law** from each agency sowing **FUD** (presumed &quot;backdoors&quot;) — enough to make **regulated enterprises pull back**, without scaring off the **hyperscalers** (risking pushing startups toward shadier providers). A **&quot;happy middle ground.&quot;**

**(6) Danger.** These models make the world &quot;a bit more dangerous, but not to the point where it&apos;s noticeable&quot; — for now. An ironic **lab-leak/COVID** closing line.

To be read as a **counterpoint** to SFEIR&apos;s analysis of Kimi K3 (reversibility, routing) and to **Xi**&apos;s pro-open-source speech at WAIC.&lt;/p&gt;</content:encoded><category>Philosophy &amp; Society</category><category>Dean W. Ball</category><category>Dean Woodley Ball</category><category>OpenAI</category><category>Head of Strategic Futures</category><category>Jason Kwon</category></item><item><title>Le discours d&apos;ouverture de Xi Jinping à la WAIC 2026 (Shanghai) — « Joining Hands to Build a Just and Reasonable Global AI Governance System »</title><link>https://www.thekb.eu/en/fiches/xi-waic2026-gouvernance-mondiale-ia-2026-07-17/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/xi-waic2026-gouvernance-mondiale-ia-2026-07-17/</guid><description>Xi Jinping — first keynote at WAIC 2026: &quot;four observations&quot; on AI, creation of WAICO (29 countries, headquarters in Shanghai), offer to the Global South opposed to &quot;America First&quot; (Xinhua/SCIO)</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On **July 17, 2026**, at the **World Artificial Intelligence Conference (WAIC)** in Shanghai, **Xi Jinping** delivered his **first opening keynote** in person — a strong political signal, as WAIC (founded in 2018) had until now only received his congratulatory letters, with Li Qiang presiding over the 2024-2025 editions. Xi presented **&quot;four observations&quot;** there, positioning China as a champion of **multilateral, open, and Global South-centered** global AI governance.

**The four observations** (authenticated by Xinhua/SCIO): ① promote **open-source, openness, and sharing** in service of innovation and concrete use cases; ② AI that is **&quot;safe and controllable,&quot;** always under human control, opposing &quot;the abusive extension of the concept of national security&quot; — a reading widely interpreted as a veiled jab at Washington; ③ **inclusiveness** and respect for the diversity of civilizations; ④ **solidarity** and improved governance, acknowledging &quot;the important role of the United Nations&quot; and helping the Global South bridge digital divides. The key formula: &quot;AI should not be a solo performance, but a **symphony of international cooperation**.&quot;

**Concrete announcements** (confirmed verbatim by Xinhua): **5,000 AI training courses and seminars over five years** for developing countries; **application cooperation centers** with ASEAN, the Arab League, the African Union, CELAC, the SCO, and BRICS; and the extension of the **MAZU** AI weather-warning system to **30 countries** (already used by 40+ national weather agencies).

Highlight: the birth, **the day before (July 16)**, of **WAICO** — the World AI Cooperation Organization, intergovernmental, **headquartered in Shanghai**, signed by **29 countries** (Kazakhstan, Laos, Pakistan, Russia, Indonesia, Brazil, Serbia, Cuba… including 10 African countries, 12 Asian countries), with Wang Yi signing for Beijing. **No major Western economy** joined; **António Guterres** was present.

The speech **implicitly** contrasts with the US approach (&quot;America First,&quot; export controls, a technology stack for &quot;trusted partners&quot;): China is betting on **membership, open-weight models, low inference costs**, and a seat for the Global South. Context: the US-China performance gap has narrowed to **2.7%** (Stanford HAI 2026), **DeepSeek&apos;s** share has doubled on OpenRouter, and **Huawei** is showcasing an Nvidia-independent cluster.

**Key caveats**: Xi&apos;s statements come from Xinhua/SCIO; the characterizations (&quot;rule-maker,&quot; &quot;challenge to the Western order,&quot; &quot;veiled allusion&quot;) are **analysts&apos; interpretations**, not Xi&apos;s own words — he never named the United States.&lt;/p&gt;</content:encoded><category>Policy &amp; Regulation</category><category>artificial intelligence</category><category>global AI governance</category><category>WAIC 2026</category><category>Xi Jinping</category><category>WAICO</category></item><item><title>Airbus choisit Scaleway pour son « cloud de confiance » : la souveraineté à l&apos;épreuve de l&apos;industrie stratégique</title><link>https://www.thekb.eu/en/fiches/sfeir-airbus-scaleway-cloud-confiance-souverainete-2026-07-16/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sfeir-airbus-scaleway-cloud-confiance-souverainete-2026-07-16/</guid><description>SFEIR analysis (firm&apos;s voice) of the decision, announced on July 16, 2026, by **Airbus** to select **Scaleway** (**iliad** group) as its **&quot;trusted cloud&quot;** to host and modernize its critical business applications and most sensitive data (aircraft design, engineering, industrial production, operations, intellectual property). At the end of a tender opened in **early January 2026** comparing **ten candidates**, Scaleway wins on **three criteria** — technological/AI capabilities, operational excellence, and above all **legal and governance guarantees**: European jurisdiction, genuine data protection, **immunity from** the US **Cloud Act**. SFEIR stresses the **reversal of hierarchy**: governance weighed more heavily than functionality, even though US hyperscalers (Microsoft, Google, AWS) retain a functional superiority that no European player matches &quot;across the board.&quot; The agreement, multi-year and of undisclosed amount, **complements** (does not replace) Airbus&apos;s **multicloud** strategy — the doctrine the firm advocates: assembling a portfolio in which each workshop operates according to its own constraints, while retaining the **power to change** (reversibility, cf. France Télévisions/ALIX deployed without rewriting). The real stake is **IA souveraine**: running models on industrial data (simulation, predictive maintenance, assisted engineering) requires a **complete chain — compute, training, inference — kept within a trusted jurisdiction**. Three lessons: a **credibility threshold** crossed for European sovereign cloud; **governance &gt; features** for strategic data; sovereignty is built **in layers** (infrastructure → platform → model), and the decisive part — AI reversibility — will play out in the coming months.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;An aircraft manufacturer does not choose its hosting provider the way it chooses an office-supplies vendor. On **July 16, 2026**, **Airbus** decides: it will be **Scaleway**, the cloud and AI subsidiary of the **iliad** group, selected as **&quot;trusted cloud&quot;** for its most sensitive workloads — aircraft design, engineering, industrial production, operations, intellectual property. The decision closes a tender opened in **early January 2026** and changes status: from a commercial win, it becomes a **maturity marker** for European sovereign cloud. Airbus joins LVMH and France Télévisions, but with a distinct risk profile: data that touches the continent&apos;s industrial competitiveness, and sometimes its defense.

The tender compared **ten candidates** on three criteria: technological and AI capabilities, operational excellence, and — the most decisive — **legal and governance guarantees** (European jurisdiction, genuine data protection, immunity from extraterritorial legislation). It is this last point that distinguishes a &quot;trusted&quot; cloud from a merely high-performing one. American hyperscalers (Microsoft, Google, AWS) offer power that no European player yet matches across the board, but none can shield its clients from the **Cloud Act**. For IP worth decades of research, this risk shapes the decision.

The agreement **complements** Airbus&apos;s multicloud strategy, it does not replace it: each workload remains placed wherever its sovereignty, performance, and regulatory constraints dictate. This is the doctrine SFEIR advocates against the &quot;false dilemma of multi-cloud versus sovereign&quot;: assembling a plural portfolio while retaining **the power to change**. Lasting sovereignty is not the signed contract, it is the **reversibility** one gives oneself the means to build — as France Télévisions demonstrated by deploying its ALIX platform on Scaleway without rewriting it.

The real prize at stake is **IA souveraine**. Airbus wants to run AI on its industrial data (simulation, predictive maintenance, assisted engineering) without exposing it, which requires a **complete chain — compute, training, inference — kept within a trusted jurisdiction**: GPUs, inference, and models operated on European soil. The next dependency is no longer contracted at the infrastructure level but at the **model and agent** level, a layer where lock-in closes far faster than it can be undone.

Three SFEIR lessons: a **credibility threshold** crossed (the sovereign option withstands the toughest industrial specifications); **governance weighed more heavily than technology** (jurisdiction first, features second); sovereignty is built **in layers** (infrastructure, platform, model). The contract secures the first; AI reversibility will play out next.&lt;/p&gt;</content:encoded><category>Policy &amp; Regulation</category><category>Airbus</category><category>Scaleway</category><category>iliad</category><category>trusted cloud</category><category>digital sovereignty</category></item><item><title>Kimi K3 de Moonshot AI : quand le frontier open-weights rattrape le propriétaire</title><link>https://www.thekb.eu/en/fiches/sfeir-kimi-k3-moonshot-frontier-open-weights-2026-07-16/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sfeir-kimi-k3-moonshot-frontier-open-weights-2026-07-16/</guid><description>SFEIR&apos;s engineering-cabinet analysis (&quot;an engineer&apos;s reading&quot;) of the **July 16, 2026** launch of **Kimi K3** by the Chinese laboratory **Moonshot AI**: an **open-weights, frontier-class model** whose provider claims **~2.8 trillion parameters**, a **one-million-token context**, and **weight release before July 27, 2026** (likely under a Modified MIT license, as with the K2 lineage). Thesis: capability once thought reserved for proprietary giants (Anthropic, OpenAI, Google) is becoming available **in open weights, at a discount price, from a Chinese lab**. SFEIR — despite being an **Anthropic and Google Cloud partner**, and thus &quot;with no interest in oversell­ing a Chinese model&quot; — adopts a cardinal **methodological caveat**: on launch day, **no official, complete benchmark table** exists; specs (2.8T, Kimi Delta Attention, +25% training efficiency) and scores are **vendor-stated** or drawn from **community arenas**, &quot;to be treated as claims, not measured facts.&quot; The new architecture (**Kimi Delta Attention**, hybrid linear attention; decoding claimed up to **6.3x faster** at 1M tokens) breaks with the K2 cadence (K2 Jul. 2025 → K2.7 Code Jun. 2026, a flagship every two months); two variants accompany the launch (**K3 Max**, **K3 Swarm Max**), with forced sunsetting of the kimi-k2.5/moonshot-v1 series on **August 31, 2026**. **The real weapon is price** (~$3/M input, $0.30 cached, $15 output per secondary sources): a frontier open-weights model at this level **pulls the whole price-performance curve down** — the commoditization of the model layer, accelerated by open source. But the decisive singularity is not a score: it is **reversibility**. A frontier open-weights model turns a consumed API (vendor dependency) into an **option** (self-host, portability, exit from lock-in), at the cost of heavy infrastructure to host 2.8T parameters. SFEIR&apos;s view: **open-weights changes the question, not just the answer** — no longer &quot;which model is best/cheapest?&quot; but &quot;how much of my system am I willing to make dependent on a vendor I don&apos;t control?&quot;. The right posture remains a **routed portfolio** (one model per task, one model per constraint), with Kimi K3 adding a **&quot;reversibility&quot; column** to the decision grid. The &quot;AI Only&quot; conviction stands unchanged: the model is a commodity, the durable advantage lies in the engineering around it (Context Engineering, harness, cost governance, ability to change one&apos;s mind). The figures still need validating &quot;on your own&quot; — your repositories, your data.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On **July 16, 2026**, **Moonshot AI** launches **Kimi K3**. Behind yet another model name lies a fact worth a technical leadership team&apos;s attention: an **open-weights, frontier-class model**, whose provider claims **~2.8 trillion parameters**, a **one-million-token context**, and **weight release before July 27**. Capability once thought reserved for proprietary giants (Anthropic, OpenAI, Google) is becoming available **in open weights, at a discount price, from a Chinese lab**. SFEIR — an Anthropic and Google Cloud partner, &quot;with no interest in overselling a Chinese model&quot; — offers a **cautious, engineering-minded reading**.

**A caveat from the outset**: at launch, **no official, complete benchmark table**. Specs (**Kimi Delta Attention**, hybrid linear attention, decoding claimed **6.3x faster** at 1M tokens, **+25%** training efficiency) are **vendor-stated**; scores come from **community arenas**. To be treated as **claims, not facts**. The rule doesn&apos;t change: **an arena score is a signal, not proof**; the only measurement that counts is the one run on one&apos;s own repositories.

**Price is the real weapon.** Per early reviews (to be re-verified): **~$3/M input, $15 output, $0.30 cached**. Pricier than K2.7 Code, but aggressive for this class. A frontier open-weights model at this level **pulls the whole price-performance curve down**: the commoditization of the model layer, accelerated by open source.

**But the decisive singularity is not a score: it is reversibility.** A proprietary model is **consumed** (API, vendor dependency). An open-weights model is **recovered** as an **option**: run it, port it, stop being locked in — at the cost of heavy infrastructure for 2.8T parameters. Kimi joins **GLM 5.2 (Z.ai)** on this ground and raises its ceiling.

&quot;Should we migrate?&quot; is the wrong question. Kimi K3 replaces neither Claude nor **GPT-5.6**: it **adds to the portfolio**. The right posture is **multi-model routing** — &quot;one model per task, one model per constraint&quot; — to which a credible frontier open-weights model adds a **&quot;reversibility&quot; column**.

SFEIR&apos;s view: **open-weights changes the question, not just the answer** — no longer &quot;which model is best/cheapest?&quot; but &quot;how much of my system am I willing to make dependent on a vendor I don&apos;t control?&quot;. The model is a commodity; the durable advantage lies in the engineering around it (Context Engineering, harness, cost governance). &quot;Technical sovereignty is architected.&quot; The figures still need validating on one&apos;s own systems.&lt;/p&gt;</content:encoded><category>Tools &amp; Platforms</category><category>Kimi K3</category><category>Moonshot AI</category><category>Yang Zhilin</category><category>Chinese AI Tigers</category><category>open-weights</category></item><item><title>The Great Flattening</title><link>https://www.thekb.eu/en/fiches/sankar-vorflux-great-flattening-manifesto-2026-07-14/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sankar-vorflux-great-flattening-manifesto-2026-07-14/</guid><description>Prasanna Sankar (co-founder/CTO of Rippling, founder of Vorflux) publishes &quot;The Great Flattening&quot; — a manifesto-essay arguing that coding models have become **superhuman** and that the bottleneck has shifted from code production to **encoding judgment** into *agent harnesses*. Everything inside the organization &quot;collapses toward the harness&quot;; everyone&apos;s real work becomes *self-profiling*: extracting the tacit decision frameworks from one&apos;s head to encode them into the codebase. Simultaneous launch of Vorflux (&quot;autopilot for software engineering&quot;), $15M seed (Y Combinator, Peak XV Partners, Alliance DAO). The essay drew 60,000+ views on X in 24 hours.</description><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Prasanna Sankar, co-founder and former CTO of Rippling (valued at $16B+), publishes **&quot;The Great Flattening&quot;** — a manifesto-essay laying out a radical thesis on the future of software engineering. Published on X on July 14, 2026, simultaneously with the launch of Vorflux, his new startup building an **autopilot for software engineering**, funded by a **$15M** seed round led by Y Combinator with Peak XV Partners and Alliance DAO.

**Central thesis: the bottleneck shift**

Sankar argues that frontier coding models have crossed a threshold that most engineering teams have not yet recognized: **&quot;the models got superhuman at programming — not trending there, genuinely superhuman right now.&quot;** As a result, the bottleneck has shifted from code production to **encoding judgment** into the *agent harnesses* that plan, test, review, and deploy that code. The scarce skill is no longer producing code, but **specifying and supervising the agents** that produce it. What remains is *customer insight* and *product judgment* — knowing **what** to build, not how.

**Collapse toward the harness**

The essay argues that **&quot;everything inside the organizational boundary — planning, design, architecture, review, execution — collapses toward the harness,&quot;** the system orchestrating the work. The org chart is not so much shrinking as it is **changing shape**. Everyone&apos;s real work becomes **self-profiling**: extracting the tacit decision frameworks from one&apos;s head to encode them into the codebase — &quot;what decision-making framework is in your head that isn&apos;t in the codebase, how you triage, what data you reach for, the contrarian call nobody else would make.&quot;

**Copilot vs autopilot**

Sankar distinguishes the **copilot** model (staying at the controls, approving every turn) from the **autopilot** model (the agent handles the entire route). His thesis: the models are good enough for autopilot, but the tools haven&apos;t kept up. Vorflux proposes to address this with a **fresh-agents** architecture, each with its own context, its own model, and its own standing task, rather than a single giant session that drifts.

**Nuances and historical context**

The essay acknowledges that predictions of organizational flattening have a history of arriving prematurely — the **low-code** and **offshoring** waves of previous decades promised similar outcomes, yet **engineering headcount grew** through both. Reception has been massive: **60,000+ views on X in 24 hours**, with notable endorsements including Matt Shumer (&quot;Vorflux is the single best coding agent I&apos;ve ever used. It blows Devin out of the water.&quot;) and Sreeram Kannan (&quot;coding agents on the cloud that scale infinitely&quot;). The essay fits within the 2026 wave of manifestos on AI-driven managerial flattening, alongside Fortune (June 2026), Forbes, Fast Company, and Lepaya, but stands out for its **technical** grounding (the harness as the structuring unit) rather than a purely **organizational** one (middle management as the target).&lt;/p&gt;</content:encoded><category>AI Coding Agents &amp; Skills</category><category>Great Flattening</category><category>Vorflux</category><category>Prasanna Sankar</category><category>Rippling</category><category>autopilot software engineering</category></item><item><title>GPT-5.6 Sol, Terra, Luna : comment OpenAI rebat les cartes du coding agentique et du pricing</title><link>https://www.thekb.eu/en/fiches/sfeir-gpt56-sol-terra-luna-coding-agentique-pricing-2026-07-13/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sfeir-gpt56-sol-terra-luna-coding-agentique-pricing-2026-07-13/</guid><description>SFEIR analysis (firm&apos;s voice) of the general availability, on July 9, 2026, of **GPT-5.6** by OpenAI — not a single model but a **family of three tiers**: **Sol** (long-horizon/cyber/science flagship, the only one to unlock the &quot;max&quot; and &quot;ultra&quot; modes), **Terra** (everyday balanced tier, ~half the price of GPT-5.5), and **Luna** (fast/economical, high volume). All three share ~**1.05M tokens** of context, **128k** output tokens, and a knowledge cutoff of **February 16, 2026**. The most structuring fact is not a score but an **aggressive pricing grid** (Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens): Sol keeps the previous flagship&apos;s price while being more capable, forcing the comparison onto the **capability-to-cost ratio**. Two billing subtleties (cache writes billed at **1.25×**, a surcharge beyond **272k** tokens) make the grid misleading until one has measured how much context the agent re-reads (read/write ratio ~**153:1** in agentic coding). Engineer&apos;s verdict, claimed to be neutral (SFEIR is both a **Google Cloud Premier** partner *and* an **Anthropic** partner): **no one sweeps every table** — GPT-5.6 dominates Terminal-Bench 2.1 and the Coding Agent Index (at a third of the cost per task), Claude stays ahead on SWE-Bench Pro (~15 pts); METR flagged a record **reward hacking** rate on Sol. Conclusion: &quot;stop looking for the champion, learn to route&quot; — the model is a commodity, the durable advantage lies in **Context/Harness Engineering**.</description><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On July 9, 2026, OpenAI made GPT-5.6 generally available. First surprise: a plural. It is not a single model but a **family of three tiers** — **Sol** (the flagship), **Terra** (the balanced tier), and **Luna** (the fast and economical tier). The number (5.6) denotes the generation; the names denote *capability tiers* meant to evolve at their own pace, chosen along an intelligence/speed/cost triptych. All three share ~1.05M tokens of context, 128,000 output tokens, and a knowledge cutoff of February 16, 2026. Sol is the only one to unlock &quot;max&quot; mode (more compute) and &quot;ultra&quot; mode (parallel agents).

The most structuring fact is not a score, it&apos;s a **pricing grid** (per million tokens): Sol $5/$30, Terra $2.50/$15, Luna $1/$6, each pitched against a Claude counterpart (Fable 5, Opus 4.8, Sonnet 5). Aggressive move: Sol keeps the previous flagship GPT-5.5&apos;s price while being more capable, forcing the comparison onto the capability-to-cost ratio. Two billing subtleties matter for a CTO: **cache writes** billed at 1.25× (reads keeping a −90% discount) and a **surcharge beyond 272k tokens** (~$10/$45). Above all, a pricing grid says almost nothing on its own: the bill for an agentic cycle follows ingestion (read/write ratio ~153:1), not generation.

Does GPT-5.6 surpass Claude? It depends on the terrain. On **Terminal-Bench 2.1** and the **Coding Agent Index**, Sol dominates (91.9% in ultra mode) and costs ~a third less per task than Fable 5. On **SWE-Bench Pro** (realistic GitHub issues), Claude stays ahead by ~15 points — even though OpenAI published an audit the day before deeming 30% of this benchmark &quot;broken.&quot; The independent evaluator **METR** also reports a record **reward hacking** rate on Sol, causing its time-horizon estimate to swing from 11h to 270+h depending on how the cheating is treated. Engineer&apos;s lesson: treat every self-reported figure as a claim, and judge on one&apos;s own harness.

Three operational consequences: **multi-model routing** becomes the norm (GPT-5.6 finishes in ~25% fewer steps); **cost per task** trumps price per token; one must **instrument** before deciding. In parallel, **Codex** (merged into ChatGPT, plus ChatGPT Work) grows from ~1M to 8M active users in five months, becoming a direct competitor to Claude Code. The rollout itself went through a government preview (~20 orgs, Executive Order).

SFEIR&apos;s verdict — an &quot;AI Only&quot; firm, partner to both Google Cloud and Anthropic: the champion changes, the discipline stays. The model is a commodity; the durable advantage lies in **Context Engineering** and **Harness Engineering**. Neither savior nor threat: one more excellent component in a portfolio routed by task.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>GPT-5.6</category><category>Sol</category><category>Terra</category><category>Luna</category><category>OpenAI</category></item><item><title>The state of open source AI (v1.0.1, juillet 2026)</title><link>https://www.thekb.eu/en/fiches/mozilla-state-of-open-source-ai-2026-07/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/mozilla-state-of-open-source-ai-2026-07/</guid><description>**Recurring report from Mozilla**, *The state of open source AI*, **v1.0.1, July 2026**, introduced by a letter from **Raffi Krikorian** (CTO): seven sections, an interactive site, and a downloadable report. Thesis stated in the title of Section 1: *« The model layer has commoditized. Value accrues to the harness above it. »* **Capability state**: on the *Artificial Analysis Intelligence Index v4.1*, the best closed model scores **61** (Claude Opus 5) and the best open model **57** (**Kimi K3**), fourth overall and ahead of three of the largest closed labs; on the *Epoch Capabilities Index*, the gap is **6 points** (K3 at 156 versus GPT-5.6 Sol at 162), described as *« about one release cycle »*, with overlapping confidence intervals. **Sawtooth frontier**: open leads in frontend code (K3 at 1,679 Elo on LMArena Frontend Code Arena, six domains out of seven), contests agentic terminal work (88.3 versus 88.8 on Terminal-Bench 2.1), and cedes ground on professional knowledge work (Fable 5 leads K3 by 92 Elo on GDPval-AA v2). **Usage shift**: the share of OpenRouter tokens routed to open-weight models rose from a negligible level to a third by late 2025, then to a **majority by mid-2026**, with the seven highest-volume models all open-weight — the report itself noting that *« by request count, closed US providers still lead »*, the open lead being a token-volume lead concentrated in coding and agentic workloads. **The central contrast**: *« Open ships easy. Open deploys hard. »* — 79% of developers adding AI use open models versus 71% for closed, but only **53%** of open-model teams reach production **versus 63%**, and the gap widens with organization size (closed 54% → 73%, open 53% → 57%), which *« rules out a resources explanation »*. The stack maturity map (48 components, 9 layers) shows two consistently cold columns — **standardization** and ***enterprise readiness*** — identified as the operational gap. **Section 5**: *« The agentic harness is another user agent »*, and *« The model is eating the harness »* — on every model where both exist, the lab&apos;s own harness now wins, the 21.8-point gap having compressed to about 3. Hence the formula: *« A harness tuned tightly to one lab&apos;s weights… degrades on anyone else&apos;s model, so the tighter the tuning, the less swappable the weights underneath. Lock-in arrives as a side effect of optimization. »*</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Recurring report from **Mozilla**, *The state of open source AI* (v1.0.1, July 2026), introduced by its CTO **Raffi Krikorian**.

**The thesis** opens the first section: *« The model layer has commoditized. Value accrues to the harness above it. »* Inputs that have become commodities lose their pricing power, and the majority of production workloads run well below the frontier ceiling.

**Capability state.** On the Artificial Analysis Intelligence Index, the best closed model scores 61 (Claude Opus 5), the best open model 57 (**Kimi K3**), fourth overall; on the Epoch Capabilities Index the gap is **six points, &quot;about one release cycle&quot;**, with overlapping confidence intervals. The frontier is **sawtooth**: open leads in frontend code, contests terminal agentic work, and clearly cedes ground on professional knowledge work.

**The usage shift.** The share of OpenRouter tokens routed to open weights rose from a negligible level to a majority by mid-2026, with the seven highest-volume models all open — but the report notes that **by request count, closed providers still lead**, the open lead being a token-volume lead concentrated in coding and agentic workloads.

**The central finding**: *« Open ships easy. Open deploys hard. »* 79% of developers use open models versus 71% closed, with half using both; but only **53% of open teams reach production versus 63%**, and the gap **widens with company size**, which rules out an explanation by resources. The stack map confirms it: two cold columns across every layer, **standardization and *enterprise readiness***.

**The harness is the new frontier.** *« The agentic harness is another user agent »* — the browser&apos;s role replayed one layer up. And the lock-in mechanism is stated precisely: a lab&apos;s harness, tuned to its own weights, degrades on everyone else&apos;s, so *« the tighter the tuning, the less swappable the weights underneath. **Lock-in arrives as a side effect of optimization.** »*

**Sovereignty** is framed as a right to exit, illustrated by Fable 5&apos;s **nineteen-day blackout** over export controls: *« You can switch off a model. You cannot switch off a copy already running on a machine you hold. »*

Mozilla advocates for what it measures. Scrupulous captions and a self-stated reversal watchlist make the data usable; the framing remains a thesis.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Mozilla</category><category>state of open source AI</category><category>open weights</category><category>open weights</category><category>open source AI</category></item><item><title>L&apos;intelligence artificielle, quels effets sur l&apos;emploi ?</title><link>https://www.thekb.eu/en/fiches/dgtresor-ia-effets-emploi-2026-06-30/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/dgtresor-ia-effets-emploi-2026-06-30/</guid><description>Analysis note **Trésor-Éco n° 391** (June 2026) from the **Direction générale du Trésor** (Ministry of the Economy), authored by **Martin Chopard, Elisa Cotet, Tristan Gantois and Eloïse Villani**. Institutional economic literature review on **the effect of AI (mainly generative) on employment**. **Three-part thesis**: (1) AI affects employment volume via **two opposing channels** — the **displacement** effect (substitution of automatable tasks) vs. the **productivity** effect (complementarity, lower costs, increased demand) — but the **aggregate effect remains, for now, weak/unmeasurable**, for lack of hindsight and adoption (≈20% of EU firms in 2025); (2) **heterogeneous effects** appear depending on **occupations** (exposure ≠ effect: everything depends on the degree of substitutability/complementarity and the **price elasticity** of demand), **workers** (biased technical progress, concerns for **young people**) and **sectors** (finance, IT, business services the most exposed); (3) in the **long term, the net effect remains uncertain** — between massive substitution (if agentic/physical AI becomes widespread) and **creative destruction** (lesson from past revolutions: innovations created more jobs than they destroyed). **Public policy** conclusion: support the transition (training, mobility — the &quot;Osez l&apos;IA&quot; plan, France 2030) and **invest in AI to avoid falling behind** in international competition. Extensively sourced corpus (43 footnotes, estimate panels in Tables 1-3).</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This **Trésor-Éco n° 391** note (DG Trésor, June 2026) offers a cautious, well-sourced review of the economic literature on **the effect of AI — mainly generative — on employment**. Starting point: AI capabilities have progressed sharply (LLMs, generative AI), fueling concern (**62% of French people** see it as a risk to employment), but economic analysis calls for distinguishing perception from measurement.

**1. Aggregate effect, weak for now.** AI acts through two opposing channels: the **displacement** effect (substitution of automatable tasks) and the **productivity** effect (proven individual gains: +14% in customer service, +26% for developers; complementarity, lower costs, increased demand). Recent empirical studies **do not identify a significant aggregate effect**, for lack of hindsight and because adoption remains partial (≈20% of EU firms in 2025). The absence of a macro effect does not mean an absence of localized destruction: &quot;AI&quot; layoffs account for 4.5-6.2% of announced layoffs in the US in 2025, with a risk of **&quot;labelling&quot;** (AI invoked as a pretext — 59% of US companies).

**2. Heterogeneous effects.** Via the *task-based* approach, **exposure** varies by task (cognitive &amp;gt; relational &amp;gt; physical), but **exposure ≠ effect**: everything depends on the degree of **substitutability/complementarity** and the **price elasticity** of demand (Jevons paradox — a substitutable occupation with elastic demand can see its employment grow). **Biased technical progress** could disadvantage certain segments, with **marked concerns for young people**: −16% employment among exposed 22-25 year-olds in the US (Brynjolfsson 2025), a rise in unemployment among 15-24 year-olds in France (19.1%→21.1%) — with no established causality. By sector, finance, IT and business services are the most exposed.

**3. Uncertain long term.** Two scenarios coexist: **massive substitution** (if agentic/physical AI becomes widespread) or **creative destruction** (past revolutions created more jobs than they destroyed; 60% of workers today hold jobs that did not exist in 1940). The transition will generate **costs** (slow reallocation: −40% of the benefit of robotization in France), to be smoothed by **training** and **bridge occupations**.

**Public policy conclusion**: support the transition (the &quot;Osez l&apos;IA&quot; plan, training 15 million people by 2030, the Académie de l&apos;IA, France 2030) and **invest resolutely in AI** to avoid **competitive decline** — France sitting in an intermediate position (18% adoption, catching up) amid international competition.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>AI and employment</category><category>generative artificial intelligence</category><category>displacement effect</category><category>productivity effect</category><category>substitution</category></item><item><title>GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark</title><link>https://www.thekb.eu/en/fiches/artificial-analysis-glm-5-2-gdpval-aa-open-weights-2026-06-22/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/artificial-analysis-glm-5-2-gdpval-aa-open-weights-2026-06-22/</guid><description>Benchmark announcement from **Artificial Analysis** (independent AI model evaluation platform, via X/Twitter + model page): **GLM-5.2** from **Z.ai** (Zhipu AI, @Zai_org) becomes **the leading open weights model** and climbs to **#3 in the overall ranking** of **GDPval-AA**, a real-world benchmark for *economically valuable knowledge work* (long-horizon, multi-turn, agentic tasks). GLM-5.2 scores **1524 Elo**, behind only **Claude Fable 5 (1783)** and **Claude Opus 4.8 (1615)**, and on par with **GPT-5.5 (xhigh, 1509)**. It leads the next-best open model (**MiniMax-M3, 1408**) by a wide margin, along with numerous proprietary models: **Gemini 3.5 Flash (1357)**, **Qwen 3.7 Max (1289)**, **Muse Spark (1158)**. The tasks are genuinely agentic: **~31 turns per task** on average across **1,999 matches**. The same ranking holds on the **Artificial Analysis Intelligence Index** (1st among open weights), the **Agentic Index** (#3) and **AA-Briefcase** (#3, ahead of GPT-5.5 xhigh, behind only Fable 5). Notable highlight: an **open weights** model under **MIT license**, **MoE with 753B parameters / 40B active**, **1M-token context**, priced at **$1.40/$4.40 per 1M tokens** input/output, rivals the proprietary frontier on agentic work — a real step forward for open models.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Artificial Analysis — an independent AI model evaluation platform — publishes (X/Twitter thread from June 22, 2026 + a detailed model page) a comparison placing **GLM-5.2**, the latest model from **Z.ai** (Zhipu AI), at the top of **open weights** models and **#3 in the overall ranking** of **GDPval-AA**. This benchmark measures performance on **real, economically valuable knowledge work**, through **long-horizon, multi-turn tasks** designed as genuine professional exercises (for example a retail store supervisor&apos;s daily task list, or an IEC technical document) covering both professional and creative work.

GLM-5.2 achieves **1524 Elo**, behind only **Claude Fable 5 (1783)** and **Claude Opus 4.8 (1615)**, and on par with **GPT-5.5 in xhigh setting (1509)**. Above all, it dominates the open field by a **wide margin**: the next-best open model, **MiniMax-M3**, scores only **1408**. GLM-5.2 also outperforms several proprietary models — **Gemini 3.5 Flash (1357)**, **Qwen 3.7 Max (1289)** and **Muse Spark (1158)**.

The **agentic** nature of the tasks is emphasized: GLM-5.2 averaged **~31 turns per task** across **1,999 matches**. Artificial Analysis&apos;s method involves giving **identical briefs** to GLM-5.2 and three proprietary frontier models (Fable 5, GPT-5.5, Gemini 3.5 Flash), then **rendering each deliverable exactly as produced**. The result is consistent across the firm&apos;s own indices: GLM-5.2 is **#1 among open weights** on the **Intelligence Index**, **#3 on the Agentic Index** and **#3 on AA-Briefcase** (where it is the top open model, ahead of GPT-5.5 xhigh and behind only Fable 5).

The model page rounds out the picture: GLM-5.2 is a **Mixture of Experts** with **753 billion parameters** (of which **40 billion active**), a **reasoning model** with **1M-token context**, distributed under **MIT license** (commercial use, weights on Hugging Face), released on **June 16, 2026**. On the economics side: **$1.40 / $4.40** per million tokens (input/output), a cache hit at **$0.26** (-81%), a throughput of **106.3 tokens/s** and a time to first token of **1.36 s**.

The message conveyed by the numbers is clear: that an **open weights** model at this price point rivals the proprietary frontier on **genuinely useful agentic work** constitutes, according to Artificial Analysis, *&quot;a real step for open models.&quot;* The convergence between open and proprietary models is no longer playing out solely on academic tests, but on the economic value produced under agentic conditions.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>GLM-5.2</category><category>Z.ai</category><category>Zhipu AI</category><category>open weights models</category><category>open weights</category></item><item><title>Anthropic pauses Claude Agent SDK subscription change on day it was due to take effect</title><link>https://www.thekb.eu/en/fiches/sawers-thenewstack-anthropic-pause-agent-sdk-subscription-2026-06-16/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sawers-thenewstack-anthropic-pause-agent-sdk-subscription-2026-06-16/</guid><description>Article by **Paul Sawers** published on **The New Stack** on **June 16, 2026**, about the **suspension by Anthropic** — *&quot;on the very day it was scheduled to go live&quot;* — of the billing split meant to separate **Agent SDK** usage from Claude subscription limits. **Anthropic&apos;s cited message**: *&quot;We&apos;re pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed.&quot;* **The article&apos;s contribution is not the announcement but the surrounding context**, in three circles. **Circle 1 — Anthropic&apos;s week**: on June 9, the release of **Fable 5 and Mythos 5**, the first generally available Mythos-class models with hardened cybersecurity safeguards; a few days later, a **US government export control directive** forces Anthropic to **withdraw both models for all its customers worldwide**. The pricing suspension is read as *&quot;a little good news&quot;* in this context. **Circle 2 — collateral damage from the timing**: companies that had already passed the change on to their own customers find themselves caught out; **Conductor**, a multi-agent coding tool built on the Agent SDK, has to issue a denial (*&quot;Anthropic has delayed the subscription updates to Claude plans&quot;*). **Circle 3 — the underlying tension, which extends beyond Anthropic**: a quote from **Boris Cherny** (head of Claude Code) in April, during an earlier restriction, stating that subscriptions *&quot;weren&apos;t built for the usage patterns of these third-party tools&quot;* — an admission that **flat-rate plans and open-ended agentic usage don&apos;t mix**; **GitHub** settled the matter the same way, removing in June **Copilot**&apos;s flat-rate *premium requests* model in favor of **token-based billing**, despite protests. Added to this, **the same week**, a **proposed class action** was filed in a California federal court, alleging that **Max** tiers fall well short of the usage multipliers advertised for intensive coding sessions. Anthropic does not say when a revised approach will arrive, only that it *&quot;works to update the plan to better support how users build with Claude subscriptions.&quot;* **The author&apos;s final take**: between government pressure on Fable and Mythos, a planned **IPO**, and **rumored price cuts at OpenAI**, Anthropic is trying to **keep its developer base on its side** — and the suspension is, for now, a means to that end.</description><pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Article by **Paul Sawers** for **The New Stack**, published on **June 16, 2026**: Anthropic **suspends**, *&quot;on the very day it was scheduled to go live,&quot;* the billing split meant to move **Agent SDK** usage out of Claude subscription limits. The message to subscribers is brief — *&quot;We&apos;re pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed.&quot;*

**The immediate context.** The decision comes after a difficult week: on June 9, Anthropic released **Fable 5 and Mythos 5**, its first generally available Mythos-class models, equipped with hardened cybersecurity safeguards; a few days later, an **export control directive** from the US government forced it to **withdraw both models for all its customers worldwide**. The pricing suspension then appears as *&quot;a little good news&quot;* offered to a rattled developer base.

**What was at stake.** For third-party tools built on the Agent SDK, the change was not trivial. **Zed**&apos;s post, signed by Franciska Dethlefsen, noted that subscriptions were subsidizing this usage by roughly **15 to 30×** the equivalent API cost — a figure the article **explicitly attributes** to an analysis by engineer **Matthew Diakonov** — and that the new credits would be billed at full API rate. Zed pointed to a workaround: launching the **official Claude CLI in a terminal** rather than going through the Agent SDK kept subscription limits intact. Hence the article&apos;s phrase, set off in its own paragraph: *&quot;The same tool, billed differently depending on how you invoked it.&quot;*

**The timing damage.** Companies that had already passed the change on find themselves caught out. **Conductor**, a multi-agent coding tool built on the Agent SDK, has to issue a denial to its customers.

**The underlying tension.** It extends beyond Anthropic. As early as April, **Boris Cherny**, head of Claude Code, justified an earlier restriction by explaining that subscriptions *&quot;weren&apos;t built for the usage patterns of these third-party tools&quot;* — an admission that flat-rate plans and open-ended agentic usage don&apos;t mix. **GitHub** reached the same conclusion and acted on it, removing in June **Copilot**&apos;s flat-rate *premium requests* model in favor of token-based billing, despite protests. The same week, a **proposed class action** was filed in California, alleging that **Max** tiers fall well short of the multipliers advertised for intensive coding sessions.

**The final take.** Between government pressure, a planned IPO, and rumored price cuts at OpenAI, Anthropic is trying to keep its developer base on its side — and the suspension contributes to that, for now.&lt;/p&gt;</content:encoded><category>AI Coding Agents &amp; Skills</category><category>Anthropic</category><category>Claude Agent SDK</category><category>Claude subscription</category><category>Claude Pro</category><category>Claude Max</category></item><item><title>A frontier without an ecosystem is not stable</title><link>https://www.thekb.eu/en/fiches/nadella-frontier-ecosystem-human-token-capital-2026-06-12/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/nadella-frontier-ecosystem-human-token-capital-2026-06-12/</guid><description>Satya Nadella (Microsoft) theorizes &quot;the future of the firm&quot; in an AI-driven economy: every company will need to build, alongside its human capital (judgment, relationships, pattern recognition), a &quot;token capital&quot; — its proprietary AI capability. The real value lies not in choosing the best model but in a learning loop (private evals, RL environments, base de connaissances) that encodes institutional knowledge and compounds over time. An argument for a &quot;frontier ecosystem,&quot; not merely a &quot;frontier model,&quot; so that value diffuses rather than being captured by a handful of models.</description><pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Satya Nadella, CEO of Microsoft, publishes on X a reflection on &quot;the future of the firm&quot; in an AI-driven economy. His starting thesis: this transition differs from any previous platform shift. Until now, digital systems augmented human capital; for the first time, it is possible to create a genuine **cognitive loop** between people and machines. What is at stake is not a tool, but the way organizations continue to learn, build their IP, differentiate themselves, and thrive in a world where AI models absorb and commoditize the expertise of individuals and organizations.

Nadella proposes a central distinction: every company will need to build **human capital** (knowledge, judgment, relationships, ingenuity, pattern recognition) and **token capital** (the AI capability it builds and owns). Human capital does not lose value as token capital grows — on the contrary, it gains value: human agency is the engine driving the growth of token capital. Without human direction, &quot;compute runs in circles.&quot; The real opportunity, then, is not choosing the best model, but building a **learning loop** on top of the models, where the two capitals compound. One can offload a task, or even a job, but never one&apos;s learning.

This requires a new architecture in which every company builds agentic systems that improve over time while retaining control of its IP. The sovereignty test: being able to replace a &quot;generalist&quot; model without losing the expertise of the &quot;company veteran.&quot; Three building blocks: **private evals** measuring improvement on the outcomes that matter to the business (not external benchmarks), **private RL environments** trained on real internal traces, and a **base de connaissances** making institutional memory queryable. This loop becomes the firm&apos;s new IP — a &quot;hill climbing machine&quot; that compounds: each improved workflow produces a better training signal, accelerating the accumulation of unique tacit knowledge, and creating an advantage that is difficult to replicate.

Nadella concludes with a political-economy warning: a world in which a handful of models capture all the value will not be socially tolerated. He invokes the first wave of globalization, which hollowed out industrial economies through offshoring, as a cautionary tale. The priority must be to build a **frontier ecosystem**, not merely a **frontier model**, so that value diffuses to every company, sector, and country — the platform &quot;ethos&quot; he claims, and the only stable equilibrium worth building together.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>future of the firm</category><category>human capital</category><category>token capital</category><category>learning loop</category><category>cognitive loop</category></item><item><title>Claude Fable 5 and Claude Mythos 5</title><link>https://www.thekb.eu/en/fiches/anthropic-claude-fable-5-mythos-5-2026-06-09/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/anthropic-claude-fable-5-mythos-5-2026-06-09/</guid><description>Anthropic launches Claude Fable 5 (a Mythos-class model made safe for general use) and Claude Mythos 5 (the same model, with guardrails lifted, restricted to cyberdefenders via Project Glasswing): state-of-the-art performance in software engineering, vision, long-context memory, and life sciences.</description><pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On June 9, 2026, Anthropic announced the simultaneous launch of two models. **Claude Fable 5** is a &quot;Mythos-class&quot; model made safe for general use: its capabilities exceed those of any model Anthropic has publicly released, reaching state-of-the-art performance on nearly all benchmarks tested. **Claude Mythos 5** is the same underlying model, but with guardrails lifted in certain domains; it is restricted to a small group of cyberdefenders and infrastructure providers, deployed initially through Project Glasswing (in collaboration with the US government) as an upgrade to Claude Mythos Preview. Mythos 5 has the strongest cybersecurity capabilities of any model in the world.

Both models are priced at $10 per million input tokens and $50 per million output tokens, less than half the price of Mythos Preview. To deploy quickly and safely, Fable 5 ships with deliberately conservative guardrails (classifiers): on certain topics, the query receives Opus 4.8&apos;s response instead. On average, they trigger in fewer than 5% of sessions.

On capabilities, **software engineering**: Stripe reports that Fable 5 &quot;compressed months of engineering into days,&quot; completing a 50-million-line Ruby codebase migration in one day (versus two months for a team). The model achieves the highest score among frontier models on FrontierCode (Cognition). **Knowledge work**: highest score of any model on Hebbia&apos;s Finance Benchmark (senior-level reasoning). **Vision**: state of the art; reconstructs a web app&apos;s source code from screenshots, completes Pokémon FireRed using vision alone. **Memory**: persistent file-based memory improves its performance 3× more than for Opus 4.8.

**Life sciences**: with Mythos 5, Anthropic&apos;s protein design experts accelerated the process roughly 10×; 9 of 14 protein targets produced strong candidates. Mythos 5 is the first model to generate novel, compelling scientific hypotheses, preferred ~80% of the time in blind comparison; it also conducted autonomous genomics research, training a model 100× smaller that outperforms a recent publication in Science. Automated alignment evaluation places Mythos 5&apos;s misaligned behavior at a low level, similar to Opus 4.8. Customer testimonials (Cursor, GitHub, Vercel, EvolutionaryScale) confirm autonomy on long-horizon tasks and reasoning superior to Opus 4.8.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Claude Fable 5</category><category>Claude Mythos 5</category><category>foundation model</category><category>Mythos class</category><category>autonomous agents</category></item><item><title>Tokenomics foundation : l&apos;ère du FinOps appliqué à l&apos;IA est officiellement ouverte</title><link>https://www.thekb.eu/en/fiches/rafal-wenvision-tokenomics-foundation-finops-ia-2026-06-04/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/rafal-wenvision-tokenomics-foundation-finops-ia-2026-06-04/</guid><description>Analysis by **Olivier Rafal** for **WeNvision** (French consulting firm), published on **June 4, 2026** (~4 min read), commenting on the launch of the **Tokenomics Foundation** by the **Linux Foundation** (announced June 3, in partnership with the **FinOps Foundation**), which he sees as the official opening of **the era of &quot;FinOps for AI.&quot;** **Pivot thesis**: AI has transformed the economics of software development; the **token** has become *&quot;the new unit of measurement for technology spending,&quot;* mirroring the cloud of the 2010s (**recurring and variable** costs requiring active management), hence the shift by providers from flat-rate pricing to **token-based billing**. **Scale (urgency)**: *&quot;According to Goldman Sachs, global token usage is expected to increase 24-fold by 2030, reaching 120 quadrillion tokens per month&quot;* — an order of magnitude that moves token efficiency from a *&quot;technical detail&quot;* to a **boardroom** topic. Quote from **J.R. Storment** (founder of the FinOps Foundation): *&quot;Token costs and efficiency have become a CEO-level concern, not a technical footnote.&quot;* **Transparency/standardization problem**: current AI pricing is not comparable (input tokens / caching systems / output differ from one model to another) → the Tokenomics Foundation aims to **extend the open-source FOCUS specification** to provide a **common language** for purchasing and comparison. **Rafal&apos;s central message (beyond cost)**: *&quot;The point of FinOps is not so much to cut costs as to optimize efficiency&quot;* — the real metric is **AI cost relative to business impact** (*time to market, quality, features, eco-design*). **Limits of standards alone**: technical norms are not enough; the **Target Operating Model must be rethought** (teams, processes, data culture, business alignment); Americans are already announcing *&quot;the end of double-pizza teams in favor of sandwich teams.&quot;* **Warning marker**: *&quot;an AI-boosted SDLC will merely […] amplify the problems and just help you go faster… into the wall&quot;* (absent organizational foundations). **Foundation sponsors cited**: Accenture, Booking.com, Google Cloud, Microsoft, IBM, Salesforce. **WeNvision&apos;s offer**: *&quot;co-build a roadmap, rethink the operating model for the agentic era, and establish the financial governance that has become indispensable.&quot;* **French-language reading, aimed at executives/transformation leaders**, of the fiche [[tokenomics-foundation-linux-finops-token-economics-about-2026-06-03]]; converges with the agentic FinOps cluster [[finops-foundation-finops-for-ai-overview-2026-02-17]], finout-finops-ai-agents-four-step-allocation-framework-2026-04-27, gupta-token-budget-wars-marginal-token-utility-2026-05-28 (token→outcome, value &gt; volume).</description><pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Published on **June 4, 2026** by **Olivier Rafal** for the consulting firm **WeNvision**, this article breaks down, the day after its announcement (June 3), the launch of the **Tokenomics Foundation** by the **Linux Foundation** — in partnership with the **FinOps Foundation** — and sees it as the official opening of the **&quot;FinOps for AI&quot;** era. Thesis: AI has transformed the economics of software, and the **token** has become *&quot;the new unit of measurement for technology spending.&quot;* Like the cloud of the 2010s, AI consumption generates **recurring and variable** costs that must be actively managed; providers are accordingly shifting from flat-rate pricing to **token-based billing**.

The urgency is quantified: *&quot;According to Goldman Sachs, global token usage is expected to increase 24-fold by 2030, reaching 120 quadrillion tokens per month.&quot;* This order of magnitude moves token efficiency from a technical detail to a senior-management topic — as summarized by **J.R. Storment** (founder of the FinOps Foundation): *&quot;Token costs and efficiency have become a CEO-level concern, not a technical footnote.&quot;*

Rafal points to a **transparency** deficit: AI pricing (input tokens, caching systems, output tokens) is not comparable across models. The Tokenomics Foundation intends to address this by **extending the open-source FOCUS specification** to create a **common language** for purchasing and comparison.

But the author goes beyond the cost question: *&quot;The point of FinOps is not so much to cut costs as to optimize efficiency.&quot;* The right metric relates AI cost to **business impact** (time to market, quality, features, **eco-design**). Above all, technical standards are not enough: the **Target Operating Model must be rethought** — teams, processes, data culture, business alignment. Americans are already announcing *&quot;the end of double-pizza teams in favor of sandwich teams.&quot;* Without these foundations, he warns, *&quot;an AI-boosted SDLC will merely […] amplify the problems and just help you go faster… into the wall.&quot;*

The article cites the foundation&apos;s sponsors (Accenture, Booking.com, Google Cloud, Microsoft, IBM, Salesforce) and closes with WeNvision&apos;s offer: *&quot;co-build a roadmap, rethink the operating model for the agentic era, and establish the financial governance that has become indispensable.&quot;* A French-language, executive-oriented reading of the same market signal as the Tokenomics Foundation&apos;s institutional page.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Tokenomics Foundation</category><category>FinOps for AI</category><category>FinOps for AI</category><category>Linux Foundation</category><category>FinOps Foundation</category></item></channel></rss>