<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>thekb.eu — Products &amp; Services</title><description>Products &amp; Services · High-fidelity tech watch — AI, coding agents, SDLC</description><link>https://www.thekb.eu/</link><language>en</language><item><title>Designing AI with character: what we learned building Berd</title><link>https://www.thekb.eu/en/fiches/block-berd-caractere-agents-open-source-2026-08-18/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/block-berd-caractere-agents-open-source-2026-08-18/</guid><description>Corporate blog post from **Block** (`block.xyz/inside`), unsigned — the displayed author is **&quot;Block&quot;** —, published on **August 18, 2026**, ~930 words, announcing **the open-sourcing of Berd**, Block&apos;s internal desktop application for working with agents, and laying out the design thesis that guided it: giving agents character *&quot;not only through roles, instructions, skills, and tools, but through distinctive visual identities&quot;* — hence the in-house animated characters, the *&quot;Gloopies&quot;*. The post starts from an observation of fragmentation (*&quot;The technology was powerful, but the experience around it was fragmented&quot;*) and a precisely named interface problem: *&quot;the product gives people little sense of how the agent is configured, which context and tools are available to it, and how it differs from another agent&quot;*. Two structuring contributions. **(A) A three-tier articulation**: **goose** remains the framework and *runtime* that holds the agent loop; **Berd** is the desktop client (projects, context, sessions, agents, configuration); the two communicate via the **Agent Client Protocol**. **Buzz** is designated as the follow-up, for when solo work becomes collaborative (*&quot;Start alone, then go multiplayer&quot;*). **(B) Six requirements handed off to Buzz**, stated as a takeaway: *&quot;private space, durable context, recognizable agent identities, reusable skills, visible configuration, and clearer visibility into an agent&apos;s configured context, tools, and capabilities&quot;* — a grid directly reusable for evaluating an agent client. The text itself distinguishes identity from capability: *&quot;The avatars make the agent recognizable. Its role, skills, and tools make it useful.&quot;* No usage figures are produced and no license is named for the open-sourcing.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Corporate blog post from **Block** (`block.xyz/inside`), **unsigned**, published on **August 18, 2026**, announcing **the open-sourcing of Berd** and laying out the design thesis that guided it.

**What Berd is.** *&quot;Berd is a desktop application our teams use to work with AI agents across projects, skills, tools, and models.&quot;* Born of an internal problem: Block had access to capable agents — **goose**, **Claude Code**, **Codex** — but each imposed *&quot;different interfaces, configuration systems, and ways of managing context&quot;*. The conclusion drawn: *&quot;we didn&apos;t need another model or agent harness, **we needed a consistent environment around them**&quot;*. Berd brings together conversations, files, folders, instructions, agents and skills around **persistent projects**, to stop rebuilding context for every task.

**The design thesis.** Give agents **character** — not only through roles, instructions, skills and tools, but through **distinct visual identities**, including a collection of animated characters, the *&quot;Gloopies&quot;*. The problem invoked is that of the empty prompt box: *&quot;the product gives people little sense of how the agent is configured, which context and tools are available to it, and how it differs from another agent&quot;*. The post places the approach in the lineage of **Square** and **Cash App** — bringing design where the category had none. **But the problem stated is a configuration-legibility problem, and the avatar solves distinguishability**; the text acknowledges this in one line it does not develop: *&quot;The avatars make the agent recognizable. **Its role, skills, and tools make it useful.**&quot;*

**The architecture.** Berd descends from **goose**, the open source agent framework launched by Block in **January 2025**, contributed to the **Agentic AI Foundation** (Linux Foundation, December 2025) alongside **MCP** and **AGENTS.md**. Explicit division: *&quot;goose remains the open agent framework and runtime. Berd is a desktop application built around it. **Berd connects to goose through the Agent Client Protocol.**&quot;* goose holds the agent loop, Berd holds the experience.

**The sequel is Buzz.** Berd served to explore **solo** work; *&quot;But work rarely stays private&quot;*. What Berd showed — *&quot;private space, durable context, recognizable agent identities, reusable skills, visible configuration&quot;* — will feed **Buzz**, the shared human+agent space. *&quot;Start alone, then go multiplayer.&quot;*

**Caveats.** **No figures, no user testing, no license named**; one isolated overreach (*&quot;create custom agents to do any task they want&quot;*); and a post whose title announces a retrospective while keeping the product in the present tense — **Berd is not declared deprecated, but the roadmap points to Buzz**.&lt;/p&gt;</content:encoded><category>Tools &amp; Platforms</category><category>Berd</category><category>Block</category><category>open source</category><category>open-sourcing</category><category>desktop application</category></item><item><title>Block explores how to price AI</title><link>https://www.thekb.eu/en/fiches/paymentsdive-block-dorsey-pricing-ia-2026-08-06/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/paymentsdive-block-dorsey-pricing-ia-2026-08-06/</guid><description>Trade-press brief (**Payments Dive**, *Dive Brief* format, **August 6, 2026**) covering **Block**&apos;s quarterly earnings release: the company has already rolled out several AI tools to its customers — **Moneybot** (Cash App) and **Managerbot** (Square) — and has not yet decided how to charge for them. **Jack Dorsey** on the analyst call: *&quot;We&apos;re in a fortunate position where we can experiment with a number of models, and then choose the right one that&apos;s going to align all of our incentives with our customers.&quot;* **The financial backdrop illuminates that stance.** Six months earlier, Block had laid off roughly **4,000 people, about 40% of its workforce**, in a reorganization explicitly framed around AI. In Q2 2026: gross profit **up 25% to $3.2B**, revenue **up 10% to $6.62B**, but **net income at $89M, down 83%** year over year due to severance costs closing out the restructuring; 2026 guidance was raised. AI&apos;s value, then, is being captured through the cost structure before it is captured through price. **The heaviest fact sits in the middle of the brief**, drawn from the shareholder letter: *&quot;Starting in June, agentic AI helped write and review nearly all of our production code changes&quot;* — writing **and** reviewing nearly all production code changes, at a publicly traded payments company, six months after cutting 40% of the workforce. A self-reported claim to investors, with no definition of *&quot;nearly all&quot;* or of what *&quot;review&quot;* covers. **The tooling**: **Goose**, an internal system built two years earlier, described as model-agnostic (it plugs in different commercial models for employees); **Buzz**, launched the previous month for *&quot;agent collaboration, communication, and code repositories.&quot;* **On the customer side**: Moneybot monitors Cash App user activity and surfaces accounts, balances, and transactions — over **one million weekly active accounts**; Managerbot runs automated marketing, margin analysis, and suggests *&quot;operational fixes&quot;* to Square merchants. **Evercore ISI** analysts list four monetization paths — SaaS bundles, direct subscriptions, enterprise offerings, usage-based pricing — **none of them tied to outcomes**. Stated order of priority: **product quality → distribution → adoption → pricing model**. Two distribution facts round out the picture: Square is rolling into **Google Maps** with a *&quot;conversational AI experience,&quot;* described as *&quot;the first step in a broader partnership between Square and Google&quot;*; and the **Tags** payment device (keychain and NFC chip wands) shows **three million people on the waitlist**. Analyst quotes: William Blair (*&quot;Block epitomizes the secular shift toward tech-forward digital finance firms&quot;*) and Bank of America on the *&quot;post-reset operating model.&quot;*</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;A **Payments Dive** brief from **August 6, 2026** on **Block**&apos;s quarterly earnings, the parent of **Cash App**, **Square**, and **Afterpay**.

**The stated topic.** Block has rolled out several AI tools to its customers and **has not yet decided how to charge for them**. **Jack Dorsey**, on the analyst call: *&quot;We&apos;re in a fortunate position where we can experiment with a number of models, and then choose the right one that&apos;s going to align all of our incentives with our customers.&quot;* The company is consulting Square merchants on their needs. **Evercore ISI** analysts list four possible paths — SaaS bundles, direct subscriptions, enterprise offerings, usage-based pricing — noting that Block is prioritizing *&quot;product quality, distribution, and adoption&quot;* first.

**The real story, left for the reader to piece together.** **Six months earlier**, Block laid off **roughly 4,000 people, ~40% of its workforce**, in a reorganization centered on AI. In Q2 2026, **gross profit rose 25% to $3.2B** while revenue climbed only 10% to $6.62B; **net income fell to $89M, down 83%**, weighed down by severance costs; **2026 guidance was raised**. Not a single dollar of AI has been billed to customers: the value has already been captured **through the cost structure**. The &quot;fortunate position&quot; that lets Dorsey take his time on pricing is exactly what the workforce cut bought.

**The buried number.** In the shareholder letter: *&quot;Starting in June, agentic AI helped write and review nearly all of our production code changes.&quot;* Writing **and** reviewing nearly all production code changes, at a publicly traded payments company. A self-reported claim to investors, with no definition of *&quot;nearly all&quot;* or of *&quot;review.&quot;*

**The tooling.** **Goose**, an &quot;agnostic&quot; internal system built two years earlier, plugging in several commercial models for employees. **Buzz**, launched the previous month, for agent collaboration, communication, and code repositories. On the customer side, **Moneybot** (Cash App) tracks activity, surfaces accounts, balances, and transactions, and has passed **one million weekly active accounts**; **Managerbot** runs automated marketing and margin analysis for Square merchants.

**Two distribution facts.** **Square is entering Google Maps** with a conversational discovery-and-ordering experience, *&quot;the first step of a broader partnership&quot;* with Google. And the **Tags** device (NFC) shows **three million people on the waitlist**.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Block</category><category>Jack Dorsey</category><category>Cash App</category><category>Square</category><category>Afterpay</category></item><item><title>Introducing Muse Code and Muse Spark 1.2</title><link>https://www.thekb.eu/en/fiches/meta-muse-code-muse-spark-1-2-2026-08-05/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/meta-muse-code-muse-spark-1-2-2026-08-05/</guid><description>Announcement from **Meta AI Research** published on **August 5, 2026** (stated reading time: 4 minutes, no individual byline): **Muse Code** in beta, *« a terminal coding agent »*, and the model that powers it, **Muse Spark 1.2**. Meta itself frames the launch: *« This marks our next step toward the frontier, with larger and much more capable models on the way. »* **Three architectural elements on the harness side.** **Asynchronous background agents** that *« remain active throughout each session, rather than being spawned for individual tasks »*, avoiding redundant information gathering and reducing the need for steering. A **local event log** where *« every model call, tool run, approval, and edit is appended »*, making the runtime a system that is *« replay-exact and restart-safe »*, able to resume exactly where it left off after a crash. And **three skills shipped out of the box**: `/plan` (turns a task into a plan submitted for approval), **`/grill`** (stress-tests the plan *« until it holds up »*), and `/goal`. **On the model side**, Meta claims **model-harness co-training** (*« to maximize harness compatibility »*, with harness trajectories sampled via rejection sampling and recipe optimizations for goals, compaction, and sub-agents), **long-horizon** training (whole-repo generation, end-to-end projects, self-research, with planning, goal conditioning, and context compaction), and a **self-improvement loop** where Muse Spark 1.1 generates the environments and instruction templates and then grades candidate solutions, producing a training set for the 1.2. **What the published charts show**, without the text commenting on it: the four comparisons — Terminal-Bench 2.1, DeepSWE 1.1, an internal Meta benchmark, and the GPU kernel optimization case study — place **Muse Spark 1.2 behind Opus 5 in all four cases**, including on Meta&apos;s own proprietary benchmark (70.6% versus 79.4%) and on the case study, where the model finishes fourth out of six (+68.7% versus +74.0%). **A reading caution on the version gain**: on the two public benchmarks, 1.1 is measured with `mini-swe-agent` and 1.2 with Muse Code, so the 6.7-point gap conflates model and harness. On the internal benchmark, the only comparison where no harness is mentioned, the 1.1 → 1.2 gap drops to **2.3 points**.</description><pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Announcement from **Meta AI Research** dated **August 5, 2026**: **Muse Code** in beta, a terminal coding agent, and **Muse Spark 1.2**, the model that powers it. Meta itself frames the launch — *« our next step toward the frontier, with larger and much more capable models on the way »*.

**On the harness side, three decisions.** **Asynchronous background agents** that *« remain active throughout each session, rather than being spawned for individual tasks »*, avoiding redundant information gathering and deciding for themselves when to escalate to the main agent. A **local event log** recording every model call, tool run, approval, and edit, which makes the runtime *« replay-exact and restart-safe »*: after a crash, the agent resumes exactly where it left off. And three **skills shipped out of the box**: `/plan` (a plan submitted for approval), **`/grill`** (stress-tests the plan until it holds up), and `/goal`.

**On the model side**, Meta claims **co-training with the harness** *« to maximize harness compatibility »*, long-horizon training (whole repo, end-to-end projects, self-research, context compaction), and a self-improvement loop where version 1.1 generates the environments and grades the solutions, producing the training set for 1.2.

**The central fact of this announcement is nowhere stated in its text.** The four published comparisons exist only as images, and they place Muse Spark 1.2 **behind Opus 5 in all four cases**: 82.9% versus 86.7% on Terminal-Bench 2.1, 59.3% versus 65.0% on DeepSWE 1.1, **70.6% versus 79.4% on Meta&apos;s own internal benchmark**, and +68.7% versus +74.0% on the GPU kernel optimization case study, where the model finishes **fourth out of six**, behind GPT 5.6 Sol and behind Anthropic&apos;s previous generation.

**And the model&apos;s own gain is smaller than it appears.** On the two public benchmarks, version 1.1 is evaluated with `mini-swe-agent` and 1.2 with Muse Code: the 6.7-point gap conflates model and harness. On the internal benchmark, the only comparison with no harness indicated, it drops to **2.3 points**.

The announcement therefore stands mainly as **empirical confirmation** of a thesis already stated: value is shifting toward the harness, and a harness co-trained with its own weights makes those weights all the more non-interchangeable.&lt;/p&gt;</content:encoded><category>AI Coding Agents &amp; Skills</category><category>Meta AI Research</category><category>Muse Code</category><category>Muse Spark 1.2</category><category>terminal coding agent</category><category>beta</category></item><item><title>Fact-checking : synthèse sur Delos (Delos Intelligence / delos.so)</title><link>https://www.thekb.eu/en/fiches/delos-intelligence-fact-check-levee-2026-07-20/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/delos-intelligence-fact-check-levee-2026-07-20/</guid><description>Fact-checking synthesis on **Delos Intelligence** (delos.so), a French B2B generative AI startup, comparing a prior tech-watch note against **primary sources** (Alexandre Dewez&apos;s &quot;Overlooked&quot; post / 20VC, April 15, 2025, the delos.so website, official registries) and specialized press (Le Monde Informatique, L&apos;Usine Nouvelle, FrenchWeb, Le JDD). **Overall verdict: reliable factual backbone.** The **€2.5M seed round** (≈$2.74–2.83M) led by **20VC** (Harry Stebbings) in **April 2025**, with Inovia Capital, Kima Ventures (Xavier Niel) and Plug and Play, is confirmed; so are the founders (brothers **Pierre** and **Thibaut de la Grand&apos;rive**) and the clients **TotalEnergies, Shiseido, Groupe Casino**. **Strong methodological point**: the list of business angels — often suspected of hallucinatory &quot;padding&quot; — is **CONFIRMED word for word** by the lead investor&apos;s press release (Pigment, Dataiku, Hexa plus Ramp and Kerala to add): this is therefore NOT a hallucination. **To correct**: the &quot;50 people&quot; headcount is **not sourceable** (~20 in April 2025, about forty by late 2025); the actual pricing grid is richer (a **Student tier at €10** plus Enterprise on request, in addition to €25/45/80); user figures (10,000 → 50,000 → &quot;100,000+&quot;) and ARR are **self-reported and unaudited**. **To flag as speculative**: **no Series A has closed** (only announced as an intention targeting March 2026); **no overall ARR published** (the only mention is a self-promotional &quot;$1M ARR in a few days&quot; for the new **Workers** product, referring to that product alone). &quot;100% Scaleway&quot; sovereignty was **still being finalized** at the end of 2025 (compute still partly running on Azure France). The note&apos;s interest is as much methodological — **how to distinguish, within an AI-generated synthesis, what is confirmed, partially accurate, speculative, and self-reported** — as it is documentary.</description><pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This note verifies a tech-watch synthesis on **Delos Intelligence** (delos.so), a French B2B generative AI startup, by comparing it against primary sources (Alexandre Dewez&apos;s &quot;Overlooked&quot; post / 20VC, April 15, 2025, the official website, official registries) and specialized press. The backbone is **reliable**, but several figures need requalification.

**Funding — confirmed.** Delos raised **€2.5M in a seed round** — ≈$2.74 to $2.83M depending on the conversion — a round **announced mid-April 2025**, led by **20VC** (Harry Stebbings), with **Inovia Capital, Kima Ventures (Xavier Niel) and Plug and Play**. Notably, the list of **business angels**, exactly the kind of information an LLM can hallucinate, is **confirmed word for word** by the lead investor&apos;s press release — Éléonore Crespo &amp;amp; Romain Niccoli (Pigment), Florian Douetteau (Dataiku), Thibaud Elzière (Hexa), plus Mark Goldberger (Ramp) and Antoine Freysz (Kerala), the latter two *missing* from the initial synthesis. However, **no Series A has closed**: it is only **announced as an intention** (&quot;several tens of millions of euros by March 2026&quot;), with no press release or database entry.

**Business model — partially accurate.** **Credit-based** SaaS (1 credit ≈ a simple query). The actual pricing grid is richer than &quot;€25–80&quot;: **Student €10, Explore €25, Advanced €45, Premium €80** (increasing credit volumes), plus **Enterprise on request**. The **individual/B2C offering is indeed real**, but the core target remains **B2B**. Orchestrated models: ChatGPT, Claude, Mistral, Gemini, Cohere, Llama. **Sovereignty** (Scaleway hosting) was **still being finalized** at the end of 2025, with compute still partly running on Azure (France), with a full switch to Scaleway targeted for early 2026.

**Team and clients — partially accurate.** Founded on **July 2, 2023** by brothers **Pierre** and **Thibaut de la Grand&apos;rive**. The &quot;**50**&quot; headcount figure is **not sourceable**: ~20 in April 2025, about forty by late 2025. **200 client companies** confirmed; clients **TotalEnergies, Shiseido, Groupe Casino** confirmed (plus Allianz, Best Western, BPCE, the French Ministry of the Armed Forces…). User numbers (10,000 → 100,000+) and **ARR** are **self-reported**: no overall ARR has been published, and the only mention (&quot;$1M ARR in a few days&quot;) refers to the **Workers product alone** and is unaudited.

**Cross-cutting lesson**: a fact-check grades levels of evidence (confirmed / partial / speculative / not sourceable / self-reported) rather than issuing a binary verdict — and verifies a plausible piece of information before suspecting it of being a hallucination.&lt;/p&gt;</content:encoded><category>Economy &amp; Market</category><category>Delos Intelligence</category><category>delos.so</category><category>fact-checking</category><category>source verification</category><category>hallucination</category></item><item><title>Netflix Q2 2026 Shareholder Letter — leveraging technology to improve every aspect of our service (zoom IA/GenAI)</title><link>https://www.thekb.eu/en/fiches/netflix-q2-2026-genai-production-personnalisation-2026-07-16/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/netflix-q2-2026-genai-production-personnalisation-2026-07-16/</guid><description>Netflix — Q2 FY2026 shareholder letter: GenAI scales up in production (≈300 titles in 2026), LLMs for discovery and natural-language search, AI tools across the entire advertising cycle (Netflix)</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;**Netflix**&apos;s **Q2 FY2026** shareholder letter (July 16, 2026) places technology — and specifically **AI/GenAI** — as one of its **three strategic pillars** (&quot;leveraging technology to improve every aspect of our service&quot;), alongside entertainment value and monetization. At the outset, management summarizes the ambition: &quot;We are leveraging AI to provide a more personalized, immersive and interactive experience for members, enhance ads capabilities for brands, and improve the quality of our series and films.&quot;

**Production: GenAI scales up.** This is the most tangible point. Across the entire production cycle — from concept and pre-visualization through post-production and delivery — the use of GenAI by Netflix&apos;s **creative partners** is &quot;growing rapidly.&quot; In 2026, **GenAI workflows were used across approximately 300 titles**, with the highest concentration in **post-production**. Netflix highlights a dual benefit: **higher quality, faster and at lower cost** than traditional methods. The strongest claim: in some cases, productions would have had to **forgo key shots or sequences** without GenAI. Three examples are named — *Glory* (India), *Brasil 70: A Saga do Tri* (Brazil), and *The American Experiment* (US) — which used GenAI for **complex sequences**: augmented crowds, historical battles, worldbuilding establishing shots.

**Product and discovery.** Netflix leverages **LLMs** to improve **title discovery** and better understand member preferences. The search experience is enhanced with new **voice search** and **AI-powered natural-language search**, serving a &quot;more personalized, immersive and interactive&quot; experience.

**Advertising.** In the ads business, Netflix has **extended its AI tools across the entire advertising cycle** — planning, creative production, campaign management, optimization, and reporting. The company is **further automating transactions** with advertisers by extending programmatic access to **Pause Ads** and live inventory, reducing the manual effort that historically limited smaller buyers. These investments (Netflix Ads Suite + programmatic capabilities) are fueling ad growth.

All of this comes within a solid quarter: **revenue of $12.6B (+13% year-over-year)**, operating margin of **33.4%**, 2026 guidance tightened to $51.0-51.4B. The framing remains cautious: GenAI is presented as an **augmentation** of creators, never as a substitute — a communication choice that avoids controversies over rights and employment.&lt;/p&gt;</content:encoded><category>Transformation &amp; Adoption</category><category>artificial intelligence</category><category>GenAI</category><category>generative AI</category><category>LLM</category><category>Netflix</category></item><item><title>ZML/LLMD : et si le « Docker des LLM » était français ?</title><link>https://www.thekb.eu/en/fiches/sfeir-zml-llmd-docker-llm-inference-souveraine-2026-07-09/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/sfeir-zml-llmd-docker-llm-inference-souveraine-2026-07-09/</guid><description>SFEIR analysis (consulting-firm voice) of the launch, on July 8, 2026, of **LLMD** by the Paris-based startup **ZML** (founded by **Steeve Morin**, former VP Engineering at Zenly): an inference server that runs LLMs across **five chip families** (NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI, Apple Metal) **from a single codebase**. Structuring thesis: training is ceding the spotlight to **inference**, where cost per token, latency, and above all **dependence on silicon** are now decided. ZML&apos;s bet — summed up by the motto *model to metal* — is to **decouple the model from the hardware** via a compiler written in **Zig + MLIR** that produces a hermetic native binary, with no Python in the execution path, exposed through an **OpenAI-compatible API**. Two components, two licenses: **ZML** (the framework, Apache-2.0, &gt;90% Zig) is open source; **LLMD** (the server) is not, free at launch. The article reads the object through three consulting-firm lenses — **token FinOps**, **architectural freedom** (Design to Exit), **sovereignty** (emerging European chips, integration into the VSORA Jotunn8 processor) — then delivers an unsparing verdict: it is an **alpha**, to be placed &quot;under active watch,&quot; not to switch to today.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;On July 8, 2026, the Paris-based startup **ZML** released **LLMD**, an inference server that runs large language models across **five chip families** (NVIDIA, AMD, Google, Intel, Apple) from **a single codebase**. SFEIR reads this as a signal: as training cedes the spotlight to **inference**, the real battleground — and cost center — shifts toward **serving**, where cost per token, latency, and dependence on silicon are decided.

ZML&apos;s bet comes down to three words, *model to metal*: not offering yet another model, but a layer that **decouples the model from the hardware**. The stack has four layers. At the top, models (Qwen, Gemma, Mistral, LLaMa) loaded **zero-copy** via a virtual file system from Hugging Face, S3, or GCS. Then **LLMD**, a server exposing an **OpenAI-compatible API** (drop-in) with continuous batching, paged attention, prefix caching, tool calling, and Prometheus metrics. Below that, **ZML** compiles the graph **upfront, once and for all**, into a **hermetic native binary** in **Zig + MLIR**, with no Python in the execution path. This binary runs on five backends: CUDA, ROCm, TPU, oneAPI, Metal. The elegance lies in being &quot;portable, not leveled&quot; — chip-specific paths (FlashAttention, AITER) are preserved. Figures announced (by the vendor): images from 1.7 GB (CUDA) to ~140 MB (Apple), cold start of 1-2 s on an 8B model, and the **DFlash** accelerator (claimed &quot;up to 10×,&quot; ~6.17× in the underlying research).

Two components, two licenses: **ZML** (the framework) is open source (Apache-2.0, &amp;gt;90% Zig); **LLMD** (the server) is not, free at launch while usage data is collected. The demo runs in two commands on Apple Silicon Macs; a 27B model in BF16 requires ≥ 64 GB of unified memory.

SFEIR reads the object through three client-facing lenses: **FinOps** (choosing the cheapest chip → acting on the cost per token), **architectural freedom** (**Design to Exit**, built-in reversibility, cf. France Télévisions/ALIX) and **sovereignty** (European chips Axelera, Kalray, SiPearl, VSORA; a VivaTech 2026 partnership with Scaleway, VSORA, and the Île-de-France Region, integration into the Jotunn8 processor).

Unsparing verdict: it is an **alpha**, not for production; support for specific local machines (DGX Spark, Ryzen AI Max+) is neither named nor benchmarked. Against vLLM (server-GPU throughput) and llama.cpp (single-user local), LLMD aims for the middle ground. Not to switch to today, but to place &quot;under active watch&quot;: a serious, *made in France* candidate to become the &quot;*docker run* of inference.&quot;&lt;/p&gt;</content:encoded><category>Tools &amp; Platforms</category><category>LLM Inference</category><category>serving</category><category>ZML</category><category>LLMD</category><category>Steeve Morin</category></item><item><title>Project Genie: Interactive World Models with Genie 3</title><link>https://www.thekb.eu/en/fiches/google-deepmind-project-genie-3-world-models-2026-02/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/google-deepmind-project-genie-3-world-models-2026-02/</guid><description>Project Genie - Real-Time Interactive World Models from Google DeepMind</description><pubDate>Sun, 01 Feb 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Google DeepMind is launching Project Genie, a web application that lets Google US Ultra subscribers create and explore interactive worlds generated by the Genie 3 model. Unlike conventional video models that produce fixed sequences, a &quot;world model&quot; generates the environment frame by frame in real time, letting the user navigate and interact.

**Workflow and creative pipeline**: The user starts by describing their world and character. Nano Banana Pro first generates a &quot;canvas&quot; image that serves as a visual starting point. Clicking &quot;Generate World&quot; prompts Genie 3 to turn this 2D image into an explorable 3D environment. The transition from 2D to immersive 3D is the &quot;wow moment&quot; identified by testers. The application also allows uploading personal photos - a photographed toy dinosaur can become a controllable character in a reconstruction of the room.

**Technical challenges**: Genie 3 tackles a more complex problem than standard video generation. A video model can retroactively adjust frames to ensure consistency; Genie 3 must generate in real time, consistent with both the past AND the user&apos;s immediate action, without knowing future inputs. The current 60-second limit results from a trade-off: the world&apos;s dynamism tends to gradually decrease, and serving costs remain high.

**Evolution since Genie 1**: Genie 1 was a research paper. Genie 2 (December 2024) offered 10 seconds at low resolution, without real time. Genie 3 (announced August 2025, now launched) reaches one minute in real time with photorealistic quality. The team notes that a year ago, a minute of real-time consistency seemed an ambitious goal; today, users are asking for more.

**Envisioned applications**: Beyond entertainment, the team is exploring education (personalized therapeutic exposures, such as a child exploring a room full of virtual spiders) and embodied intelligence. The Simmer project already uses Genie 3 to train AI agents capable of accomplishing goals in arbitrary 3D worlds - a step toward embodied AGI.

**Outlook**: The roadmap includes multiplayer (complex because of latency), more interaction controls, a developer API, and expansion to other surfaces. The team estimates it has reached 50% of its vision, with &quot;enormous headroom&quot; for further improvements. The ultimate vision: a simulation indistinguishable from reality, &quot;a copy of the universe where you can do whatever you want&quot;.&lt;/p&gt;</content:encoded><category>Products &amp; Services</category><category>World models</category><category>Project Genie</category><category>Genie 3</category><category>Google DeepMind</category><category>Google Labs</category></item><item><title>Tech predictions for 2026 and beyond</title><link>https://www.thekb.eu/en/fiches/vogels-tech-predictions-2026-allthingsdistributed-2025-11-25/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/vogels-tech-predictions-2026-allthingsdistributed-2025-11-25/</guid><description>Werner Vogels - Tech Predictions 2026 - All Things Distributed - Amazon CTO - AI Trends - Companionship Revolution - Education Transformation - Healthcare Innovation - Human-AI Collaboration</description><pubDate>Tue, 25 Nov 2025 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In his eagerly awaited annual essay, Werner Vogels, Amazon&apos;s CTO, presents his technology predictions for 2026 and beyond. His central thesis: placing AI &quot;in the human loop&quot; (&quot;AI in the human loop&quot;) rather than the reverse, to amplify human capabilities and address the most pressing social crises. Three areas structure the essay.

**The companionship revolution.** Loneliness affects 1 in 6 people worldwide and has been designated a public health crisis by the OMS; it increases the risk of dementia by 31% and stroke by 30%. Vogels predicts the rise of companion robots (Pepper, Paro, Lovot) in long-term care: 95% of dementia patients have beneficial interactions with Paro, with reduced agitation, depression, and medication use. Because humans are biologically wired to project intention onto autonomous movement (MIT study; 50-80% of Roomba owners name their vacuum), robots like Amazon Astro create genuine emotional bonds through mobility and facial expressions.

**The transformation of education.** Faced with a global teacher shortage and a system designed for compliance rather than curiosity, personalized AI tutors adapt to each student&apos;s style, pace, and language. The impacts are measurable: +65% willingness to attempt difficult tasks, up to 17 IQ points gained in autistic children (Duke study), and 5.9 hours saved per week for teachers — six weeks per year reinvested in creativity and individual support. NextGenU already produces culturally adapted textbooks at 1/100th of the traditional cost.

**The reinvention of healthcare.** Amid distrust of medical institutions and misinformation, trustworthy AI agents can provide reliable medical information and improve treatment adherence, with humans remaining in control.

Vogels also identifies challenges: social acceptance of companion robots, emotional dependence on machines, the regulatory framework for AI agents in healthcare, and equity of access. But he sees massive opportunities in them: an aging population, educational innovation, preventive health, and new economic models.

This vision represents a fundamental shift in our relationship with technology: the move from transactional tools to partners that help us solve the deepest human problems, with AI amplifying humans rather than replacing them.&lt;/p&gt;</content:encoded><category>Products &amp; Services</category><category>Tech Predictions</category><category>AI Trends</category><category>Werner Vogels</category><category>Amazon CTO</category><category>Human-AI Collaboration</category></item><item><title>Perplexity Integrates Directly into Chrome Browser, Challenging Google Search Dominance</title><link>https://www.thekb.eu/en/fiches/perplexity-chrome-integration-browser-ai-search-2025-10-22/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/perplexity-chrome-integration-browser-ai-search-2025-10-22/</guid><description>Perplexity - Chrome integration - Browser AI - Search - Google competition - Native integration - AI-powered search</description><pubDate>Wed, 22 Oct 2025 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Perplexity announced a **native integration into the Google Chrome browser**, allowing users to **set Perplexity as their default search engine** and access AI search directly from the address bar. This move represents a **bold competitive challenge** to Google within its own browser, **dramatically reducing the friction** for users seeking an AI-native search experience without having to visit the Perplexity site separately.

**Integration mechanics**

The implementation allows users to: **set Perplexity as the default engine** in Chrome settings (alongside Google, Bing, DuckDuckGo), **search directly from the omnibox** (address bar queries are routed to Perplexity), **receive AI-generated answers** (instead of traditional link lists), **view source citations** (preserving transparency), **continue the conversation** (follow-up dialogue on the searched topic). The technical implementation relies on Chrome&apos;s **OpenSearch protocol**, which allows alternative engines to integrate without friction.

**Strategic scope: distribution**

A tech industry truism: **distribution determines the winner**. Google Search dominates partly thanks to Chrome integration — the browser&apos;s default search generates a massive volume of queries. Perplexity&apos;s integration **solves a critical distribution challenge**: fewer steps from query to result (no separate site visit), integration into the existing workflow (the address bar is already the primary search interface), lower adoption friction (a single settings change versus repeated visits), increased visibility (a constant reminder that an alternative exists).

**Competitive dynamics with Google**

The move places Perplexity **directly on Google&apos;s own turf** — the Chrome browser that Google controls. This creates interesting tensions: **Google could block the integration** (but faces antitrust scrutiny — already under challenge over default engines), **Google could add similar AI features** (Gemini integration is the likely response), **user choice becomes key** (default-settings battles intensify), **direct quality comparison** (users easily switch between engines and compare results).

**User experience transformation**

Traditional search: query → list of links → clicks through multiple results → manual synthesis of information. **Perplexity search**: query → AI-synthesized answer with source citations → optional follow-up questions → refined understanding. This **fundamentally different paradigm** is especially valuable for: research tasks (synthesis of multiple sources), fact-checking (visible citations), complex queries (multi-step reasoning), exploratory learning (natural follow-up dialogue).

**Monetization challenges**

The integration raises business-model questions: **fewer website visits** (publishers potentially lose traffic), **attribution complexity** (how to credit sources cited by the AI?), **advertising disruption** (traditional search ads live in link lists — where do ads go in AI-generated answers?), **premium features** (how to differentiate free and paid offerings?). Perplexity must balance **user value against ecosystem sustainability**.

**Technical requirements and performance**

Chrome integration requires: **low latency** (users expect instant results as with Google), **reliability** (downtime is unacceptable for a default search engine), **query understanding** (handling the full diversity of search intent), **scalable infrastructure** (potentially massive increase in query volume if adoption grows), **cross-device synchronization**.

**Google&apos;s potential responses**

Likely actions from Google: **accelerate Gemini integration** into search, **leverage Chrome control** (promoting Google&apos;s own AI search features), **improve search quality** (narrowing the advantage of synthesized answers), **adjust commercial terms** (restricting alternative engines where legally possible), **acquire or form partnerships**.

**Regulatory context**

The timing is notable given the ongoing **antitrust scrutiny** of Google&apos;s search dominance. The U.S. DOJ and European regulators are examining default search agreements and browser integration practices. Perplexity&apos;s move could **strengthen antitrust arguments**: it demonstrates that viable alternatives exist, shows that Google&apos;s control limits competition, and illustrates the value of user choice.

**Broader industry implications**

Success could inspire: **other AI search engines** (You.com, Phind) pursuing browser integration, **browser diversification** (Firefox, Edge offering multiple AI search options), **a search paradigm shift** (acceleration toward synthesized answers versus link lists), **new entrants** (lowered barriers encouraging innovation).

**Adoption unknowns**

Perplexity&apos;s success depends on: actual adoption rates (will people change their default settings?), quality consistency at scale, monetization viability, the speed of Google&apos;s competitive response, and the evolving regulatory environment.

The integration represents a **significant milestone** in AI search competition, taking the battle directly to Google&apos;s most powerful distribution channel.&lt;/p&gt;</content:encoded><category>Products &amp; Services</category><category>Perplexity</category><category>Chrome integration</category><category>browser AI</category><category>AI search</category><category>Google competition</category></item><item><title>google-agentic-commerce/AP2: Building a Secure and Interoperable Future for AI-Driven Payments</title><link>https://www.thekb.eu/en/fiches/google-agentic-commerce-ap2-payment-protocol-2025-09-16/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/google-agentic-commerce-ap2-payment-protocol-2025-09-16/</guid><description>Agent Payments Protocol (AP2) - Google Agentic Commerce - Secure AI-Driven Payments - GitHub</description><pubDate>Tue, 16 Sep 2025 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The GitHub repository `google-agentic-commerce/AP2` introduces the **Agent Payments Protocol (AP2)**, an initiative aimed at establishing a secure and interoperable framework for AI-driven payment systems. The project&apos;s core mission is to enable a future in which artificial intelligence can handle payment transactions smoothly and safely across diverse platforms and agents.

**Practical resources and flexibility**

The repository serves as a practical resource, offering a collection of code samples and demos illustrating AP2&apos;s key components and features. While the provided samples rely on specific technologies such as the **Agent Development Kit (ADK) and Gemini 2.5 Flash**, the documentation explicitly states that the Agent Payments Protocol itself is **not tied** to these tools. This flexibility allows developers to integrate AP2 with their preferred agent development kits and AI models, fostering broad adoption and customization.

**Architecture and selected scenarios**

The repository is structured to guide users through various selected scenarios, designed to demonstrate different aspects of the protocol in action. These scenarios are organized under `samples/android/scenarios` for Android applications and `samples/python/scenarios` for Python implementations. Each scenario is self-contained, with a `README.md` file providing a detailed description along with clear installation and local execution instructions. A `run.sh` script is included to simplify execution.

**Authentication and configuration**

Running the scenarios requires **Python 3.10 or higher**. The repository describes two main authentication methods: the **Google API key**, recommended for development environments for its simplicity, or **Vertex AI** configuration, advised for production deployments as it is more robust and scalable. Instructions are provided for setting environment variables or using a `.env` file for credentials.

**Core objects and future plans**

AP2&apos;s core objects and definitions reside in the `src/ap2/types` directory, with a **planned dedicated PyPI package release** to simplify installation and dependency management. The demos involve **multiple agents and servers**, with most of the source code residing in `samples/python/src`. Scenarios that include an Android application (shopping assistant) have their dedicated source code in `samples/android`.

**Open and collaborative approach**

The project is licensed under **Apache-2.0**, reflecting an open and collaborative approach. This initiative represents Google&apos;s commitment to building the infrastructure enabling safe and efficient AI-driven commercial transactions. The distributed multi-agent/server architecture illustrates scalability considerations and the protocol&apos;s real-world applicability. The repository positions AP2 as a foundational technology for the emerging agentic commerce ecosystem, where AI agents autonomously manage financial transactions with appropriate security safeguards and interoperability standards.&lt;/p&gt;</content:encoded><category>Products &amp; Services</category><category>payments</category><category>agents</category><category>a2a</category><category>generative-ai</category><category>gen-ai</category></item><item><title>Google DeepMind Unveils Genie 3: Revolutionary Interactive Video Generation Model</title><link>https://www.thekb.eu/en/fiches/google-genie-3-video-generation-model-deepmind-2025-08-05/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/google-genie-3-video-generation-model-deepmind-2025-08-05/</guid><description>Google DeepMind Genie 3 — interactive video generation model: world models, controllable generation, playable AI-generated games (deepmind.google)</description><pubDate>Tue, 05 Aug 2025 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Google DeepMind announces **Genie 3**, a revolutionary **interactive video generation** model capable of creating **controllable, temporally coherent video** that responds to user actions in real time. Unlike previous video models producing fixed sequences, **Genie 3 functions as a world model** — understanding spatial relationships, physics, and causality — and enables **AI-generated interactive experiences**, including playable games created from text descriptions or images.

**Core innovation: controllable generation**

The fundamental advance is **user control during generation**. Genie 3 accepts continuous inputs — arrow keys, mouse movements, action commands — and generates video that responds appropriately. Example: the user requests &quot;a platform game in a forest,&quot; Genie generates the first frame, then the user **controls the character&apos;s movements**, with the model generating subsequent frames (jumps, movement, environment interactions). This **interactive loop** creates playable experiences rather than passive videos.

**World model architecture and training**

Genie 3 implements a **latent world model**: a compressed representation of an environment&apos;s physics, understanding of spatial relationships and object permanence, prediction of action consequences, temporal coherence over extended sequences. The model **does not run pre-programmed physics**: it learned physical rules by observing vast volumes of video game sequences (2D platform games as primary data, action annotations, varied visual styles), developing an **emergent understanding** of gravity, collisions, and motion dynamics. Its **11 billion parameters** allow it to capture fine-grained relationships between actions and visual consequences.

**Temporal coherence and applications**

Video models struggle to maintain object appearance, position, and physics across frames. Genie 3 addresses this through long-term memory mechanisms, physics-informed priors, spatial attention, and action conditioning, with markedly improved coherence. Applications: **rapid game prototyping**, custom educational games, accessibility, procedural content, no-code creative tools — a **democratization of game development**.

**Limitations and competition**

Acknowledged limitations: a ceiling on mechanics complexity, coherence degradation over very long sequences, imperfect control fidelity, high inference cost, training data bias. Against **Runway Gen-3**, **OpenAI Sora**, or Meta&apos;s Make-A-Video, Genie 3&apos;s **interactive control** is the key differentiator, a step toward **general-purpose world models**. In the long run, Genie 3 charts a trajectory toward general-purpose world simulators, AI-driven interactive experiences beyond gaming, and AI-generated virtual worlds responsive to user agency.&lt;/p&gt;</content:encoded><category>Products &amp; Services</category><category>Google</category><category>Genie 3</category><category>DeepMind</category><category>video generation</category><category>generative AI</category></item><item><title>Expanding AI Overviews and introducing AI Mode</title><link>https://www.thekb.eu/en/fiches/google-ai-mode-search-personalized-sites-2025-03-05/</link><guid isPermaLink="true">https://www.thekb.eu/en/fiches/google-ai-mode-search-personalized-sites-2025-03-05/</guid><description>Google AI Mode - Search transformation - Personalized sites - Generative search - Generative web</description><pubDate>Wed, 05 Mar 2025 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Google launches **AI Mode**, a feature that profoundly transforms the online search experience. For each search result, a **personalized site is automatically generated**, tailored specifically to the user&apos;s query. This innovation represents a **paradigm shift** in how we access information online, and could mark the **end of traditional websites** as we know them.

**Fundamentals of the generative approach**

Rather than simply indexing and ranking existing web content, AI Mode **dynamically generates content** tailored to the search query. Two users searching for the same topic can thus receive entirely different generated &quot;sites,&quot; depending on context, search history, and the specifics of the query. The system analyzes the intent behind the search and builds the response from scratch, synthesizing information from multiple sources.

**Transformation of the web ecosystem**

Google&apos;s generative approach could **fundamentally redefine the web ecosystem**. If Google generates personalized content for each query, the traditional concept of a &quot;website&quot; potentially becomes obsolete. Content creators face an existential question: why maintain traditional sites if users primarily interact with versions generated by Google?

**Redefinition of SEO**

SEO, an entire industry built around ranking in Google results, faces a **complete redefinition**. Traditional tactics — keyword optimization, backlinks, technical SEO — could become irrelevant if Google generates instead of indexing. A new form of &quot;generative SEO&quot; could emerge, aimed at ensuring that AI systems correctly represent brands&apos; messages and information.

**Implications for business models**

The web&apos;s business model — advertising, subscriptions, affiliate marketing — relies on the assumption that users actually visit sites. **AI Mode threatens this model**. If users consume information directly from Google-generated results without visiting source sites, how do content creators get paid? Google will likely need to address this fundamental challenge to maintain a healthy content ecosystem.

**Content attribution and copyright**

The generated personalized sites raise complex questions of **attribution and copyright**. When AI synthesizes information from multiple sources to create a new &quot;site,&quot; who owns the resulting content? How are original sources credited and compensated? These legal and ethical questions will likely spark significant debate, even litigation.

**Transformation of the user experience**

From the user&apos;s perspective, AI Mode promises an **optimal experience**: information perfectly tailored to the query, no more need to navigate between multiple sites, reduced time between question and answer. But this also raises concerns about **filter bubbles, echo chambers**, and the loss of serendipitous discovery that characterizes traditional web browsing.

**Competitive implications**

AI Mode represents a **major escalation** in the race for search innovation. Competitors such as Microsoft (Bing with ChatGPT), Perplexity, and other AI-powered search engines will need to respond. A future where search engines generate instead of indexing could reshape the entire infrastructure of the Internet.

**Implementation questions**

Practical questions remain: how does Google guarantee accuracy? What mechanisms exist against misinformation? How does the system handle controversial topics? How fresh is the generated content? Can users verify sources?

The announcement of AI Mode signals **Google&apos;s vision for the future of search**: moving from a passive index to an active content generator. Whether this vision benefits the broader Internet ecosystem remains to be seen, but the transformation is already underway.&lt;/p&gt;</content:encoded><category>Products &amp; Services</category><category>Google Search</category><category>AI Mode</category><category>Generative search</category><category>Personalization</category><category>Generative web</category></item></channel></rss>