Skip to content

root / tags / glm-5-2

#GLM-5.2

6 fiches

Quality & Security Auto-verified translation

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Announcement post published on the **official Z.ai blog** (formerly Zhipu AI, Chinese lab) on **August 14, 2026**, **with no individual byline**, ~2,000 words plus footnotes. It announces **GLM-5.3**, successor to GLM-5.2, opening with a methodological thesis: *« Scaling post-training is all we did for GLM-5.3. »* Same base model as GLM-5.2 — *« every gain comes from post-training »*. Three announcements. **(A) An open-weights coding model**: +50% claimed on **Z.ai Code Bench**, an unpublished in-house benchmark. **(B) A cyber capability presented as "emergent"**, which the body of the text traces to a training choice — *« As part of post-training, we introduced vulnerability discovery data and environments into the training mix. We expected this to make the model better at finding and reasoning about vulnerabilities »* — what came as a surprise was the speed and the change in nature: the model moves from identifying isolated flaws to *« coherent plans for complete exploitation chains »*. Gains grow with position in the exploitation chain: CyberGym 77.2 → **84.5%**, ExploitBench 24.4 → **54.4%** (×2.2), ExploitGym 29 → **105** tasks in 2h (×3.6), with the gap to the closed frontier remaining wide (181 and 247 tasks). Z.ai puts it this way: *« Capability is growing fastest exactly where we are furthest behind. »* The post also publishes a **Z.ai Security Disclosure Ledger**: **2,436 vulnerabilities identified across 269 open source projects** — kernels, OSes, browser engines, infrastructure, web applications, network protocols — the oldest introduced in **1981**, average lifetime before discovery **26.6 years**, of which **53 disclosed** and **2,383 under embargo**. **(C) A weight release** *« within two weeks of launch, once safety evaluation and hardening are complete »*. The most reusable methodological contribution: **environment and verifier synthesis**, the latter produced without access to the reference solution and admitted only after a triptych of negative controls — **oracle**, **no-op**, **unsolved-state**. All agentic evaluations are conducted **in Claude Code 2.1.207**.

#GLM-5.3#GLM-5.2#Z.ai

**Z.ai** (anciennement **Zhipu AI**) · laboratoire d'IA chinois · éditeur de la famille **GLM**. Billet **institutionnel et non signé** : aucun auteur nommé · aucun chercheur mis en avant · aucun lien vers un rapport technique ou une carte de modèle. Publié le **14 août 2026**. La page est une SPA React — le HTML servi est un `<div id="root">` vide · et le texte comme les scores ont dû être extraits du bundle `glm-5.3-BCnx8T5_.js` · où ils figurent en valeurs source.

Economy & Market Auto-verified translation

Mistral AI wants to build 1 gigawatt of European compute by 2030 — and lock in customers now.

News article analyzed, published on **VentureBeat** on **August 11, 2026** by **Michael Nuñez**, based on an **exclusive interview with Timothée Lacroix**, co-founder and CTO of **Mistral AI**, conducted ahead of the announcement, ~2,000 words. Mistral is expanding its infrastructure offering in three parts: **Mistral Regional Endpoints** in general availability (pinning inference and its associated processing to Europe or the United States), a **Priority Tier** in public preview (committed service levels, custom quotas, availability SLA), and a **coalition of European enterprises** whose multi-year commitments are meant to fund **200 MW by the end of 2027** and **1 GW by the end of 2030**. The vehicle is called the **European Compute Unit (ECU)**: a claim on capacity built by Mistral, fungible across inference, training, model adaptation, or managed Kubernetes, over a targeted five-year horizon. Lacroix describes the mechanism bluntly — *"The whole point of compute units is to have commitment"* — and, on early exit: *"There is no getting out."* The article scales the ambition: Mistral states it operates *"less than 200 MW"* and details three sites totaling **77 MW** (44 MW near Paris, 23 MW in Sweden with EcoDataCenter, 10 MW in Les Ulis); **Epoch AI** puts the initial capex for a one-gigawatt AI datacenter at **~$38B**, and **Goldman Sachs Research** puts next-generation facilities at **$15-20M/MW excluding chips**, against the **~$4B** Mistral has raised in total (PitchBook). Added to this is a decision that *"is likely to raise a few eyebrows among sovereignty purists"*: Mistral is starting to **host third-party open models**, beginning with **GLM-5.2** from **Z.ai**, a Chinese lab — *"It's a great model. Everyone loves it. It's open-weight, so there was no good reason for us not to do it."* The article digs into the fine print of Mistral's documentation, which mentions *"limited, controlled transfers"* to subcontractors outside the region; pressed for detail, Lacroix points to **tool calls**, web search in particular, and states that **gating is the feature, not the bug**. The author's framing: *"full regional control is available, but the moment an AI agent reaches out to the open web, sovereignty becomes a configuration decision, not a default."* Two dependencies remain: **GPUs** come from Nvidia, and **Microsoft** — anchor tenant of Mistral's European datacenters since July — is presented as what de-risks the buildout.

#Mistral AI#digital sovereignty#AI sovereignty

**Michael Nuñez** — journaliste **VentureBeat** · couvre l'IA et l'infrastructure ; déjà présent au corpus. L'article est bâti sur un **entretien exclusif avec Timothée Lacroix** · cofondateur et CTO de Mistral AI · conduit **avant l'annonce** · et fait suite à un entretien de juin avec le même interlocuteur. Publié le **11 août 2026**.

AI Coding Agents & Skills Auto-verified translation

The Token Manifesto

Nicolas Martignole (Le Touilleur Express), co-written with **GLM-5.2** and **MiniMax-M3**, publishes **« The Token Manifesto »**: a pastiche of the **Manifeste Agile** (2001) transposed to the LLM era, where the unit of value is no longer the engineer-hour but the **token**. Four values: *short system prompts over clever system prompts*, *one clear example over three paragraphs of explanation*, *iterating in small steps over dumping the whole spec at once*, *outputting in a defined format over letting the model freestyle*. Twelve principles subvert those of Agile one by one — "simplicity, the art of maximizing the amount of work **not done by the model**," "self-organizing teams that spot repetition and document it once," "regular reflection **before the monthly bill arrives**." Beneath the humor ("staring at a usage bar nervously") lies a serious thesis: the real economic constraint of AI-assisted dev is no longer velocity but the **token budget** and the **context-window economy**. Two punchlines close the text: **« You don't have a prompt problem. You have a context-window problem. »** and **« Everyone's a prompt engineer until they run out of monthly quota. »** Worth noting, the meta wink: a manifesto on token frugality co-written *with* models.

#The Token Manifesto#Nicolas Martignole#Le Touilleur Express

Nicolas Martignole (Le Touilleur Express) · avec GLM-5.2 et MiniMax-M3

Tools & Platforms Auto-verified translation

Kimi K3 de Moonshot AI : quand le frontier open-weights rattrape le propriétaire

SFEIR's engineering-cabinet analysis ("an engineer's reading") of the **July 16, 2026** launch of **Kimi K3** by the Chinese laboratory **Moonshot AI**: an **open-weights, frontier-class model** whose provider claims **~2.8 trillion parameters**, a **one-million-token context**, and **weight release before July 27, 2026** (likely under a Modified MIT license, as with the K2 lineage). Thesis: capability once thought reserved for proprietary giants (Anthropic, OpenAI, Google) is becoming available **in open weights, at a discount price, from a Chinese lab**. SFEIR — despite being an **Anthropic and Google Cloud partner**, and thus "with no interest in oversell­ing a Chinese model" — adopts a cardinal **methodological caveat**: on launch day, **no official, complete benchmark table** exists; specs (2.8T, Kimi Delta Attention, +25% training efficiency) and scores are **vendor-stated** or drawn from **community arenas**, "to be treated as claims, not measured facts." The new architecture (**Kimi Delta Attention**, hybrid linear attention; decoding claimed up to **6.3x faster** at 1M tokens) breaks with the K2 cadence (K2 Jul. 2025 → K2.7 Code Jun. 2026, a flagship every two months); two variants accompany the launch (**K3 Max**, **K3 Swarm Max**), with forced sunsetting of the kimi-k2.5/moonshot-v1 series on **August 31, 2026**. **The real weapon is price** (~$3/M input, $0.30 cached, $15 output per secondary sources): a frontier open-weights model at this level **pulls the whole price-performance curve down** — the commoditization of the model layer, accelerated by open source. But the decisive singularity is not a score: it is **reversibility**. A frontier open-weights model turns a consumed API (vendor dependency) into an **option** (self-host, portability, exit from lock-in), at the cost of heavy infrastructure to host 2.8T parameters. SFEIR's view: **open-weights changes the question, not just the answer** — no longer "which model is best/cheapest?" but "how much of my system am I willing to make dependent on a vendor I don't control?". The right posture remains a **routed portfolio** (one model per task, one model per constraint), with Kimi K3 adding a **"reversibility" column** to the decision grid. The "AI Only" conviction stands unchanged: the model is a commodity, the durable advantage lies in the engineering around it (Context Engineering, harness, cost governance, ability to change one's mind). The figures still need validating "on your own" — your repositories, your data.

#Kimi K3#Moonshot AI#Yang Zhilin

SFEIR (voix éditoriale du cabinet)

Transformation & Adoption Auto-verified translation

AI4IT vs AI4Business : le renversement, et ce qu'il fait à vos budgets 2027

In-depth opinion piece (point of view) published on **sfeir.com** on June 24, 2026, by **Didier Girard** (Managing Director, SFEIR). **Central thesis**: in 2024 everyone was betting on **AI4Business** (AI in business processes) as the great value reservoir; by 2026 the picture has **reversed** — it is **AI4IT** (AI to produce the information system: code, SDLC, software factory) that is creating **measurable** value. The article *grounds* this thesis in the firm's tech watch: AI4Business disappointment (the MIT study "95% of pilots without ROI," contested but revealing; an **organizational** blockage / Mollick's Hayekian problem) versus quantified AI4IT evidence (Salesforce, Intercom, Raiffeisen, AWS/Bedrock, Atlassian, DORA). Mechanistic explanation: **code verifies itself** (compilation, tests, CI) whereas business processes have neither a compiler nor an immediate feedback loop. **2027 budget consequence**: a **CapEx→OpEx** shift, token price dynamics (rising peak — Fable 5 at 2× Opus — vs inference ÷280 and downward pressure from open weights/desktop), and **AI FinOps** driven by **cost per outcome**. Closes with **4 recommendations for the COMEX**.

#AI4IT#AI4Business#reversal

**Didier Girard** — Managing Director (CTO / DG) de **SFEIR** · ESN française (~1 000 personnes, France · Belgique · Luxembourg · Suisse). Auteur de l'article ; voix éditoriale du cabinet sur la transformation IA des DSI.

Economy & Market Auto-verified translation

GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark

Benchmark announcement from **Artificial Analysis** (independent AI model evaluation platform, via X/Twitter + model page): **GLM-5.2** from **Z.ai** (Zhipu AI, @Zai_org) becomes **the leading open weights model** and climbs to **#3 in the overall ranking** of **GDPval-AA**, a real-world benchmark for *economically valuable knowledge work* (long-horizon, multi-turn, agentic tasks). GLM-5.2 scores **1524 Elo**, behind only **Claude Fable 5 (1783)** and **Claude Opus 4.8 (1615)**, and on par with **GPT-5.5 (xhigh, 1509)**. It leads the next-best open model (**MiniMax-M3, 1408**) by a wide margin, along with numerous proprietary models: **Gemini 3.5 Flash (1357)**, **Qwen 3.7 Max (1289)**, **Muse Spark (1158)**. The tasks are genuinely agentic: **~31 turns per task** on average across **1,999 matches**. The same ranking holds on the **Artificial Analysis Intelligence Index** (1st among open weights), the **Agentic Index** (#3) and **AA-Briefcase** (#3, ahead of GPT-5.5 xhigh, behind only Fable 5). Notable highlight: an **open weights** model under **MIT license**, **MoE with 753B parameters / 40B active**, **1M-token context**, priced at **$1.40/$4.40 per 1M tokens** input/output, rivals the proprietary frontier on agentic work — a real step forward for open models.

#GLM-5.2#Z.ai#Zhipu AI

Artificial Analysis (@ArtificialAnlys)