# artificial-analysis-glm-5-2-gdpval-aa-open-weights-2026-06-22

## Veille

Benchmark announcement from **Artificial Analysis** (independent AI model evaluation platform, via X/Twitter + model page): **GLM-5.2** from **Z.ai** (Zhipu AI, @Zai_org) becomes **the leading open weights model** and climbs to **#3 in the overall ranking** of **GDPval-AA**, a real-world benchmark for *economically valuable knowledge work* (long-horizon, multi-turn, agentic tasks). GLM-5.2 scores **1524 Elo**, behind only **Claude Fable 5 (1783)** and **Claude Opus 4.8 (1615)**, and on par with **GPT-5.5 (xhigh, 1509)**. It leads the next-best open model (**MiniMax-M3, 1408**) by a wide margin, along with numerous proprietary models: **Gemini 3.5 Flash (1357)**, **Qwen 3.7 Max (1289)**, **Muse Spark (1158)**. The tasks are genuinely agentic: **~31 turns per task** on average across **1,999 matches**. The same ranking holds on the **Artificial Analysis Intelligence Index** (1st among open weights), the **Agentic Index** (#3) and **AA-Briefcase** (#3, ahead of GPT-5.5 xhigh, behind only Fable 5). Notable highlight: an **open weights** model under **MIT license**, **MoE with 753B parameters / 40B active**, **1M-token context**, priced at **$1.40/$4.40 per 1M tokens** input/output, rivals the proprietary frontier on agentic work — a real step forward for open models.

## Titre Article

GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark

## Date

2026-06-22

## URL

https://artificialanalysis.ai/models/glm-5-2

## Keywords

GLM-5.2, Z.ai, Zhipu AI, open weights models, open weights, GDPval-AA, agentic benchmark, knowledge work, Elo, multi-turn tasks, long-horizon, Artificial Analysis Intelligence Index, Agentic Index, AA-Briefcase, Mixture of Experts, MIT license, 1M-token context, reasoning, cost per token, proprietary frontier, Claude Fable 5, MiniMax-M3

## Authors

Artificial Analysis (@ArtificialAnlys)

## Ton

Profile: a factual, data-driven benchmark announcement from the perspective of an **independent third-party evaluator** (Artificial Analysis), distributed as an X/Twitter thread backed by a detailed model page, neutral English data-driven register, high technical level, targeting ML engineers, technical decision-makers, API buyers and foundation-model market observers. The tone is one of **objectified performance comparison**: rankings (Elo, indices), precise figures (1524, ~31 turns, 1,999 matches, $1.40/$4.40) and an explicit methodology (identical briefs given to multiple models, deliverables rendered as-is). Authority rests on Artificial Analysis's **reputation for neutrality** and methodological transparency (rendering of actual deliverables, multiplicity of cross-referenced indices). The rhetoric, free of commercial emphasis, lets the numbers carry the implicit thesis: the gap between **open** and **proprietary** models is closing on the most demanding ground — economically useful agentic work. Implicit metaphor: the benchmark as a "professional exercise" (a store supervisor's tasks, etc.) rather than an abstract academic test. The only explicit value judgment — *"a real step for open models"* — is measured and grounded in the price/performance ratio.

## Pense-betes

- **Core fact**: **GLM-5.2** (Z.ai / Zhipu AI) = **#1 open weights model** and **#3 overall** on **GDPval-AA** with **1524 Elo**.
- **GDPval-AA podium**: 1. **Claude Fable 5 (1783)**; 2. **Claude Opus 4.8 (1615)**; 3. **GLM-5.2 (1524)** — on par with **GPT-5.5 (xhigh, 1509)**.
- **Open weights dominance**: the next-best open model, **MiniMax-M3**, sits at just **1408** (wide margin).
- **Outperforms proprietary models**: **Gemini 3.5 Flash (1357)**, **Qwen 3.7 Max (1289)**, **Muse Spark (1158)**.
- **What GDPval-AA measures**: performance on **real, economically valuable knowledge work**, via **long-horizon, multi-turn tasks**; covers both professional AND creative work (e.g. a retail store supervisor's daily task list, an IEC document…). Method: identical briefs given to GLM-5.2 + 3 proprietary frontier models (Fable 5, GPT-5.5, Gemini 3.5 Flash), **deliverables rendered exactly as produced**.
- **Genuinely agentic**: GLM-5.2 averaged **~31 turns per task** across **1,999 matches**.
- **Cross-index consistency**: #1 among open weights on the **Artificial Analysis Intelligence Index**, **#3 on the Agentic Index**, **#3 on AA-Briefcase** (on AA-Briefcase: top open weights, ahead of GPT-5.5 xhigh, behind only Fable 5).
- **Specs (model page)**: **MoE 753B params / 40B active**, **reasoning model** (extended thinking), **1M-token context**, **text-to-text**, **MIT license** (commercial use), weights on Hugging Face, **released June 16, 2026**.
- **Economics**: **$1.40 / $4.40** per 1M tokens (input/output), **cache hit $0.26** (-81%), blended rate **~$0.90/1M**; throughput **106.3 tokens/s**, TTFT **1.36 s**; total eval cost **$982.90**. → Price/performance argument: agentic frontier at open weights pricing.
- **Implicit thesis**: the open/proprietary convergence is now playing out on **useful agentic work**, not just academic benchmarks.
- **Related**: extends the frontier-model race (Opus 4.8, Fable 5); market signal for "Chinese open weights" (cf. Qwen / MiniMax fiches); relevant to the **agent cost** debate and sovereignty/self-hosting.

## RésuméDe400mots

Artificial Analysis — an independent AI model evaluation platform — publishes (X/Twitter thread from June 22, 2026 + a detailed model page) a comparison placing **GLM-5.2**, the latest model from **Z.ai** (Zhipu AI), at the top of **open weights** models and **#3 in the overall ranking** of **GDPval-AA**. This benchmark measures performance on **real, economically valuable knowledge work**, through **long-horizon, multi-turn tasks** designed as genuine professional exercises (for example a retail store supervisor's daily task list, or an IEC technical document) covering both professional and creative work.

GLM-5.2 achieves **1524 Elo**, behind only **Claude Fable 5 (1783)** and **Claude Opus 4.8 (1615)**, and on par with **GPT-5.5 in xhigh setting (1509)**. Above all, it dominates the open field by a **wide margin**: the next-best open model, **MiniMax-M3**, scores only **1408**. GLM-5.2 also outperforms several proprietary models — **Gemini 3.5 Flash (1357)**, **Qwen 3.7 Max (1289)** and **Muse Spark (1158)**.

The **agentic** nature of the tasks is emphasized: GLM-5.2 averaged **~31 turns per task** across **1,999 matches**. Artificial Analysis's method involves giving **identical briefs** to GLM-5.2 and three proprietary frontier models (Fable 5, GPT-5.5, Gemini 3.5 Flash), then **rendering each deliverable exactly as produced**. The result is consistent across the firm's own indices: GLM-5.2 is **#1 among open weights** on the **Intelligence Index**, **#3 on the Agentic Index** and **#3 on AA-Briefcase** (where it is the top open model, ahead of GPT-5.5 xhigh and behind only Fable 5).

The model page rounds out the picture: GLM-5.2 is a **Mixture of Experts** with **753 billion parameters** (of which **40 billion active**), a **reasoning model** with **1M-token context**, distributed under **MIT license** (commercial use, weights on Hugging Face), released on **June 16, 2026**. On the economics side: **$1.40 / $4.40** per million tokens (input/output), a cache hit at **$0.26** (-81%), a throughput of **106.3 tokens/s** and a time to first token of **1.36 s**.

The message conveyed by the numbers is clear: that an **open weights** model at this price point rivals the proprietary frontier on **genuinely useful agentic work** constitutes, according to Artificial Analysis, *"a real step for open models."* The convergence between open and proprietary models is no longer playing out solely on academic tests, but on the economic value produced under agentic conditions.

## GrapheDeConnaissance

- Z.ai —a_créé→ GLM-5.2 (TECHNOLOGIE, 0.98)
- GLM-5.2 —est_variante_de→ GLM (TECHNOLOGIE, 0.97)
- GLM-5.2 —remplace→ GLM-5.1 (TECHNOLOGIE, 0.95)
- GLM-5.2 —mesure→ 1524 Elo sur GDPval-AA (#3 au général, #1 open weights) (MESURE, 0.95)
- Claude Fable 5 —surpasse→ GLM-5.2 (TECHNOLOGIE, 0.92)
- Claude Opus 4.8 —surpasse→ GLM-5.2 (TECHNOLOGIE, 0.9)
- GLM-5.2 —surpasse→ MiniMax-M3 (TECHNOLOGIE, 0.92)
- GLM-5.2 —surpasse→ Gemini 3.5 Flash (TECHNOLOGIE, 0.88)
- GLM-5.2 —concurrence→ GPT-5.5 (TECHNOLOGIE, 0.85)
- GLM-5.2 —mesure→ ~31 tours par tâche sur 1 999 matchs (MESURE, 0.85)
- GDPval-AA —mesure→ travail de connaissance économiquement valorisable en tâches multi-tours longue-horizon (AFFIRMATION, 0.9)
- Artificial Analysis —publie→ GDPval-AA (DOCUMENT, 0.9)
- GLM-5.2 —est_instance_de→ modèle à poids ouverts (CONCEPT, 0.95)
- GLM-5.2 —utilise→ architecture Mixture of Experts (753 Mds params / 40 Mds actifs) (CONCEPT, 0.9)
- GLM-5.2 —mesure→ tarif 1,40 $ / 4,40 $ par 1M tokens (entrée/sortie) (MESURE, 0.9)
- Artificial Analysis —affirme_que→ qu'un modèle open weights rivalise avec la frontière propriétaire sur le travail agentique est un vrai progrès (AFFIRMATION, 0.85)

---
Canonical: https://www.thekb.eu/en/fiches/artificial-analysis-glm-5-2-gdpval-aa-open-weights-2026-06-22/
