Skip to content
Technology

GPT-5.5

GPT-5.5 — Technology. score_expert_pass1: 71,4% (+-8,0%) · score_expert_pass5: 90,5% (+-12,9%)

The AI Safety Institute of the UK measured GPT-5.5 at 71.4% pass@1 (±8.0% standard error) on expert-level capture-the-flag tasks, and 90.5% pass@5 (±12.9%), the highest score AISI had recorded at the time of its pre-deployment evaluation published 30 April 2026. The comparison set matters more than the absolute number: GPT-5.4 scored 52.4% and Opus 4.7 48.6%, while Mythos Preview sat at 68.6%, statistically level with GPT-5.5. Basic tasks had been saturated at 100% by every frontier model since February 2026, so only the harder tiers still separate anything.

On "The Last Ones (TLO)", a 32-step cyber range built with SpecterOps across four subnets and roughly twenty machines that AISI estimates would take a human expert about 20 hours, GPT-5.5 completed the full attack chain in 2 attempts out of 10, second to Mythos Preview at 3/10. AISI red-teamers then found a universal jailbreak in six hours, eliciting offensive content across the entire set of malicious cyber requests OpenAI supplied. A configuration problem prevented AISI from verifying the safeguards actually shipped.

The same model reads differently on other terrain. Dan Shipper cites 62/100 on the Senior Engineer benchmark against a human range of 80-90. Artificial Analysis places GPT-5.5 at xhigh setting on 1509 Elo on GDPval-AA, tied with the open-weights GLM-5.2 at 1524 and well behind Claude Fable 5. By July 2026 GPT-5.5 was the outgoing flagship, its price inherited by GPT-5.6 Sol.

Type
Technology
score_expert_pass1
71,4% (+-8,0%)
score_expert_pass5
90,5% (+-12,9%)
relations
9
Cited in
1 fiches

Adoption measures

Neighborhood

GPT-5.4 The Last Ones (TLO) Opus 4.7 jailbreak universel Mythos Preview GPT-5.6 GLM-5.2

→ outperforms

GPT-5.4 TECHNOLOGIE high confidence stable Source ↗
Opus 4.7 TECHNOLOGIE high confidence stable Source ↗

→ solves

The Last Ones (TLO) TECHNOLOGIE high confidence stable Source ↗

← outperforms

jailbreak universel CONCEPT high confidence stable Source ↗

→ converges with

Mythos Preview TECHNOLOGIE high confidence stable Source ↗

← replaces

GPT-5.6 TECHNOLOGIE high confidence stable Source ↗

← competes with

GLM-5.2 TECHNOLOGIE high confidence evolving Source ↗

Cited in (1)