- thekb.eu
- Knowledge graph
- Technology
- Opus 4.7
Opus 4.7
Opus 4.7 — Technology. capability: Refactor 100k lines, find zero-days — but jagged on simple questions · default_reasoning_effort: xhigh (after correction) · score_expert_pass1: 48,6% (+-10,0%)
48.6% (+-10.0%) at pass@1 on expert capture-the-flag tasks: that is where the UK AI Safety Institute placed Opus 4.7 in its 30 April 2026 evaluation, behind GPT-5.4 (52.4%) and well behind GPT-5.5 (71.4%). The same model refactors 100k lines of code and finds zero days. Andrej Karpathy uses exactly that gap as his marker for jaggedness: Opus 4.7 handles the refactor, then advises walking 50m to the car wash. His account of why is verifiability, since labs run reinforcement learning on domains where answers can be checked, which builds peaks in math and code and troughs elsewhere.
Anthropic's post-mortem of April 2026 shows behaviour swinging on configuration rather than weights. The default reasoning effort dropped from high to medium between 4 March and 7 April to fix interface freezes; users said Claude Code felt less intelligent, and Anthropic reversed course, setting xhigh for Opus 4.7 and high for the others. A separate system-prompt instruction capping verbosity cost 3% intelligence in ablation testing, for Opus 4.6 and Opus 4.7 alike. The same document credits Opus 4.7's Code Review tool, loaded with full repository context, with the identification rétrospective du bug de cache that Opus 4.6 had failed to produce.
Thariq Shihipar leans on a different property. HTML output costs more tokens than Markdown, and the 1MM context window of Opus 4.7 absorbs the difference. Generation still runs two to four times slower.
- Type
- Technology
- capability
- Refactor 100k lines, find zero-days — but jagged on simple questions
- default_reasoning_effort
- xhigh (after correction)
- score_expert_pass1
- 48,6% (+-10,0%)
- relations
- 4
- Cited in
- 3 fiches