# sfeir-code-review-anneau-contraintes-2026-07-30

## Veille

Episode "Phase 5 · Review" of the SFEIR series on the augmented SDLC, published **the same day** as the Addy Osmani LinkedIn post that it translates into a phase specification. Thesis: **quality has changed address** — it is no longer read in the code (agents produce more of it than anyone can review) but in **the ring of constraints surrounding the agent**. Osmani's ring (seven dimensions — correctness, security, performance, accessibility, maintainability, **economic efficiency**, **comprehensibility** — linked by the **back-pressure** rule: "a loop is only granted the autonomy that can be verified cheaply and reliably, not an inch more") is redrawn, translated, and attached to phase 5 of SFEIR's 11-phase cycle. The structuring corollary: **the bottleneck has never been generation, it is verification** — "generation is a wide mouth, verification a narrow neck; speeding up the mouth thickens the pile at the neck." **The most interesting design decision is a cycle-architecture choice**: Review is deliberately **outside the three human gates** (Define, Plan, Ship), because making Review the gate would put human attention — a finite resource — as the control point of a generation capacity that itself scales: "you would have built a pipeline whose maximum throughput is the number of diffs a senior can read before the end of the day." Hence the split: **Review instruments, Ship decides** — Review delivers an *opposable body of evidence*, Ship decides on the evidence, not on the full diff. A position staked against Monperrus (from whom SFEIR retains the diagnosis — human inspection of every diff cannot withstand agentic speed — but rejects the conclusion: acceptance cannot be delegated). The named trap is **circular validation** (the agent that writes the code writes the tests that validate it: "you built a mirror, not a ring"), with five countermeasures drawn from Anthropic (independent gates in separate context windows, deterministic + agentic never substituting for one another, shadow mode, risk-based tiering, logging to the SIEM) and Compare the Market's warning (**AST graph ~70% vs vector RAG ~58%**, with RAG performing *worse than no context at all*). The firm's own extension is **the ratchet**: "every escape becomes a constraint" — a defect that has crossed the ring is closed *within the ring* (test, lint rule, review rubric, harness guardrail) at Compound-1, "the only asset in the chain that appreciates while the models depreciate" (an unaudited internal measurement: **−30% fix iterations after ten cycles**). It closes by reformulating the question: "is this code good?" has become unanswerable; what remains is **"what does my system refuse to let through?"**

## Titre Article

Code review dans le SDLC augmenté : l'anneau de contraintes autour des agents

## Date

2026-07-30

## URL

https://www.sfeir.com/articles/sdlc-ai-review-anneau-contraintes/

## Keywords

ring of constraints, constraints around agents, Review phase, phase 5, augmented SDLC, SFEIR 11-phase cycle, human gates, Define Plan Ship, Review instruments Ship decides, opposable body of evidence, back-pressure, autonomy ≤ verifiability, low-cost verifiability, verification bottleneck, wide mouth narrow neck, more code than you can read, Addy Osmani, harness, harness engineering, Software Factories Light and Dark, dark factory, lights-out factory, seven constraint dimensions, correctness, property tests, mutation tests, green/red oracle, security, SAST, DAST, dependency scanning, secrets detection, performance, performance budget, load testing, SLO, accessibility, axe, contrast, keyboard navigation, maintainability, coverage, complexity, blast radius, component boundaries, architectural debt, economic efficiency, token budget, cost per change, FinOps token, TCO, CapEx OpEx, comprehensibility, comprehension debt, comprehension debt, decision log, decision log attached to the PR, intent is not lost it is discarded, Martin Monperrus, end of code review, false sense of security, from creator to verifier, circular validation, circular validation, mirror, Augment Code, Paula Hingel, DORA 2025, throughput vs stability, Anthropic, Jason Clinton, 80% of merged code, independent gates, separate context windows, deterministic and agentic, shadow mode, shadow mode, red team, risk-tiered codebase, risk-weighted sampling, SIEM logging, 16% to 54%, a third of incidents, what would we run if scanning cost almost nothing, Compare the Market, context retrieval, structural knowledge graph, AST analysis, 70% vs 58%, 79 merge requests, vector RAG, structural understanding, callers signatures hierarchies, ratchet, ratchet, every escape becomes a constraint, a bug seen twice is a hole in the system, Compound-1, compounding static lessons, memory reloaded at Plan, −30% fix iterations, ten cycles, where to place the switch, lights off, cheap high-frequency hard-to-bypass immediate non-drifting control, short loops, three to ten steps, context accumulation, authentication billing public API contract, execution output vs agent's declaration, constraints the model cannot negotiate, what does my system refuse to let through

## Authors

SFEIR (voix éditoriale du cabinet, article non signé individuellement) — construit sur Addy Osmani (Google) ; cite Martin Monperrus, Paula Hingel (Augment Code), DORA/Google Cloud, Jason Clinton (Anthropic), l'équipe Engineering de Compare the Market

## Ton

**Profile**: an operational-use doctrine article, an episode in a numbered series ("Phase 5 · Review"). Audience: engineering leadership, architects, CISOs who already tool agents and are looking for where to place controls. Dense professional register, polished French, zero undefined jargon, no anglicism left untranslated (*back-pressure* and *comprehension debt* are kept but glossed).

**Style**: the **clincher line** at the end of each paragraph is the text's engine, and it is almost always a concrete image rather than an abstraction — "generation is a wide mouth, verification a narrow neck; speeding up the mouth thickens the pile at the neck"; "you built a mirror, not a ring of constraints"; "a lights-out factory doesn't pay down that debt, it takes it on at full throttle, tests green"; "intent is not lost, it is discarded"; "a loop that sprawls hides its errors in the corners"; "a static ring is a leaking ring." A funnel-shaped rhetorical architecture: a borrowed diagram → its translation into a two-column table (mechanizable / irreducibly human) → the named failure mode → the firm's own extension → five self-diagnostic questions → a closing question that replaces the opening one. Two rigor markers unusual for the genre: **every external claim carries a numbered note** with a status label (*Industry · Anthropic*, *Measured · SFEIR*), and the central diagram is **explicitly credited** ("SFEIR diagram after Addy Osmani's diagram, © Addy Osmani, redrawn, translated, and attached to phase 5").

**Epistemic position**: assumed as prescriptive, but with two gestures of probity. The first is **explicit, argued disagreement** with Monperrus: SFEIR takes his diagnosis, rejects his conclusion, and says why (acceptance cannot be delegated — that's the Ship gate) rather than ignoring or caricaturing it. The second is the **usage caveat repeated on Anthropic's figures** ("the figures come from Anthropic about itself"), already set out in the July 26 fiche. The text's limitation mirrors its strength: **it is published the same day as the Osmani post it comments on**, leaving little room for the test of time or for contradiction — and its only original figure (−30% iterations after ten cycles) is first-party, unaudited, with no described protocol or scope. The final CTA ("Instrument your Review phase before opening the agentic floodgates") is a reminder that the doctrine is also an offer.

## Pense-betes

- **Date / source**: **July 30, 2026**, SFEIR, unsigned (firm voice), tags `sdlc`, `ia-agentique`, `software-factory`, `harness-engineering`, `code-review`. Eight numbered sources with status labels.
- **Nature**: **framing article**, not a primary source of results. Three original contributions; everything else is sourced synthesis. This fiche should not be re-cited for Anthropic's or Compare the Market's figures — go to the source fiches. ### The three original contributions 1. **Review outside the human gates**, with the split *"Review instruments / Ship decides"*. 2. **Attaching the ratchet to Compound-1** as the phase explicitly responsible for it. 3. The firm's own figure of **−30%**. ### The dimension-by-dimension grid Directly transposable into a Review phase specification: dimension → mechanizable constraint → residual human judgment. | Dimension | Mechanizable constraint | Residual human judgment | |---|---|---| | Correctness | unit, property, mutation testing, green/red oracle | functional acceptance | | Security | SAST/DAST, dependencies, secrets, dedicated agentic review | arbitrating residual risk | | Performance | performance budget, load, measured regression | defining the SLO | | Accessibility | axe, contrast, keyboard | the actually lived experience | | Maintainability | coverage, complexity, blast radius, component boundaries | assumed architectural debt | | Comprehensibility | the agent logs what it tried **and what it discarded**, decision log attached to the PR | reconstructing intent | | Economic efficiency | tokens/compute budget per task, cost per change | TCO, the CapEx/OpEx trade-off | The rule linking all seven: **back-pressure** — autonomy ≤ cheaply verifiable — whose residual judgment is *where to place the switch*. ### The architecture argument If Review carried the human gate, *"the system's control point would be human attention, a finite resource that doesn't scale, facing a generation capacity that does scale"*: the neck would never widen, and *"you would have built a pipeline whose maximum throughput is the number of diffs a senior can read before the end of the day."* This is what justifies **moving** the gate to Ship rather than removing the human. Operative distinction: Review has a **deliverable** (an opposable body of evidence), Ship has a **decision** — *"and that decision is made on the evidence, not the full diff."* ### The forgotten dimension **Comprehensibility** is systematically omitted *"because it doesn't break CI."* The remedy is the cheapest on the grid and the least applied: asking the agent to write down what it tried and discarded, since *"reviewing an agentic PR is the first time a human reconstructs the why."* A phrase worth keeping: *"intent is not lost, it is discarded."* This is the structural guardrail that [[osmani-cognitive-surrender-comprehension-debt-2026-05-05]] called for. ### Circular validation The agent writes the code, the same agent writes the tests, CI is green: *"a mirror, not a ring."* The phenomenon is measured by DORA 2025 — AI adoption correlated **positively with throughput** and **negatively with stability** when the foundations don't keep up. The article's most cost-effective diagnostic question: *"are your tests written by the agent that writes the code?"* ### Qualifying criteria for an autonomous loop Control must be **cheap, high-frequency, hard to bypass, immediate, and non-drifting**. Qualifying: green/red oracle, type gate, property tests, an agentic reviewer with a real rubric. Corollary: short loops verify better than long ones — an agent holds up for 3 to 10 steps and loses the thread beyond about twenty; *"a loop that sprawls hides its errors in the corners."* Where to keep the light on: subtle bugs invisible to tests, wide blast radii, decisions structuring a year of work — namely **authentication, billing, public API contracts**. Governance warning: *"the real risk is setting every switch the same way"* — everything off leads to dismantling four months later, everything on halts delivery. ### The ratchet, SFEIR's own extension *"Every escape becomes a constraint"*: a defect that has crossed the ring is not just fixed in the code, it is **closed within the ring** (test, lint rule, review rubric, harness guardrail), at **Compound-1**, with memory reloaded at the Plan of the next cycle. This is Osmani's *ratchet principle* wired into a phase of the cycle explicitly responsible for it. Two formulas: *"a checklist is written once and goes stale, the ring thickens with every cycle"* and *"a bug seen twice is not a bug, it's a hole in the system."* Associated investment argument: the ring is *"the only asset in the chain that appreciates while the models depreciate."* ### The firm's own figure **−30% fix iterations after ten cycles**, labeled *Measured · SFEIR*, first-party material, 2026. No protocol, scope, sample size, or definition of "fix iteration" is given. This is the article's only original figure and it supports its most commercially useful thesis. Do not reuse it without qualification. ### Economic efficiency in a quality grid *"A per-task token budget is a constraint on the same footing as a performance budget, and it produces the same virtue: bounding autonomy by the cost of verifying it."* A junction between FinOps and quality. ### A citation variance worth knowing SFEIR references Monperrus under the title *"The End of Code Review: How AI Agents Supersede Human Code Review."* The title carried by the corpus (arXiv 2606.13175) is *"The End of Code Review: Coding Agents Supersede Human Inspection."* Same paper, same date (June 11, 2026); it's SFEIR's label that drifts — use the arXiv title for formal citation. SFEIR takes the diagnosis from [[monperrus-end-of-code-review-agents-supersede-2026-06-11]] and rejects its conclusion. ### Cited sources absent from the corpus, candidates for addition Two Osmani texts: *Set the constraints around your agents* (LinkedIn, **July 30, 2026** — the original diagram) and above all ***Software Factories, Light and Dark*** (addyosmani.com, July 2026), which alone carries three of the structuring concepts taken up here: the back-pressure principle, operationalized *comprehension debt*, and the exploitable length of loops. Also absent: *Code review à l'ère de l'IA : du créateur au vérificateur* (SFEIR, April 1, 2026). **Watch-item to note**: published the same day as the commented Osmani post.

## RésuméDe400mots

Fifth episode in SFEIR's series on the augmented SDLC, devoted to the Review phase, published the same day as the Addy Osmani LinkedIn post it converts into a phase specification.

The starting observation: quality used to be read in the code; agents now produce more of it than anyone can review. It has therefore **changed address** — it now lives in **the ring of constraints** surrounding the agent, that is, in the harness. Seven dimensions make up this ring (correctness, security, performance, accessibility, maintainability, economic efficiency, comprehensibility), linked by the **back-pressure** rule: a loop is only granted the autonomy that can be verified cheaply and reliably. The corollary overturns the dominant intuition: the bottleneck has never been generation, it is verification — "generation is a wide mouth, verification a narrow neck; speeding up the mouth thickens the pile at the neck."

Hence the central architecture decision: in the eleven-phase cycle, **Review is not a human gate**, and this is deliberate. The three inviolable gates are Define, Plan, and Ship. Making Review carry the gate would put human attention — a finite resource — as the control point of a generation that itself scales: the neck would never widen. **Review instruments, Ship decides**; Review produces an opposable body of evidence, and the decision is made on the evidence, not the full diff. SFEIR retains from Monperrus that human inspection of every diff cannot withstand agentic speed, but rejects his conclusion: acceptance cannot be delegated.

The operational translation is a dimension-by-dimension table, separating what can be mechanized from irreducibly human judgment. The dimension systematically forgotten is **comprehensibility**, "because it doesn't break CI" — hence the cheapest remedy on the grid: having the agent log what it tried and discarded, since "intent is not lost, it is discarded."

The named failure mode is **circular validation**: the agent that writes the code writes the tests that validate it, CI is green, "you built a mirror, not a ring." Five countermeasures are drawn from Anthropic (independent gates, deterministic + agentic, shadow mode, risk-based tiering, SIEM logging), and Compare the Market warns that a reviewer built on vector RAG degrades review quality (~70% for an AST graph versus ~58%).

The firm's own extension is **the ratchet**, attached to Compound-1: every escape becomes a constraint. The ring thickens with every cycle — "the only asset in the chain that appreciates while the models depreciate" (−30% fix iterations after ten cycles, internal measurement). Only one question remains: **what does my system refuse to let through?**

## GrapheDeConnaissance

- SFEIR —affirme_que→ la qualité logicielle ne se lit plus dans le code mais dans l'anneau de contraintes qui entoure l'agent (AFFIRMATION, 0.98)
- anneau de contraintes —est_basé_sur→ Addy Osmani (PERSONNE, 0.97)
- SFEIR —affine→ anneau de contraintes (CONCEPT, 0.96)
- anneau de contraintes —fait_partie_de→ cycle à 11 phases (METHODOLOGIE, 0.95)
- back-pressure —fait_partie_de→ anneau de contraintes (CONCEPT, 0.96)
- Addy Osmani —affirme_que→ "on ne confie à une boucle que l'autonomie qu'on sait vérifier à faible coût et de façon fiable, pas un pouce de plus" (CITATION, 0.96)
- SFEIR —affirme_que→ le goulot n'a jamais été la génération mais la vérification : accélérer la génération ne fait qu'épaissir le tas au col de la vérification (AFFIRMATION, 0.97)
- phase Review (SDLC) —fait_partie_de→ cycle à 11 phases (METHODOLOGIE, 0.97)
- SFEIR —affirme_que→ "Review instrumente. Ship décide." — Review livre un faisceau de preuves opposable, Ship décide sur les preuves et non sur le diff intégral (CITATION, 0.97)
- SFEIR —affirme_que→ placer le gate humain sur Review ferait de l'attention humaine le point de contrôle d'une génération qui scale, plafonnant le débit au nombre de diffs qu'un senior peut lire dans une journée (AFFIRMATION, 0.96)
- SFEIR —s_oppose_à→ The End of Code Review: Coding Agents Supersede Human Inspection (DOCUMENT, 0.93)
- Martin Monperrus —affirme_que→ le modèle hybride "l'agent écrit, l'humain relit" est intenable et générateur d'une fausse sécurité (AFFIRMATION, 0.95)
- validation circulaire —est_instance_de→ mode d'échec de la revue agentique (CONCEPT, 0.95)
- Augment Code —référence→ validation circulaire (CONCEPT, 0.94)
- DORA 2025 —mesure→ "l'adoption de l'IA est corrélée positivement au débit de livraison et négativement à la stabilité quand les fondations ne suivent pas" (MESURE, 0.93)
- compréhensibilité —réduit→ comprehension debt (CONCEPT, 0.94)
- journal de décision —permet→ reconstruction de l'intention derrière une PR agentique (CONCEPT, 0.95)
- SFEIR —affirme_que→ "l'intention n'est pas perdue, elle est jetée" (CITATION, 0.95)
- budget de tokens par tâche —est_instance_de→ contrainte de qualité au même titre qu'un budget de performance (CONCEPT, 0.93)
- cliquet de l'anneau —est_basé_sur→ ratchet principle (CONCEPT, 0.95)
- cliquet de l'anneau —fait_partie_de→ Compound-1 (CONCEPT, 0.94)
- SFEIR —recommande→ toute échappée devient une contrainte : un défaut qui a franchi l'anneau se referme dans l'anneau sous forme de test, de règle de lint, de rubrique de revue ou de garde-fou de harnais (AFFIRMATION, 0.97)
- SFEIR —affirme_que→ "l'anneau s'épaissit à chaque cycle, et c'est le seul actif de la chaîne qui s'apprécie pendant que les modèles se déprécient" (CITATION, 0.95)
- SFEIR —mesure→ "− 30 % d'itérations de correction après dix cycles" (mesure interne first-party, protocole non publié) (MESURE, 0.78)
- Compare the Market —mesure→ "le graphe de connaissance structurel construit par analyse de l'AST place un commentaire en ligne pertinent dans environ 70 % des cas contre 58 % pour le RAG vectoriel, sur 79 merge requests" (MESURE, 0.92)
- graphe de connaissance structurel —surpasse→ RAG vectoriel (TECHNOLOGIE, 0.93)
- SFEIR —recommande→ qualifier une boucle pour l'autonomie sur cinq critères de contrôle — peu coûteux, à haute fréquence, difficile à contourner, immédiat et non dérivant (AFFIRMATION, 0.95)
- boucles courtes —améliore→ vérifiabilité d'un agent (CONCEPT, 0.92)
- SFEIR —recommande→ garder la revue humaine sur l'authentification, la facturation et les contrats d'API publics, et réviser chaque interrupteur à chaque Compound (AFFIRMATION, 0.94)
- SFEIR —recommande→ remplacer la question "ce code est-il bon ?" par "qu'est-ce que mon système refuse de laisser passer ?" (AFFIRMATION, 0.96)

---
Canonical: https://www.thekb.eu/en/fiches/sfeir-code-review-anneau-contraintes-2026-07-30/
