# denisov-blanch-stanford-quantify-ai-roi-software-engineering-2025-11-23

## Veille

Stanford Research - AI ROI Measurement - Developer Productivity - Metrics Framework - AI Adoption Gap

## Titre Article

How to Quantify AI ROI in Software Engineering (Stanford Study / 120k Devs)

## Date

2025-11-23

## URL

https://www.youtube.com/live/cMSprbJ95jg?si=4HnxK8w1ELvSr4tz&t=9705

## Keywords

Stanford, AI ROI, Developer Productivity, Metrics, Code Quality, Tech Debt, AI Adoption, Engineering Output, Rework

## Authors

Yegor Denisov-Blanch (Researcher, Stanford)

## Ton

**Profile:** Academic-Research | Data-Driven | Analytical | Objective

The tone is that of rigorous academic research applied to industry. The approach is grounded in quantitative evidence (longitudinal and cross-sectional studies across 120k developers). The discourse aims to demystify the "hype" through solid data, proposing sophisticated measurement frameworks (ML models to assess output, code cleanliness indices). It is a call for metric rigor addressed to technical leaders.

## Pense-betes

- **Widening Gap**: Teams that perform well with AI improve even further ("rich gets richer"), while struggling teams fall further behind. Knowing which cohort you belong to is critical.
- **Quality of use > Quantity**: The correlation between the number of tokens used and productivity is weak. What matters is *how* AI is used (engineering patterns).
- **Code Hygiene (Cleanliness Index)**: Clean, modular, tested code amplifies AI's gains. Conversely, AI used on messy code accelerates entropy and technical debt.
- **Measuring ROI**:
- Do not rely on business results (too much noise/confounding factors).
- Do not rely on simplistic metrics (lines of code, number of PRs). An increase in PRs can mask a drop in quality and an increase in "rework" (rework).
- **Proposed metrics framework**:
- **Primary metric**: Engineering Output (assessed by an ML model simulating a panel of experts, not just volume).
- **"Guardrail" metrics**: Quality, Rework/Refactoring, Team health (DevOps metrics). The goal is to maximize output while keeping guardrails healthy.
- **Case study (Warning)**: A company saw +14% PRs with AI, but a 9% drop in quality and a 2.5x explosion in rework. Without fine-grained measurement, this looked like a success when the ROI was likely negative.

## RésuméDe400mots

Yegor Denisov-Blanch, a researcher at Stanford, presents the results of a large-scale study on the impact of AI on the productivity of more than 100,000 developers. The objective is to move past the "hype" and quantify the actual return on investment (ROI).

The study reveals a **widening gap**: the performance difference between teams that succeed with AI and those that fail is growing. Contrary to popular belief, the amount of AI used (number of tokens) correlates poorly with productivity. The determining factor is the **quality of the code environment** ("Codebase Hygiene"). Clean, well-tested, well-documented code allows AI to perform well. Conversely, on a degraded codebase, AI risks accelerating entropy and technical debt, requiring more human effort to correct errors.

Denisov-Blanch warns against simplistic metrics such as the number of Pull Requests (PRs). He cites the example of a company that observed a 14% increase in PRs, initially interpreted as a success. A more detailed analysis revealed a 9% drop in code quality and a 2.5x increase in "rework" (rework on recently written code). The gain in volume was offset by the drop in quality, making the ROI potentially negative.

To measure ROI correctly, Stanford proposes a framework:
1.  **Measure Usage**: Distinguish theoretical access from actual usage (via fine-grained telemetry) and identify usage "patterns" (personal use vs. agentic orchestration).
2.  **Measure "Engineering Outcomes"**: Use a primary **"Engineering Output"** metric (based on an ML model trained to replicate human expert evaluation, rather than lines-of-code volume) coupled with **"Guardrail" metrics** (quality, rework rate, team health) that must be kept at a healthy level.

The conclusion is that AI is an amplifier: it accelerates both good and bad practices. Leaders must precisely measure impact to course-correct and invest in code hygiene to unlock AI's potential.

## GrapheDeConnaissance

- Yegor Denisov-Blanch —a_créé→ étude ROI IA développement (EVENEMENT, 0.95)
- Stanford —emploie→ Yegor Denisov-Blanch (PERSONNE, 0.95)
- étude Stanford —mesure→ 120 000 développeurs analysés (MESURE, 0.95)
- hygiène du code —améliore→ gains IA (CONCEPT, 0.92)
- code sale —permet→ dette technique avec IA (CONCEPT, 0.9)
- métriques simplistes —permet→ masquage de la baisse qualité réelle (CONCEPT, 0.88)
- Engineering Output —remplace→ métriques volume (lignes, PRs) (CONCEPT, 0.85)
- IA développement —permet→ creusement écart équipes performantes/faibles (CONCEPT, 0.9)
- étude Stanford —affirme_que→ la quantité de tokens corrèle faiblement avec la productivité développeur (AFFIRMATION, 0.88)
- Yegor Denisov-Blanch —recommande→ framework métriques guardrails (METHODOLOGIE, 0.88)

---
Canonical: https://www.thekb.eu/en/fiches/denisov-blanch-stanford-quantify-ai-roi-software-engineering-2025-11-23/
