# pragmatic-engineer-measure-ai-impact-dev-2025-09-16

## Veille

Pragmatic Engineer - Measuring AI Impact - Developer Productivity - Metrics - GitHub Copilot - DX - Engineering Efficiency

## Titre Article

HOW TECH COMPANIES MEASURE THE IMPACT OF AI ON SOFTWARE DEVELOPMENT

## Date

2025-09-16

## URL

https://newsletter.pragmaticengineer.com/p/how-tech-companies-measure-the-impact-of-ai?utm_source=tldrnewsletter

## Keywords

AI impact, software development, engineering efficiency, developer productivity, AI tools, metrics, GitHub Copilot, Google, Microsoft, Dropbox, Monzo, Atlassian, DX, AI Measurement Framework, Change Failure Rate, PR throughput, developer experience, CSAT, time savings

## Authors

Gergely Orosz and Laura Tacho

## Ton

**Profile:** Professional-Analytical | Expert co-authors | Educational-Prescriptive | Intermediate-Expert

Orosz and Tacho adopt a collaborative expert voice combining reporting with methodological framework building. Data drawn from 18 companies (Google, GitHub, Microsoft, Dropbox) empirically grounds the recommendations. The systematic structure blending core metrics and AI-specific metrics reveals framework-oriented thinking. Concrete case studies (90% adoption at Dropbox, BDD at Microsoft, Monzo's challenges) illustrate abstract principles. The prescriptive language provides actionable guidance. Explicit warnings (quality risks, maintainability debt, limits of acceptance rate) demonstrate intellectual honesty. The article targets engineering leaders with a mix of strategic thinking and tactical implementation. Typical of Pragmatic Engineer's in-depth analyses combining industry research with practical recommendations.

## Pense-betes

- **18 major tech companies** studied (Google, GitHub, Microsoft, Dropbox, Monzo, Atlassian...)
- **85% of engineers use AI tools**, but lack clear metrics to justify the investment
- **Core + AI-specific metrics**: combine existing ones (CFR, PR throughput, PR cycle time, developer experience) with new ones (adoption rate, CSAT, time saved, AI spend)
- **Dropbox results**: **90% AI adoption**, engineers merge **20% more PRs** with a reduced CFR
- **Segment the data**: AI users vs non-AI users, before/after AI, by role/seniority/language
- **Balance speed and quality**: track metrics that check each other (PR throughput + CFR)
- **Developer experience as a priority**: measuring satisfaction and experience is crucial for sustainable adoption
- **3-layer data collection**: system data + periodic surveys + experience sampling
- **Experimental mindset**: approach measurement with a clear goal, test predictions
- **"Bad developer days" (BDD)**: Microsoft metric assessing AI's impact on daily friction
- **Decline of acceptance rate**: no longer a benchmark metric, captures neither maintainability nor bugs
- **Agent telemetry**: emerging area set to evolve significantly
- **Monzo case**: objective measurement is difficult (data retention by vendors), subjective sentiment + specific use cases (code migrations) demonstrate clear value

## RésuméDe400mots

This in-depth analysis explores how **18 major tech companies**, including Google, GitHub, Microsoft, and Dropbox, measure the impact of AI on software development, amid the challenge of justifying growing investments in AI coding tools. Written by Gergely Orosz and Laura Tacho (CTO of DX), the article notes that while **85% of engineers use AI tools**, many engineering leaders struggle to assess their real value, lacking clear metrics beyond superficial measures such as lines of code (LOC).

**Central message: combine metrics**

Effectively measuring AI impact requires **combining existing 'core' engineering metrics with new AI-specific metrics**. Companies should not abandon traditional metrics such as Change Failure Rate, PR throughput, PR cycle time, and developer experience, since the ultimate goal of AI is precisely to improve these software delivery fundamentals. These core metrics must be tracked alongside AI adoption rates, satisfaction (CSAT) with the tools, time saved per engineer, and AI spend. **Dropbox**, for example, reached **90% AI adoption** and saw its engineers merge **20% more pull requests** with a reduced change failure rate.

**Segmentation and an experimental mindset**

A crucial aspect is **breaking down metrics by level of AI usage**: comparing AI users to non-AI users, and analyzing trends over time. This breakdown by role, seniority, or programming language helps identify which groups benefit most from AI or need additional training. The article emphasizes an **experimental mindset**, where data is used to answer specific questions and test predictions about AI's influence.

**Quality, maintainability, developer experience**

Vigilance over **code quality, maintainability, and developer experience** is paramount. The authors warn that AI-assisted development can create "the biggest pile of technical debt" if not managed carefully. It is essential to track metrics that check each other, such as speed alongside quality (PR throughput and CFR). Beyond system metrics, self-reported data on "confidence in changes," "code maintainability," and "perceived quality" are vital for capturing long-term impacts. Developer experience, often wrongly reduced to superficial perks, is critical for reducing friction across the entire development cycle.

**Emerging trends and challenges**

Microsoft uses **"bad developer days" (BDD)** to assess AI's impact on daily friction, while Glassdoor measures experimentation outcomes (A/B tests). The **acceptance rate** of AI suggestions, once a benchmark metric, is declining because it is too narrow: it captures neither maintainability, nor bug introduction, nor overall productivity. Cost analysis, still rarely practiced so as not to discourage usage, is expected to receive greater scrutiny as AI budgets grow. **Agent telemetry** and measurement beyond code writing are identified as areas set to evolve significantly.

**AI Measurement Framework and data layers**

The article introduces the **AI Measurement Framework**, a recommended set of metrics blending AI metrics with core engineering metrics, with developer experience at its center. It advocates layered data collection: quantitative system data (AI tools, GitHub, JIRA, CI/CD), periodic qualitative surveys, and in-the-moment experience sampling. **Monzo Bank**'s experience serves as a case study: objective measurement is difficult (data retention by vendors), but engineers' subjective sentiment and specific use cases such as code migrations demonstrate clear value.

## GrapheDeConnaissance

- Gergely Orosz —publie→ AI Measurement Framework (METHODOLOGIE, 0.97)
- Laura Tacho —publie→ AI Measurement Framework (METHODOLOGIE, 0.97)
- Laura Tacho —travaille_chez→ DX (ORGANISATION, 0.98)
- DX —permet→ mesure de l'efficacité ingénierie en entreprise (CONCEPT, 0.95)
- AI Measurement Framework —recommande→ combiner métriques d'ingénierie core et métriques spécifiques IA (AFFIRMATION, 0.96)
- Dropbox —mesure→ 90% taux d'adoption IA (MESURE, 0.98)
- Dropbox —mesure→ augmentation 20% des PRs fusionnées (MESURE, 0.95)
- Microsoft —utilise→ Bad Developer Days (METHODOLOGIE, 0.97)
- difficultés de mesure objective de l'IA —observé_dans→ Monzo Bank (ORGANISATION, 0.93)
- acceptance rate —s_oppose_à→ mesure pertinente de productivité IA (CONCEPT, 0.88)
- LeadDev —publie→ AI Impact Report 2025 (DOCUMENT, 0.96)
- METR study —s_oppose_à→ perception de gain de vitesse IA (CONCEPT, 0.9)
- LOC —s_oppose_à→ mesure pertinente productivité (CONCEPT, 0.92)
- Gergely Orosz —prédit→ une évolution significative de la télémétrie d'agents (AFFIRMATION, 0.82)

---
Canonical: https://www.thekb.eu/en/fiches/pragmatic-engineer-measure-ai-impact-dev-2025-09-16/
