# ng-the-batch-350-coding-agents-software-work-acceleration-2026-04-24

## Veille

Andrew Ng's editorial in The Batch #350 sets out an **acceleration hierarchy for coding agents** by type of software work: **Frontend (max) > Backend (moderate) > Infrastructure (low) > Research (minimal)**. The rationale rests on implicit *verifiability* (fluency in TypeScript/JavaScript plus an autonomous agent–browser test loop on the frontend) and on the LLMs' blind spots (corner cases / security / DB migrations for backend, opaque network tradeoffs for infra, irreducible hypothesis formation for research). The issue is rounded out by 4 structuring news items: **GLM-5.1 (Z.ai)**, a 754B/40B-active-parameter MIT-licensed model capable of autonomous tasks lasting 8 hours (SWE-Bench Pro leader at 58.4%); **Digit (Agility Robotics) at Schaeffler**, the first industrial deployment of humanoids (5'9"/143lb, $10–25/h vs $20/h for a human); the **anti-data-center revolt** (~$64B blocked May 2024 – March 2025, Maine moratorium on 20MW+ facilities, molotov cocktail at Sam Altman's home); and the **"assistant axis"** (Christina Lu, MATS / Oxford / Anthropic), which reduces persona drift and jailbreaks (Qwen3 32B: 83%→41%; Llama 3.3 70B: 65%→33%) without degrading IFEval/GSM8k/MMLU-Pro/EQ-Bench.

## Titre Article

The Batch n°350 — How Coding Agents Accelerate Different Types of Software Work (Andrew Ng) + GLM-5.1, Digit chez Schaeffler, anti-data-center revolt, assistant axis

## Date

2026-04-24

## URL

https://www.deeplearning.ai/the-batch/issue-350/

## Keywords

Andrew Ng, The Batch, DeepLearning.AI, coding agents, differential acceleration, frontend development, backend development, infrastructure, research, TypeScript, JavaScript, autonomous browser, corner cases, security, DB migrations, hypothesis formation, GLM-5.1, Z.ai, 754B parameters, 40B active, MoE, SWE-Bench Pro, Arena Code, CyberGym, GPQA Diamond, MIT license, HuggingFace, prompt caching, Agility Robotics, Digit, Schaeffler, humanoid robot, manufacturing, South Carolina, RGB depth cameras, LiDAR, IMU, McKinsey 5 million humanoids 2040, anti-data-center, Maine moratorium, Port Washington Wisconsin, Festus Missouri, Ohio constitutional amendment, molotov cocktail Altman, Indianapolis shooting, closed-loop cooling, off-grid generation, Christina Lu, MATS, Oxford, Anthropic, assistant axis, activation capping, persona drift, jailbreak resistance, Qwen3 32B, Llama 3.3 70B, Gemma 2 27B, IFEval, GSM8k, MMLU-Pro, EQ-Bench, acceleration hierarchy, verifiability, jagged frontier

## Authors

Andrew Ng (édito principal — fondateur DeepLearning.AI, Stanford, ex-Google Brain, ex-Baidu) ; rédaction The Batch (DeepLearning.AI) pour les sections actualités

## Ton

**Profile**: Weekly AI newsletter, 5-section digest format (1 editorial + 4 news items), third-person journalistic perspective for the news + reflexive first person for the editorial. Target audience: ML/AI practitioners, engineers, tech leaders, the DeepLearning.AI community. Level: intermediate-accessible, without excessive jargon but with concrete technical specifications (parameters, benchmarks, API pricing).

**Style**: Ng's editorial adopts a pedagogical-prescriptive tone typical of the author — he sets out a **taxonomy** (frontend / backend / infra / research) and justifies it with structural arguments (LLMs' language fluency, autonomous test-loop capability, depth of domain knowledge required). No hype: Ng calibrates expectations, *"calibrate expectations and team organization around AI capabilities"*. Direct style, short propositions, key quotes in italics (*"coding agents are fluent in popular frontend languages like TypeScript and JavaScript"*, *"relatively limited"* for infra). This is **managerial education** more than a technical tutorial: Ng addresses decision-makers who must allocate resources and organize AI-augmented teams.

The news sections adopt a "fact-sheet" format: developer, innovation, specs, performance, market context. They are dense with figures (754B/40B, $1.40/$0.26/$4.40 per million tokens, 200 humanoids in 2026 → 5M by 2040, $64B blocked, 83%→41% jailbreaks). This numerical density is The Batch's signature: **structured watch reporting for practitioners** rather than long-form editorial analysis.

Taken together, the issue forms a **snapshot of applied AI in April 2026**: where agents accelerate (software by segment), where they gain autonomy (8-hour models), where they take physical form (industrial robotics), where they meet physical resistance (data centers), where they drift (alignment / persona).

## Pense-betes

- **Issue 350 of The Batch** published April 24, 2026, ~15-minute read, editorial by Andrew Ng.
- **Acceleration hierarchy for coding agents** (from most to least accelerated): 1. **Frontend** — *"coding agents are fluent in popular frontend languages like TypeScript and JavaScript"*. Autonomous agent-browser loop (the agent tests its outputs in a browser, iterates). When the design is specified, implementation is fast. 2. **Backend** — **moderate** acceleration. The developer must guide the model through *corner cases*, security considerations, and database migrations. Subtle bugs + downstream effects = experienced human oversight required. 3. **Infrastructure** — **the weakest** acceleration. LLMs have *"relatively limited"* knowledge of infra complexities and tradeoffs. Testing, experimentation, and debugging network misconfigurations demand deep engineering expertise beyond current agent capabilities. 4. **Research** — **minimal** acceleration despite the code-related gains. Agents speed up code generation and experiment orchestration, but the conceptual work (hypothesis formation, interpretation, iteration) remains human.
- **Management takeaway**: *"Understanding these distinctions helps organizations calibrate expectations and team organization around AI capabilities."*
- **Mirror reading with Karpathy** (fiche [karpathy-vibe-coding-agentic-engineering-software-3-0-2026-04-29](karpathy-vibe-coding-agentic-engineering-software-3-0-2026-04-29.md)): Ng proposes a hierarchy *by domain*, Karpathy an explanation *by verifiability* — domains with a strong verification signal (frontend visual rendering, math/code) peak; fuzzy domains (infra, conceptual research) lag. **Both frameworks are congruent.** ### News item 1 — GLM-5.1 (Z.ai): 8-hour autonomous agent
- **Architecture**: MoE, **754B total parameters**, **40B active per token**.
- **Context**: 200,000 tokens input, 128,000 output.
- **License**: MIT, open weights (HuggingFace).
- **API pricing**: $1.40 / $0.26 (cached) / $4.40 per million tokens (input/cached/output).
- **Distinctive capability**: autonomously sustains a single task for up to **8 hours**, with a plan→execute→evaluate loop and adaptive abandonment after hundreds of tool calls if the approach fails (instead of terminating prematurely).
- **Performance**: **SWE-Bench Pro** leader at 58.4% (vs. 54-57% for competitors); 3rd on Arena Code (1530 Elo); CyberGym record at 68.7; trails on GPQA Diamond at 86.2% vs. 94.3% for Gemini 3.1.
- **Market**: Z.ai raised its API pricing ~40% and doubled its coding subscription cost — narrowing the competitive gap with proprietary offerings. ### News item 2 — Digit (Agility Robotics) on Schaeffler's production line
- **First operational deployment** of humanoids in industry (South Carolina, auto parts).
- **Digit specs**: 5'9" (1.75m), 143lb (65kg), biped inverted-knee legs, 4-finger grippers, RGB depth + LiDAR + IMU sensors.
- **Regimen**: two 4-hour shifts with recharging; tasks specified as *workflows* (not direct motor commands) — bin transfer.
- **Economics**: Agility puts operating cost at **$10-25/h** vs. **~$20/h** for an entry-level human position. Schaeffler plans **hundreds of deployments** in the US + Europe by 2030.
- **Global context**: ~200 humanoids in factories in 2026; **McKinsey projection: 5 million by 2040**.
- **Employment effect**: research suggests **restructuring rather than elimination** — promotion into supervisory roles. ### News item 3 — Anti-data-center revolt
- **Scale**: ~**$64B** in data-center projects blocked or delayed between **May 2024 and March 2025**.
- **Legislation**: Maine — moratorium on facilities ≥20MW through 2027 (law pending governor's signature). Wisconsin (Port Washington) — first US referendum requiring a **popular vote** for tax incentives on megaprojects. Missouri (Festus) — voters **ousted** city council members who had approved a $6B data center. Ohio — proposed constitutional amendment banning facilities ≥25MW.
- **Grievances**: pressure on the power grid, rising residential energy rates, water consumption, noise nuisance, neighborhood impact, environmental footprint.
- **Violent incidents**: (1) **molotov cocktail thrown at Sam Altman's home** in San Francisco; (2) gunfire at the residence of an Indianapolis council member who had backed a $500M data center.
- **Technical mitigations**: water-efficient closed-loop cooling; growing private off-grid power generation.
- **Strategic tension**: tech companies view the data center as AI-sovereignty infrastructure against China — hence rapid expansion despite local resistance. ### News item 4 — "Assistant axis" (Christina Lu, MATS / Oxford / Anthropic)
- **Problem**: LLMs trained as assistants undergo **persona drift** in long or emotionally charged conversations — adopting alternate traits.
- **Solution**: an **"assistant axis"** = a vector derived from layer outputs that measures adherence to the trained assistant persona. Enables both detection AND correction of the drift.
- **Methodology**: 1,200 character-probing questions + 1,375 alternate system prompts; measurements on Gemma 2 27B / Qwen3 32B / Llama 3.3 70B; **"activation capping"** = constraining outputs within the assistant persona's parameters at inference.
- **Jailbreak results**:
- Qwen3 32B: harmful responses **83% → 41%**
- Llama 3.3 70B: harmful responses **65% → 33%**
- **Performance preservation**: IFEval, GSM8k, MMLU-Pro, EQ-Bench **stable or improved** — the strengthened alignment does not compromise capability.
- **Impact example**: a 30-turn conversation around suicidal ideation — the unmodified model drifts into an inappropriate tone; the *capped* version maintains therapeutic boundaries + compassionate guidance.
- **Implication**: a practical, lightweight means of **stabilizing persona** without retraining — adjacent to Anthropic's work on character training (cf. [anthropic-measuring-political-bias-claude-2025-11-13](../2025-11/anthropic-measuring-political-bias-claude-2025-11-13.md)).

## RésuméDe400mots

The 350th issue of **The Batch**, DeepLearning.AI's weekly newsletter published on April 24, 2026, opens with an editorial by **Andrew Ng** laying out an **acceleration hierarchy for coding agents** by type of software work. Ng sets out an explicit ranking: **frontend** benefits from maximum acceleration (agents are *"fluent in popular frontend languages like TypeScript and JavaScript"* and can iterate in an autonomous browser loop); **backend** sees **moderate** acceleration (corner cases, security, and DB migrations require experienced human oversight); **infrastructure** benefits **little** from agents (LLMs have *"relatively limited"* knowledge of network and system tradeoffs); and **research** remains largely human on the conceptual work of hypothesis formation, interpretation, and iteration. Ng draws a management takeaway from this: calibrate expectations and team organization according to these differentials.

The issue then covers four structuring news items. **Z.ai's GLM-5.1** is a MoE model with 754B parameters (40B active), MIT-licensed, capable of autonomously sustaining a single task for up to **8 hours** via a plan-execute-evaluate loop. It takes the lead on **SWE-Bench Pro at 58.4%** (vs. 54-57% for competitors) and tops CyberGym (68.7), while trailing on reasoning (GPQA Diamond 86.2% vs. 94.3% for Gemini 3.1). Z.ai has meanwhile raised its API pricing by roughly 40%.

**Agility Robotics** is deploying its **Digit** humanoids on **Schaeffler**'s production lines in South Carolina — the first operational industrial deployment. Operating cost is put at $10-25/h against roughly $20/h for an entry-level human position. McKinsey projects 5 million humanoids in factories by 2040 (vs. ~200 in 2026).

The **anti-data-center revolt** is gaining momentum: roughly $64B in projects blocked or delayed between May 2024 and March 2025, a moratorium in Maine for facilities ≥20MW, the first popular referendum in Wisconsin, and ousted council members in Missouri. Two violent incidents stood out: a molotov cocktail thrown at Sam Altman's home in San Francisco, and gunfire at an Indianapolis council member's residence. Grievances center on the power grid, energy rates, water consumption, and nuisance.

Finally, researchers (Christina Lu, MATS, Oxford, Anthropic) introduce the **"assistant axis"** — a vector of adherence to the trained persona that enables **activation capping**. Results: harmful responses in Qwen3 32B drop from 83% to 41%, in Llama 3.3 70B from 65% to 33%, without degrading IFEval/GSM8k/MMLU-Pro/EQ-Bench.

## GrapheDeConnaissance

- Andrew Ng —affirme_que→ les coding agents accélèrent le frontend plus que le backend, l'infra et la recherche (AFFIRMATION, 0.98)
- Andrew Ng —dirige→ DeepLearning.AI (ORGANISATION, 0.97)
- DeepLearning.AI —publie→ The Batch (DOCUMENT, 0.99)
- Coding agents —améliore→ frontend development (CONCEPT, 0.95)
- Coding agents —utilise→ TypeScript et JavaScript (TECHNOLOGIE, 0.96)
- Andrew Ng —affirme_que→ l'infrastructure est peu accélérée par les LLMs actuels (AFFIRMATION, 0.92)
- Andrew Ng —affirme_que→ le travail conceptuel de recherche reste majoritairement humain (AFFIRMATION, 0.9)
- Z.ai —publie→ GLM-5.1 (TECHNOLOGIE, 0.98)
- GLM-5.1 —est_variante_de→ GLM (TECHNOLOGIE, 0.97)
- GLM-5.1 —utilise→ architecture MoE 754B / 40B actifs (CONCEPT, 0.97)
- GLM-5.1 —permet→ tâches autonomes jusqu'à 8 heures (CONCEPT, 0.94)
- GLM-5.1 —mesure→ 58,4% sur SWE-Bench Pro (leader) (MESURE, 0.95)
- Agility Robotics —a_créé→ Digit (TECHNOLOGIE, 0.98)
- Schaeffler —utilise→ Digit (TECHNOLOGIE, 0.97)
- Agility Robotics —mesure→ coût opérationnel Digit 10-25$/h vs ~20$/h humain entry-level (MESURE, 0.9)
- McKinsey —prédit→ 5 millions d'humanoïdes en usine en 2040 (vs ~200 en 2026) (MESURE, 0.85)
- Mouvement anti-data-center —s_oppose_à→ projets data-centers US (~64 Md$ bloqués mai 2024 – mars 2025) (CONCEPT, 0.93)
- Maine —publie→ moratoire data centers ≥20MW jusqu'en 2027 (DOCUMENT, 0.9)
- The Batch —référence→ cocktail molotov au domicile de Sam Altman (SF) (EVENEMENT, 0.92)
- Christina Lu —a_créé→ Assistant axis (CONCEPT, 0.95)
- Christina Lu —a_créé→ Activation capping (METHODOLOGIE, 0.95)
- Activation capping —réduit→ jailbreaks Qwen3 32B de 83% à 41% (CONCEPT, 0.96)
- Activation capping —permet→ préservation des performances IFEval, GSM8k, MMLU-Pro, EQ-Bench (CONCEPT, 0.94)
- Hiérarchie d'accélération —converge_avec→ verifiability framework de Karpathy (CONCEPT, 0.85)

---
Canonical: https://www.thekb.eu/en/fiches/ng-the-batch-350-coding-agents-software-work-acceleration-2026-04-24/
