# gupta-token-budget-wars-marginal-token-utility-2026-05-28

## Veille

Viral X thread (**230.5K views**, May 28, 2026, 1:51 AM) by **Jaya Gupta** (@JayaGup10, investor — likely Foundation Capital, author of the *Context Graphs* framework) titled ***"Token Budget Wars"***. **Pivot thesis**: ***"Enterprise AI has moved from adoption to allocation"*** — phase 1 of enterprise AI proved that models can work; phase 2 will decide **how much of that work is worth it**. The new currency at the top of the enterprise is the **ability to quantify AI ROI**: *"show me the value"*. Canonical concept: ***marginal token utility*** = *"the business value created by each additional dollar of inference"* — the number that matters at scale, and that **most companies cannot see**. Timeline: **Claude shipped November 2025**, after the 2026 annual budgets were locked → as early as **Q1**, companies *"running multiples ahead of plan"* → inference stops being an experimentation line item and becomes a **recurring operating cost**. Shift from **experimentation (a few $100K) → infrastructure (seven figures, $1M+)**: at infrastructure scale, **technical variance produces material P&L swings — two runs of the same workflow on the same input can differ by 5-10× in token cost** with nothing visibly broken, *"a number the CFO has to explain to the CEO"*. **AI competes with labor**: 3 types of budget requests (replace outsourced work / replace internal work / generate revenue) → shift toward the ***cost of a completed outcome*** (cost per resolved ticket, processed claim, reviewed contract, completed invoice, avoided hire, retained customer, dollar of revenue moved). **BPO = the easiest baseline to benchmark against** (already priced in completed units); internal work is much harder (multi-skilled employees, diffuse gains, HR resistance to headcount reduction). **Why it's different from SaaS**: SaaS learned to treat usage as a proxy for value; AI breaks that proxy — *"the signal and the noise share the same unit"* (the token), *"SaaS usage told you the software had been adopted. AI usage tells you the meter is running. It doesn't tell you whether your company is cooking."* **Three causes of marginal token utility's invisibility**: (1) ***retry tails*** — tokens per resolved workflow ≈ **T/p**; going from 90% to 70% completion increases effective cost by ~**28%**, not 20%, because failures compound; (2) ***context inflation*** — inference cost ≈ **O(n²)** in context length (attention), doubling the context **quadruples** reasoning cost (over-retrieval: 50 docs when 5 would do); (3) ***routing*** — by default the most powerful model is used (basic classification run on a complex reasoning model); across millions of calls, the difference between routing easy tasks to a small model and sending everything to the frontier model = *"the difference between a manageable bill and a board-level problem."* **Sector split**: **software** companies = a **productivity measurement** problem (already instrumented: PRs, commits, deploys, incidents, cycle time, MTTR — tracks *"AI layoffs"*); **non-software** companies = a **transformation** problem (operational work: claims, underwriting, support, compliance reviews, supply chain exceptions, payment disputes — *right under audit, not just right on average*). **The missing layer = token-to-outcome attribution**: a conversion layer linking inference spend → work performed → business outcome, answering 3 questions (real cost including retries/corrections; which parts of the trace mattered vs. thrashing; did the work change the operating model). ***Measurement becomes memory***: linking a token to an outcome requires capturing **decision traces** (what the agent saw, retrieved, called, ignored, where it retried, when a human overrode it) — *"decision rationale is one of the most perishable assets in a company"* (lives in Slack, emails, escalation calls, people's heads). Agents **create** these traces; captured first to justify the spend, they become *"more valuable than the cost report"* → a **context graph** (*"although I am so tired of that word these days"*). **The allocation layer is the prize**: whoever owns token-to-outcome attribution makes the **allocation calls** (which workflows deserve more compute, which are capped, which move to cheaper models, which stay human, which replace BPO). Companies won't do this on their own — they'll **buy it as a transformation** (Fortune 500 playbook: McKinsey + Palantir alumni + top-down CEO, in the manner of ERP/BI/digital transformation, a *"program"* with an executive sponsor and infrastructure that becomes the **new source of truth**). Framed by **Charlie Munger**: *"show me the incentive and I will show you the outcome."* Organizational sub-thesis: the decades-old executive instinct that *big teams = big jobs/scope/power* → once intelligence becomes the **scarce resource**, the new marker is *"how much of it you're orchestrating."* Direct relevance to the **Cost Optimization / agentic FinOps positioning**: empirically confirms the levers (model routing, prompt caching, context hygiene, sub-agents) and shifts the KPI toward **cost per completed outcome**. Strong convergence with Bain's *cross-system labor* (execution data moat, Cursor), Ng's *No AI jobpocalypse* (pricing anchored on the replaced employee's salary), DORA ROI (cost per feature), Mensch/Mistral (electron→token), Ensarguet (economics of computation), Foundation Capital's *Context Graphs* (decision traces, same author), Wescale's *Token Burning*, BFM/Girard (token = value fuel).

## Titre Article

Token Budget Wars

## Date

2026-05-28

## URL

https://x.com/JayaGup10/status/2059784669382791653

## Keywords

Token Budget Wars, marginal token utility, token-to-outcome attribution, adoption to allocation, allocation layer, cost per completed outcome, cost of a completed outcome, retry tails, context inflation, model routing, Haiku Sonnet Opus routing, prompt caching, context hygiene, AI ROI quantification, show me the value, inference recurring operating cost, experimentation to infrastructure, seven figures, technical variance 5-10x token cost, AI spend competes with labor, BPO benchmark, business process outsourcing, cost per resolved ticket processed claim reviewed contract, avoided hire retained customer, SaaS usage proxy value, signal and noise share the same unit, is your company cooking, T over p retry economics, completion rate 90 to 70 percent 28 percent, O(n²) attention context length, over-supplied retrieval, frontier model overpowered, software productivity measurement problem, non-software transformation problem, AI layoffs, PRs commits deploys incidents cycle time MTTR, claims underwriting compliance reviews supply chain exceptions payment disputes, right under audit, decision traces, measurement becomes memory, decision rationale perishable asset, systems of record what not why, context graph, allocation calls, capped compute, transformation program, Fortune 500 playbook, McKinsey Palantir top-down CEO, ERP BI digital transformation, new source of truth, intelligence scarce resource, orchestrating intelligence, big teams big jobs, Charlie Munger show me the incentive, Jaya Gupta, Foundation Capital, tokenmaxxing, metamates, CFO CEO board-level problem, agentic FinOps, TCO cost per feature

## Authors

**Jaya Gupta** (@JayaGup10) — investisseuse / VC. Très probablement **Foundation Capital** (le thread s'auto-réfère au cadre ***Context Graphs*** — *« ahem, context graph, although I am so tired of that word these days »* — concept porté par Foundation Capital, cf. fiche `bain-100b-saas-opportunity` qui cite *Foundation Capital — Context Graphs trillion-dollar opportunity, 2025-12-22*). Thread publié sur X le **28 mai 2026 à 1h51**, **230,5K vues**, format essai long en un seul post. Une réponse notable de **@tuning_engines** (*« DevSecFinOps for the Agentic Era »*) : *« Tokens will basically have to be managed like headcount […] model hierarchies too »*.

## Ton

**Profile**: An investor's essay-thesis aimed at founders, CFOs/CEOs, CIOs, and enterprise AI operators. A discreet first-person narrative perspective, **VC-analyst** register, **medium-high** technical level (assumes *O(n²)*, retry tails, two-tower, but stays readable for a decision-maker). Long X-thread format (a single dense post, ~1,200 words), structured with internal subheadings (*Why this is different from SaaS*, *Why marginal token utility is hard to see*, *Measurement becomes memory*, *The allocation layer is the prize*).

**Style**: Sharp VC prose in the style of an investment memo — pivot statements laid down in one sentence then unpacked, a blend of **economic rigor** (formulas like *T/p*, *O(n²)*, numeric ranges 5-10×, 28%) and **Silicon Valley/Meta cultural register** (*tokenmaxxing*, *#metamates*, *#bearish* Jim Cramer/CNBC, *TikTok breaks at lunch*, *is your company cooking*). Self-deprecation about jargon (*"context graph, although I am so tired of that word these days"*). Closes on a totem quote (Munger) — a VC-memo signature. **Prescriptive without being activist** tone: no product sold outright, but a market thesis ("the allocation layer is the prize") that implicitly sketches a category of startups worth funding.

**Key aphorisms**:
- ***"Enterprise AI has moved from adoption to allocation."*** (pivot thesis).
- ***"Marginal token utility: the business value created by each additional dollar of inference."*** (canonical concept).
- ***"The signal and the noise share the same unit."*** (why SaaS no longer applies).
- ***"SaaS usage told you the software had been adopted. AI usage tells you the meter is running. It doesn't tell you whether your company is cooking."***
- ***"Measurement becomes memory."*** (decision traces → context graph).
- ***"The allocation layer is the prize."*** (the layer that makes the allocation calls).
- ***"Show me the incentive and I will show you the outcome."*** (Charlie Munger, closing).

**Elaborated metaphors**:
- ***The meter is running*** — inference as a running taxi meter: usage no longer proves value, it proves the spend keeps running. (A direct echo of Girard/BFM's *taxi without fuel* metaphor.)
- ***Is your company cooking*** — slang for "is this actually producing anything": two identical token bills can cover two radically different operations (one converts, the other pays for *thrash*).
- ***Token budget wars*** as **internal allocation warfare** — the battle over token ownership collides with the decades-old executive instinct (team size = power); the new marker = *"how much intelligence you're orchestrating."*
- ***Tokens managed like headcount*** (picked up in a reply) — tokens become a resource to manage like FTEs, with model hierarchies.

**Epistemic stance**: **market analysis + economic framework** positing an emerging product category (*token-to-outcome attribution* / *allocation layer*); a mix of formal models (retry economics, attention scaling) and field observation (Q1 2026 multiples ahead of plan). Candid about the jargon she helped popularize (context graph).

**Authority**: built from (a) the **investor's vantage point** seeing multiple companies at scale, (b) **authorship of the Context Graphs framework** (self-reference), (c) **precision of the models** (T/p, O(n²)), (d) **timing** (May 28, 2026, right after the Q1 budget shift), (e) **virality** (230.5K views) validating the topic's resonance with operators.

## Pense-betes

- **Date / source**: **May 28, 2026** (1:51 AM), X thread @JayaGup10, **230.5K views**. Long-essay format in a single post.
- **Author**: **Jaya Gupta**, investor (likely **Foundation Capital** — author of the *Context Graphs* framework, self-cited).
- **Pivot thesis**: ***"Enterprise AI has moved from adoption to allocation"*** — phase 1: models can work; phase 2: **how much of that work is worth it**. ### The core concept — marginal token utility > ***"the business value created by each additional dollar of inference. It's the number that matters at scale, and the number most companies cannot see."***
- It is the **derivative** of ROI: not total cost, but the **value of the marginal dollar of inference**.
- Invisible because **token utility is not quantified**: the bill doesn't say whether the spend replaced work, generated revenue, reduced risk, sped up a workflow… or just funded engineers *tokenmaxxing* on the leaderboard. ### Timeline of the shift | Moment | Fact | |--------|------| | **Nov. 2025** | Claude shipped **after** the 2026 annual budgets were locked | | **Q1 2026** | Companies *"running multiples ahead of plan"* | | **~a few $100K** threshold | Still experimentation | | **Seven figures ($1M+)** threshold | Becomes **infrastructure** → material P&L swings |
- **Canonical technical variance**: *"two runs of the same workflow on the same input can differ in token cost by 5-10x without anything visibly going wrong"* → at infrastructure scale, *"a number the CFO has to explain to the CEO."* ### AI competes with labor (not with SaaS)
- **3 types of budget requests**: replace **outsourced** work / replace **internal** work / generate **revenue**.
- Shift in unit: from the token to the ***cost of a completed outcome*** → cost per **resolved ticket, processed claim, reviewed contract, completed invoice, avoided hire, retained customer, dollar of revenue moved**.
- **BPO = the easiest baseline** (already priced in completed units: price per ticket/claim/invoice/review). **Internal** work = much harder (multi-skilled employees, diffuse gains = avoided hiring/capacity, HR resistance).
- ⚠️ Pitfall: *"a claim that requires three retries, human correction, and a frontier model may be more expensive than the outsourced labor it was supposed to replace."* ### Why SaaS no longer applies > ***"The signal and the noise share the same unit."*** The token (billing unit) is **stable**, but the work it represents **is not**. > ***"SaaS usage told you the software had been adopted. AI usage tells you the meter is running. It doesn't tell you whether your company is cooking."*** ### The 3 causes of invisibility — directly actionable (FinOps) | # | Cause | Mechanism | Lever | |---|-------|-----------|--------| | 1 | **Retry tails** | tokens/resolved workflow ≈ **T/p**; 90%→70% completion = **+~28%** cost (not 20%, failures compound) | improve first-pass completion reliability | | 2 | **Context inflation** | cost ≈ **O(n²)** in context length; doubling the context **×4**s reasoning cost; over-retrieval (50 docs instead of 5, entire email threads, stale history) | **context hygiene**, targeted retrieval | | 3 | **Routing** | default = strongest model; basic classification run on a complex reasoning model | **model routing** (small model for easy tasks) = *"manageable bill vs board-level problem"* | → **Maps exactly** onto the levers from the "Cost Optimization" slot of the Claude Code morning session (Haiku/Sonnet/Opus routing, prompt caching, context hygiene, sub-agents).
- ### Sector split | | **Software** | **Non-software** | |--|-------------|------------------| | Nature of the problem | **Productivity measurement** | **Transformation** | | Why | Work already instrumented (PRs, commits, deploys, incidents, cycle time, MTTR) | Operational work (claims, underwriting, support, compliance, supply chain, payment disputes) | | Requirement | *right on average* | ***right under audit*** | | Symptom | tracks *"AI layoffs"* | unit of work ≠ unit of cost ≠ same organization | ### The missing layer — token-to-outcome attribution
- **Conversion layer** linking: inference spend → work performed → business outcome.
- **3 questions**: (1) real cost **including retries/corrections**? (2) which parts of the trace **mattered** vs. **thrashing**? (3) did the work **change the operating model** (fewer tickets/agent, shorter claims cycles, reduced BPO line, delayed hiring)?
- Attribution **in the language of the business**: not *"this workflow cost $2.13"* but *"this class of claims is cheaper with agents than BPO, except when the policy requires exception documents, in which case the retry tail destroys the economics."* ### Measurement becomes memory > ***"Decision rationale is one of the most perishable assets in a company"*** — lives in Slack threads, email chains, escalation calls, and people's heads (who leave).
- Systems of record capture **what** happened, rarely **why** (a CRM says a deal slipped, not the unwritten judgment behind the forecast).
- **Agents create traces** (retrieval, tool call, retry, escalation, human correction, final decision).
- Captured first to **justify the spend**, they become *"more valuable than the cost report"* → a **context graph** (jargon the author says she's *"so tired"* of). ### The allocation layer is the prize
- Whoever owns token-to-outcome attribution makes the **allocation calls**: which workflows → get more compute / are capped / move to cheaper models / stay human / replace BPO.
- *"And once you make those calls, you control where AI spend goes inside the enterprise and get to have the trust to allocate."*
- **Bought as a transformation** (Fortune 500 playbook): McKinsey + Palantir alumni + top-down CEO; arrives like ERP/BI/digital transformation, a *"program"* with an executive sponsor + infrastructure = **new source of truth**.
- The founders capable of doing this will be *"different people than the classic archetype."* ### To leverage for
- **"Cost Optimization" slot (Claude Code morning session)**: canonical citation for the shift from *token cost → cost per completed outcome*; the 3 causes (retry/context/routing) structure the levers section; *"the meter is running"* and *"is your company cooking"* = punchlines for decision-makers.
- **Agentic FinOps consulting offering**: the *token-to-outcome attribution* layer is an **emerging product/consulting category** — possible positioning for SFEIR (instrument the trace, link it to the P&L).
- **CFO/CEO narrative**: *"a number the CFO has to explain to the CEO"* — frames exactly the 2026 AI budget conversation.
- **Decision traces / context graph argument**: convergence with Foundation Capital (same author), Bain (execution data moat), Talisman (ontology/governance) — *measurement becomes memory* = the 2026 data moat thesis. ### Connections to the watch corpus
- **Bain — cross-system labor** (2026-05): the same pairing of **execution data = moat** + **cost of completed outcome**; Bain sizes the market, Gupta sizes the **measurement problem** that unlocks it. Both cite AI as **labor cost substitution**.
- **Ng — No AI jobpocalypse** (2026-05-08): Ng describes **pricing power** (vendors anchor pricing on the **replaced employee's salary**); Gupta describes the **buyer-side counterpart** (AI is benchmarked against BPO/salary). Two faces of the same *AI spend competes with labor* mechanic.
- **DORA ROI** (2026-04-21): *"we don't measure AI by code it writes but by bottlenecks it clears"* + *"code is a liability"* → Gupta = the **token-level** version of the same rejection of activity proxies.
- **Mensch / Mistral** (2026-05-13): *"we're turning electricity into intelligence, into token generation"* — an electron→token economy on the supply side; Gupta = a token→outcome economy on the demand side.
- **Ensarguet — Economics of Computation** (2026-03-11): the *kilowatt-hour moment*, the end of the billable-hour brain; Gupta extends this on the **unit value of compute** side.
- **Foundation Capital — Context Graphs** (2025-12-22, **same author**): *measurement becomes memory* = an explicit bridge to the Context Graphs framework; decision traces = the new system of record.
- **Wescale — Augmented Software Factory** (2026-05-03): the *Token Burning* concept + Agent Manager = the French operational counterpart to Gupta's *thrash*.
- **BFM / Girard** (2026-05-05): *"token = value fuel,"* NVIDIA bonuses paid in tokens, the taxi metaphor — direct convergence with *the meter is running*.
- **@tuning_engines** (reply — *"DevSecFinOps for the Agentic Era"*): a **governance/organizational** extension of the thesis. Three ideas: (1) ***"Tokens will basically have to be managed like headcount"*** — the token becomes a resource managed like a headcount (budget, allocation, justification); (2) ***model hierarchies*** — *"which model reports to which user (meaning which user can use which model in essence!)"* = **role-based access control (RBAC) on models**: who is allowed to use Opus vs. Sonnet vs. Haiku; (3) *"many organizational FTE management techniques will need application to the tokens as well"* — importing HR **workforce management** techniques (capacity planning, allocation, review) into token management. → A direct bridge to the morning's **Governance** slot (per-user and per-model quotas/permissions) and convergence with **Uber Engineering** (agent identity, *which agent can do what*).

## RésuméDe400mots

On May 28, 2026, **Jaya Gupta** (investor, likely Foundation Capital) published a viral essay-thread on X (230.5K views): ***"Token Budget Wars"***. **Pivot thesis**: *"Enterprise AI has moved from adoption to allocation."* Phase 1 proved that models can work; phase 2 will decide **how much of that work is worth it**. The new currency at the top of enterprises is **AI ROI quantification** — *"show me the value."*

Canonical concept: ***marginal token utility*** = *"the business value created by each additional dollar of inference"* — the number that matters at scale, **invisible** to most companies because the bill doesn't say whether the spend replaced work, generated revenue, or funded *tokenmaxxing*. Timeline: **Claude shipped November 2025**, after 2026 budgets were locked; as early as **Q1**, companies *"multiples ahead of plan."* Shift from **experimentation ($100K) → infrastructure ($1M+)**: *"two runs of the same workflow on the same input can differ in token cost by 5-10x"* — *"a number the CFO has to explain to the CEO."*

**AI competes with labor**: the unit shifts from the token to the ***cost of a completed outcome*** (per resolved ticket, processed claim, reviewed contract, avoided hire…). **BPO** is the easiest baseline (already priced in completed units). **Why SaaS no longer applies**: *"the signal and the noise share the same unit"*; *"SaaS usage told you the software had been adopted. AI usage tells you the meter is running. It doesn't tell you whether your company is cooking."*

**Three causes of invisibility**: (1) **retry tails** — tokens/resolution ≈ T/p, 90%→70% = +~28%; (2) **context inflation** — cost ≈ O(n²), doubling the context ×4s reasoning; (3) **routing** — sending everything to the frontier model = *"board-level problem."* **Split**: software = a **productivity measurement** problem; non-software = a **transformation** problem (*right under audit*).

**Missing layer**: ***token-to-outcome attribution*** linking inference → work → outcome. ***Measurement becomes memory***: agents create **decision traces** (*"decision rationale is one of the most perishable assets"*) that become *"more valuable than the cost report"* → a **context graph**. **The allocation layer is the prize**: whoever owns it makes the *allocation calls* and controls where AI spend goes — bought as a **transformation** (McKinsey + Palantir + top-down CEO, in the manner of ERP/BI). Closing with Munger: *"show me the incentive and I will show you the outcome."*

## GrapheDeConnaissance

- Jaya Gupta —publie→ Token Budget Wars (DOCUMENT, 0.98)
- Jaya Gupta —travaille_chez→ Foundation Capital (ORGANISATION, 0.75)
- Jaya Gupta —affirme_que→ « Enterprise AI has moved from adoption to allocation » (CITATION, 0.96)
- Marginal token utility —est_instance_de→ valeur business créée par dollar marginal d'inférence (CONCEPT, 0.97)
- Jaya Gupta —affirme_que→ Claude a été shippé en novembre 2025, après le lock des budgets annuels 2026 (AFFIRMATION, 0.93)
- Inférence —est_instance_de→ coût opérationnel récurrent (CONCEPT, 0.95)
- Jaya Gupta —mesure→ variance de 5-10× en coût de tokens entre deux exécutions du même workflow sur le même input (MESURE, 0.94)
- AI spend —concurrence→ le travail (labor) (CONCEPT, 0.95)
- Cost of completed outcome —remplace→ coût du token comme unité pertinente (CONCEPT, 0.94)
- BPO —est_instance_de→ baseline benchmarkable (unités complétées) (CONCEPT, 0.93)
- Jaya Gupta —affirme_que→ « the signal and the noise share the same unit » (CITATION, 0.95)
- Jaya Gupta —mesure→ retry tails : coût par résolution ≈ T/p ; 90%→70% de complétion = +~28% de coût (MESURE, 0.92)
- Context inflation —est_basé_sur→ scaling O(n²) de l'attention en longueur de contexte (CONCEPT, 0.92)
- Model routing —permet→ facture gérable (vs board-level problem) (CONCEPT, 0.93)
- Problème de mesure de productivité —s_applique_à→ entreprises software (CONCEPT, 0.9)
- Problème de transformation —s_applique_à→ entreprises non-software (CONCEPT, 0.9)
- Token-to-outcome attribution —est_instance_de→ la couche manquante (CONCEPT, 0.96)
- Agents —permet→ decision traces (CONCEPT, 0.95)
- Jaya Gupta —affirme_que→ les decision traces deviennent un context graph plus précieux que le cost report (AFFIRMATION, 0.93)
- Jaya Gupta —affirme_que→ « the allocation layer is the prize » (CITATION, 0.94)
- Jaya Gupta —prédit→ le token-to-outcome attribution sera acheté comme une transformation (McKinsey + Palantir + top-down CEO) (AFFIRMATION, 0.91)
- Charlie Munger —affirme_que→ « show me the incentive and I will show you the outcome » (CITATION, 0.96)
- @tuning_engines —affirme_que→ « tokens will basically have to be managed like headcount » (CITATION, 0.95)
- Model hierarchies (gouvernance tokens) —permet→ contrôle d'accès par rôle sur les modèles (quel utilisateur peut utiliser quel modèle) (CONCEPT, 0.92)
- Techniques de gestion FTE —s_applique_à→ la gestion des tokens (CONCEPT, 0.91)
- Token Budget Wars —converge_avec→ Bain cross-system labor, Ng pricing power, DORA ROI, Foundation Capital Context Graphs (DOCUMENT, 0.9)

---
Canonical: https://www.thekb.eu/en/fiches/gupta-token-budget-wars-marginal-token-utility-2026-05-28/
