The question AI providers hope VPs of Engineering never ask
April 20, 2026 – 10:00 am
AI coding adoption is exploding. But most engineering leaders are still measuring usage instead of outcomes. That creates a costly blind spot. There is a question that nobody in the AI industry wants you to ask:
How much of the code your AI agents generate actually reaches production?
Not how much was generated, not how many prompts were run, not how many seats are active. But how much survived code review, passed CI, got merged, deployed, and reached a customer. Most engineering leaders cannot answer this. And AI providers have no incentive to help them find out.
The spend is real; the visibility is not. According to the Stanford AI Spend Index, the median company now spends $86 per developer per month on AI coding tools (that’s across 140 companies and over 113,000 developers). The top quartile spends more than $195. Some companies spend over $28,000 per developer per month.
Anthropic just crossed $30 billion in annualized revenue. Up from $9 billion four months ago. According to SemiAnalysis, 4% of all public GitHub commits are now authored by Claude Code. That is projected to exceed 20% by year-end. Linear’s CEO declared issue tracking dead in March.
Coding agents are installed in more than 75% of Linear’s enterprise workspaces. The money is flowing, and the code is flowing, but nobody is tracking how much of that code actually ships.
The incentive problem nobody talks about
AI providers bill by tokens; the more tokens your engineers consume, the more revenue the provider earns. They get paid when a token is consumed, not when the code it generated passes review, gets merged, or deploys. This creates a structural misalignment:
- A developer who prompts an AI agent ten times to generate a function that gets rewritten by a human reviewer costs you ten times more than one who gets it right on the first prompt.
- The provider earns ten times more from the first developer; the second developer is worth ten times more to your organization.
Right now, most engineering leaders cannot tell the difference. They see a single line item on the AI bill and have no idea which tokens produced production code and which produced waste.
This is not a conspiracy; it’s a structural incentive problem. And it’s the VP of Engineering’s problem to solve because the provider has no reason to solve it for them.
We’ve seen this before, in the early days of cloud computing. Companies moved to AWS and Azure aggressively, spending heavily on the promise of efficiency. The reality was waste; it took years for FinOps discipline to emerge, as companies realized they were overspending by 30 to 40% on cloud infrastructure because nobody was measuring what was actually being used.
AI spend is following the exact same pattern, except the growth rate is faster.