The AI Revolution: Optimizing for the Wrong Outcomes
The AI revolution is optimizing for the wrong things, with a focus on coding, logic, and autonomous agent tasks, while neglecting general writing. This is the warning from Maz Ahmadi, founder of Wizard Labs.
Regression in AI Writing Performance
Ahmad notes that despite improvements in other areas, large language models are getting worse at writing, with measurable regression seen in client-specific prose benchmarks across model upgrades. While 44% of organizations are scaling AI enterprise-wide, only 37% report seeing an impact on earnings before interest and taxes (EBIT).
The Gap Between Deployment and Value
The widening gap between AI deployment and realized value raises questions about the criteria used to select models. Ahmadi advocates for building custom evaluation suites specific to a company’s needs before development begins, instead of relying solely on public benchmarks.
AI in Everyday Work
AI is now integrated into the daily work of millions, with nearly nine in ten companies using it in functions like report drafting, customer service, legal document preparation, coding, research analysis, and decision-making. Optimizing models for a narrow set of criteria can have widespread negative consequences.
AI Text Watermarking Debate
The debate over AI text watermarking further highlights the issue. Anthropic plans to watermark generated text in Claude models to comply with the EU AI Act, claiming it has no effect on quality or readability. However, enterprise users should consider the potential impacts of optimizing models around multiple, sometimes conflicting, criteria.
The Paradox for Enterprises
Enterprises are beginning to encounter a paradox: as AI spreads faster, the value it delivers grows slower. Selecting the "best" model based on public benchmarks and then expecting it to perform well in specific workflows is a common mistake. Custom evaluation suites are essential to ensure models meet actual business needs.