OpenAI's New Ultrafast Mode: Running GPT-5.6 Sol 14 Times Faster
OpenAI wants its cleverest model to also be its quickest. The company has previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol up to 14 times faster, reaching around 750 output tokens per second, on hardware built by Cerebras, a wafer-scale chipmaker.
Ultrafast is not a new model but rather a way to serve an existing one. It leverages Cerebras’s outsized chips to strip out the latency that has long hindered frontier AI, and OpenAI opened a limited preview on August 13th to a small group of customers with plans to expand access as capacity allows.
The key pitch revolves around a trade-off OpenAI says it can finally dissolve: Until now, genuine real-time responses required using smaller, less capable models, sacrificing intelligence for speed. Ultrafast aims to deliver frontier-grade reasoning and near-instant answers simultaneously, rather than forcing a choice between them.
This combination matters most for agentic software, the industry’s ultimate goal. An AI agent that takes thirty seconds to think before each step is merely a demo; one that answers in real-time begins to feel like a product. Speed is thus becoming as important a feature as raw cleverness.
OpenAI is targeting Ultrafast for time-sensitive applications including incident response, financial research and fraud detection, customer support, and e-commerce. Early testers like Jane Street, Podium, Basis, and Rogo describe the change as qualitative rather than incremental, with one noting that speed "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence."
For Cerebras, the deal is a significant endorsement at a critical moment. The company went public this year but has struggled to convince investors of its profit potential. Ultrafast provides much-needed validation. It also highlights that exotic chip architectures once dismissed as science projects are now handling real work for major AI players.
The speed boost comes from unusual engineering: Cerebras builds processors the size of a dinner plate, cut from single silicon wafers, allowing entire models to sit on one chip rather than being split across racks of Nvidia GPUs that constantly shuttle data. Stripping out this internal traffic significantly reduces the latency between prompt and reply—the basis for Cerebras’s argument that its design is better suited for inference than general-purpose chips.
This move aligns with a broader shift: As raw model capability levels off, the race shifts to who can run those models fastest and most cheaply. Leaders like Groq and others are already making latency a selling point, and the market increasingly rewards whoever can make a given model respond quickest, not just train it largest.