OpenAI Previews 'Ultrafast' Mode, Running GPT-5.6 Sol at Up to 14x Standard Speed

OpenAI opened a limited preview on Thursday of Ultrafast, a new API service tier that runs GPT-5.6 Sol, its most capable model, at speeds the company says reach 14 times its standard processing rate. The tier peaks at 750 output tokens per second, a jump OpenAI is pitching as a shift in what real-time AI products can actually do rather than a routine performance patch.

What Ultrafast changes and what it doesn't

Ultrafast is not a new model. It runs the existing GPT-5.6 Sol, the flagship model OpenAI launched in early July, on different hardware under a different service tier. OpenAI and its hardware partner both describe the change as latency-only, meaning the model's underlying intelligence stays the same whether a request runs on Standard or Ultrafast.

"Until now, getting real-time speed typically meant choosing a smaller or more specialized model," OpenAI wrote in its announcement. "Ultrafast points to progress in a new direction: more useful work per second." That framing matters because it's the crux of OpenAI's pitch. Historically, a team that needed sub-second responses had to accept a lighter, less capable model to get there. OpenAI is arguing Ultrafast breaks that trade-off rather than just moving it.

The speed comes from Cerebras, the chipmaker whose wafer-scale processors avoid the memory-bandwidth bottlenecks that slow conventional GPU inference. Cerebras's chip carries 44 gigabytes of on-die SRAM, which lets it keep more of a model's working data close to the compute itself instead of shuttling it back and forth from external memory. That architecture is the whole reason the speed jump is possible without dropping to a smaller model.

The partnership behind it

This didn't happen overnight. OpenAI agreed in January to buy up to 750 megawatts of Cerebras compute over three years, a deal that quietly set up the infrastructure this preview now runs on. Ultrafast is the first publicly visible product of that agreement.

Cerebras put its own numbers behind the launch. The company says Ultrafast completed Humanity's Last Exam, a 2,500-question benchmark covering graduate-level chemistry, economics and literature, in just over 11 hours, reaching accuracy similar to Anthropic's Claude Fable 5, which needed more than three days of continuous compute to finish the same test. On GDP-Val, a benchmark built around paid knowledge work such as legal briefs and financial models, Cerebras claims a 5.6x end-to-end speedup with no drop in output quality.

Cerebras also ran head-to-head speed comparisons against competing models. According to the company, Ultrafast processes output 5x faster than Claude Opus 4.8 running in Fast mode, and 11x faster than Claude Fable 5. Anthropic does offer its own accelerated tier, documented as Fast Mode in its developer platform, but nothing in Anthropic's public materials claims speeds in the range OpenAI and Cerebras are reporting here.

Reading the numbers with some care

A caveat worth sitting with: every one of these multiples originates from OpenAI or Cerebras themselves. There's no independent baseline disclosed, no published prompt set, and no stated reasoning-effort configuration behind the "up to 14x" or "up to 750 tokens per second" figures. Vendor benchmarks aren't automatically wrong, but "up to" language typically describes a ceiling reached under favorable conditions rather than what a typical customer request will see.

It's also, for now, a tightly controlled preview. There's no public price attached, no general availability date, and access is limited to a small group of API customers OpenAI selected directly. The company says it will expand access as capacity grows, without committing to a timeline.

Where OpenAI expects it to matter first

OpenAI named a cluster of workflows it thinks benefit most from cutting response time to near-real-time: incident response, customer support, financial market analysis, and e-commerce, along with fraud detection mentioned by Cerebras.

OpenAI is also testing the tier on itself. One internal team is using Ultrafast for live incident response, feeding alerts and logs through the model fast enough to help engineers diagnose problems as they happen rather than after the fact. On the research side, the company says Ultrafast is tightening experimentation loops enough that a researcher can run several hypothesis tests inside a single workday, work that previously ran as overnight batch jobs.

That internal dogfooding is arguably a more honest signal than the benchmark charts. A company routing its own incident response through a new inference tier, rather than saving that job for a slower and better-tested system, tells you something about how much confidence sits behind the preview.

Cerebras CEO Andrew Feldman framed the stakes in broader terms, calling speed a prerequisite for wider AI adoption rather than a nice-to-have layered on top of intelligence. Whether that holds depends less on the demo numbers than on what happens once Ultrafast moves past a hand-picked group of early customers and into workloads OpenAI doesn't control the conditions for.

Comments

Join the discussion and share your perspective.