GPT‑5.6 Sol Ultrafast Preview Claims Up to 750 Output Tokens per Second

OpenAI is previewing a Cerebras-powered API tier that runs GPT‑5.6 Sol up to 14 times faster than Standard processing for selected customers.

Frontier intelligence enters real-time workflows

On August 13, OpenAI announced a limited preview of Ultrafast, a new API service tier running GPT‑5.6 Sol on Cerebras infrastructure. The company reports up to 14× the speed of Standard processing and up to 750 output tokens per second. It highlights incident response, voice support, financial analysis, commerce and rapid research as early use cases where latency can change the workflow itself.

What the speed figure does—and does not—show

Both numbers are vendor-reported maxima, not guaranteed end-to-end latency for every prompt. Time to first token, input size, tool calls, network conditions, concurrency and output length still affect the user experience. Even if model quality is preserved, price and available capacity will determine whether the tier is economical.

Ultrafast is currently a limited API preview for selected customers; broad availability and public pricing have not been announced. Teams should compare total task completion time, success rate, cost and retry rate against Standard processing on identical workloads rather than treating tokens per second as the only metric.

Official source