R to @OpenAI: Powered by @Cerebras, Ultrafast generates up to 750 tokens per second, bringing our…
TL;DR - OpenAI previewed Ultrafast, an API service tier running GPT-5.6 Sol at up to 750 output tokens per second—up to 14× faster—on Cerebras infrastructure. It targets latency-sensitive enterprise applications where response speed directly affects utility.
- Supports real-time use cases such as voice, customer support, coding, design, and security response.
- Extends to commerce and financial research workflows requiring rapid model output.
- The announcement specifies generation speed, but provides no details on pricing, availability, or benchmark methodology.