Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
TL;DR - OpenAI previewed Ultrafast, an API service tier that runs GPT-5.6 Sol up to 14Ă— faster. Powered by Cerebras, it targets latency-sensitive applications with output speeds up to 750 tokens per second.
- Supports GPT-5.6 Sol through the OpenAI API.
- Delivers up to 750 output tokens per second.
- Advertises speed improvements of up to 14Ă—.
- Uses Cerebras infrastructure for accelerated inference.