AI-Assisted Software DevelopmentAug 14, 2026

OpenAI previews an Ultrafast mode running GPT-5.6 Sol on Cerebras at 750 tokens per second

On 13 August 2026 OpenAI previewed Ultrafast, a mode it says runs GPT-5.6 Sol at 14x standard processing speed and up to 750 output tokens per second, powered by its partnership with chipmaker Cerebras. Access is limited to a small group of enterprise customers and will widen as capacity grows. OpenAI named incident response, customer service and support, financial market analysis and e-commerce as the target deployments. No pricing was disclosed.

What it means At 750 tokens per second a multi-step agent loop stops feeling like a batch job, which moves the design question from 'can the agent do this' to 'can it do this while the user waits' — but the gate is Cerebras capacity, not the model.

Where it came from TechCrunch

Back to the Stream