OpenAI is previewing Ultrafast, a new API tier powered by Cerebras that accelerates its most capable GPT-5.6 Sol model up to 14 times faster, generating 750 output tokens per second.
OpenAI has announced Ultrafast, an API tier designed to dramatically accelerate inference speeds for GPT-5.6 Sol, the company's most advanced model. The new tier leverages infrastructure from Cerebras, a specialist in AI chip design and deployment.
The performance gains are substantial. Ultrafast delivers up to 14× faster speeds compared to standard OpenAI API tiers, while generating up to 750 output tokens per second. This throughput represents a significant leap for real-time applications that demand rapid model responses.
The preview signals OpenAI's push to address latency concerns that have limited enterprise adoption of its most capable models. Real-time AI applications—from customer service to code generation—often require faster response times than standard inference provides.
Cerebras' involvement indicates OpenAI's broader infrastructure strategy. The partnership leverages Cerebras' specialized hardware and software stack optimized for transformer-based models, complementing OpenAI's existing computational resources.
The Ultrafast tier will likely carry premium pricing, positioning it as a specialized offering rather than a replacement for existing API tiers. OpenAI has historically tiered its services by capability and speed, allowing customers to optimize cost versus performance based on their needs.
This development reflects intensifying competition in the inference optimization space. Other AI providers and specialized hardware companies are similarly pursuing faster, more efficient model deployment. Faster inference directly impacts user experience and operational costs for AI applications at scale.
The preview phase suggests the offering remains under evaluation before broader availability. OpenAI typically gathers feedback and optimizes offerings during preview periods before general release.
For developers and enterprises relying on GPT-5.6 Sol for latency-sensitive applications, Ultrafast could enable new use cases previously constrained by inference speed.
A critical examination of an AI-generated film found that its most compelling moments came from human-created elements, highlighting current limitations in machine-generated entertainment.
A new analysis reveals significant variance in how different AI models respond to identical prompts, highlighting the importance of model selection for specific use cases.
Google released Gemini 3.7 Flash just three weeks after its predecessor, positioning the model as its strongest coding and AI agent tool. The company claims it outperforms Claude Sonnet 5 and GPT-5.6 Terra at half the price.
Anthropic researchers deployed multiple AI agents on identical tasks and observed them clash, collude, and coordinate in unexpected ways. The findings suggest current safety tests may not adequately capture risks posed by multi-agent AI systems.