Gimlet Labs Adds Cerebras to Deliver Ultrafast AI Inference Through Gimlet Cloud Deployment
Combining Cerebras wafer-scale compute with Gimlet's inference cloud to deliver up to 3,000 tokens per second
This is a Press Release edited by StorageNewsletter.com on October 6, 2026 at 2:01 pmGimlet Labs and Cerebras Systems announced a collaboration to deliver a new class of ultrafast AI inference at massive scale.
The collaboration brings together Cerebras’ wafer-scale compute with the Gimlet Cloud to deliver a purpose-built disaggregated inference cloud spanning datacenter infrastructure to developer APIs. Together, the companies plan to deliver speeds of up to 3,000 tokens per second for demanding agentic and real-time applications, with the first Cerebras-powered Gimlet Cloud datacenter expected to come online later this year.
In AI, speed drives user experience and engagement and shapes what users can build. For real-time applications, from voice and video AI to agents and assistants, latency can be the difference between an interaction that feels seamless and one that feels slow. Real-time AI feels like an active collaborator that is immediate, fluid and responsive. When AI responds in real time, users do more with it, stay longer and run higher value workloads. As a result, fast tokens are more valuable tokens.
Gimlet Cloud combines the Cerebras Wafer Scale Engine with GPUs into an integrated inference solution. It uses advanced inference disaggregation technology to orchestrate model execution so each phase of inference runs on the silicon best suited to it. Together, these capabilities deliver a new class of fast inference optimized for production-scale agentic and real-time workloads.
“Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets. By combining Gimlet’s multi-silicon software with the Cerebras Wafer Scale Engine, we can run each phase of inference on the hardware best suited to it and plan to deliver up to 3,000 tokens per second at production scale,” said Zain Asgar, co-founder and CEO, Gimlet Labs.
“Combining the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics for everyone,” said Sean Lie, co-founder and CTO, Cerebras. “Everyone wants more high value tokens. Cerebras delivers the fastest AI inference in the world, and GPUs deliver high throughput. By making Cerebras a native part of its inference cloud, Gimlet will bring our industry leading speed and intelligent AI to more developers at production scale. We’re excited to build with Gimlet as a launch partner for CS-4, giving customers a direct path to our latest technology.”
The collaboration builds on joint customer engagements underway since last year and an integrated solution already serving tokens in private deployments.
Gimlet Labs will expand the collaboration with Cerebras to integrate software, infrastructure design, APIs, developer tooling, optimization, validation and production operations to make ultrafast inference broadly available through Gimlet Cloud.













