What are you looking for ?
VergeIO
RAIDON

CoreWeave Fully Connected 2026: CoreWeave Delivers Nvidia Vera Rubin NVL72 Performance at Production Scale, Starting with Cognition

Maintaining industry-leading AI cloud performance across multiple compute generations with its software platform

CoreWeave Inc., an essential cloud for AI, announced the availability of Nvidia Vera Rubin NVL72 on CoreWeave, with Cognition as the first customer anywhere running production workloads on the system. Customers, like Cognition, run the system under the same operating model and tooling as their existing Nvidia GB200 NVL72 and GB300 NVL72 fleets, with performance engineering from CoreWeave’s team.The news was shared during Fully Connected, CoreWeave’s AI cloud conference, which brings together more than 4,500 customers, partners, developers and AI leaders to share how they are building and running AI in production.

“Bringing up Nvidia Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU gens,” said Chen Goldberg, EVP, product & engineering, CoreWeave. “With customers like Cognition, that investment shows up in the ability to get production workloads running within days. When it comes to agentic tasks, long contexts, repeated model calls and thousands of concurrent tasks put pressure on the entire platform. Our job is to make compute, networking and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves.”

Cognition is the first customer in production with Vera Rubin NVL72
Cognition, the applied AI lab behind Devin, runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled from bridge capacity to thousands of GPUs for training and inference in less than nine months. Cognition worked with CoreWeave to stand up a Vera Rubin NVL72 cluster in early September, and Cognition’s own engineers ran the first customer-executed Vera Rubin inference benchmark, measured vs. a GB 200 NVL72 cluster baseline.

“Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning.” said Silas Alberti, SVP, research & founding team, Cognition. “By deploying the Nvidia Vera Rubin NVL72 on CoreWeave, our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done. CoreWeave continues to deliver the bleeding-edge rack-scale acceleration we need to push the boundaries of AI.”

Cognition’s engineers benchmarked Vera Rubin
In independent benchmarks run on CoreWeave Cloud, Cognition measured a 4.8 times increase in total token throughput for its SWE-2 inference workloads on Nvidia Vera Rubin NVL72 compared to a GB 200 NVL72 baseline. Additionally, the team recorded a 3.8 times boost in output token throughput for reinforcement learning workloads. For Cognition, that translates to more concurrent Devin sessions per GPU, drastically accelerated research loops and lower cost per session, with no loss in gen speed.

Proven across every Nvidia generation
The relationship between CoreWeave and Nvidia dates to 2017, beginning with the Nvidia Volta generation, which is still in commercial service today on CoreWeave Cloud, demonstrating the long useful life of Nvidia compute and the value CoreWeave’s platform can continue to draw from it. CoreWeave’s full-stack software platform – including CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and serverless inference – gives customers a consistent way to deploy and manage workloads across GPU gens. Customers can match each workload to appropriate capacity, keep using existing infrastructure as needs evolve and adopt new Nvidia architectures through a familiar operating environment.

“CoreWeave has consistently demonstrated the infrastructure expertise required to bring each new gen of Nvidia accelerated computing into production,” said Ian Buck, VP, hyperscale and HPC, Nvidia.  “Its full-stack expertise across multiple gens of Nvidia infrastructure is helping AI innovators like Cognition quickly put Nvidia Vera Rubin NVL72 to work on demanding production workloads.”

CoreWeave, in collaboration with Dell Technologies, was among the first cloud providers to deploy Dell PowerRack systems featuring NvidiaGB 200 and GB 300 NVL72, and is one of the first to deploy Nvidia Vera Rubin. CoreWeave has also published the industry’s first measured silicon performance numbers on the platform, showing 10 times the token throughput per megawatt over Nvidia GB 200 NVL72 on the DeepSeek R1 reasoning model at matched interactivity.

Read also :
Articles_bottom
AIC