What are you looking for ?
VergeIO
RAIDON

Benchmark Study from DeepInfra Confirms Nvidia Vera CPU Outperforms Leading CPUs on Agentic Workloads

Production-scale testing on DeepInfra's agent infrastructure showed the Nvidia Vera CPU hosting up to 1.6x more concurrent AI agents at the same quality of service

DeepInfra, a purpose-built cloud platform for high-throughput AI inference, announced the results of an independent benchmark of the Nvidia Vera CPU, Nvidia’s next generation CPU purpose-built to power agentic workloads.As one of a select group of collaborators granted early access to Vera hardware through Nvidia’s open AI ecosystem, DeepInfra designed and ran its own production-grade benchmark to test the infrastructure with a real-world agent workload. The benchmark found Vera led by up to 2.2x vs. competitive CPUs.

DeepInfra processes nearly five trillion tokens per week, with close to 30% driven by agentic systems. At that scale, CPU performance is critical to delivering the cost efficiency, low latency, and throughput that agentic workloads demand. 

“We built DeepInfra’s infrastructure from the ground up for inference. We have tuned every layer of it for cost, latency, and throughput at production scale,” said Nikola Borisov, co-founder and CEO, DeepInfra. “Nvidia is building for where AI workloads are headed, and we believe Vera is exactly the kind of hardware solution the next iteration demands.”

DeepInfra tested Vera using its own production AI agent and real captured traffic, benchmarking vs. three leading CPUs from AMD, and Intel under identical conditions. 

Key results from the benchmark include:

  • Fastest in every workload category, doubling Nvidia’s claim: Against Nvidia’s published claim of 80% faster agentic CPU performance, Vera measured up to 2.2x faster orchestration than the x86 baseline – and was the fastest of all four architectures tested in every workload category
  • Up to 1.6x more concurrent agents at the same quality of service: On identical CPU partitions, Vera sustained 256 concurrent agents within strict response-time and error targets, vs. 160-192 for the competing chips
  • Spare capacity served a full LLM: With all 256 agents running at full load, Vera’s leftover cores simultaneously served a 20-billion-parameter open-source model faster than an entire previous-gen CPU socket dedicated solely to that task, without slowing the agents

“Bringing Vera to market is an incredible opportunity, but performance is ultimately proven in production,” said Ian Finder, director, data center CPU products, Nvidia. “DeepInfra is pushing Vera with demanding, real-world agentic workloads that reflect the needs of production environments. These results demonstrate exactly what Vera was designed to deliver.”

Read also :
Articles_bottom
AIC