What are you looking for ?
VergeIO
RAIDON

Xinnor Demonstrates How GPU Node NVMe can Become Shared, High-Performance, Resilient Tier-0 Storage

Using xiRAID Opus to transform NVMe already installed inside GPU and compute nodes into a shared, protected storage resource

Modern GPU servers contain significant amounts of high-performance NVMe drives, but that capacity is typically tied to individual nodes and used as local scratch storage. Building a conventional shared Tier-0 environment usually means adding a separate high-performance storage tier alongside the compute cluster.The Xinnor‘s architecture demonstrated at the Weizmann Institute of Science takes a different approach. xiRAID Opus disaggregates local NVMe drives across the cluster and composes those resources into protected, software-defined storage pools that can be accessed from any participating node. This allows organizations to turn storage capacity already present in their compute infrastructure into shared storage, while adding RAID protection, multipathing, failover, and flexible resource placement in software.

The implementation used IBM Storage Scale, the established file-system environment at the Weizmann Institute of Science, above the xiRAID Opus block-storage layer to provide the shared file system and global namespace. This separation is central to the architecture: xiRAID Opus provides the distributed, protected storage foundation, while the file-system layer provides shared file access and data-management capabilities.

With xiRAID Opus providing the distributed block layer beneath IBM Storage Scale, the system reached 131.7GiB/s sequential read throughput, effectively matching the same performance measured at the raw NVMe-oF layer. Sequential write reached 46.9GiB/s, retaining approximately 98.5% of the raw block-level bandwidth. For random 4KiB IO impressive 5.8M IOPS were achieved for read and 3.1M IOPS for write.

The same architecture was then tested through the controlled loss of an entire node. With four of five nodes remaining, the architecture retained approximately 70–80% of healthy-cluster sequential and random-I/O performance while average random-I/O latency increased by only 0.6% and average sequential latency increased by 10-15%. Therefore, validated solution can continue serving data with predictable degradation rather than turning a node loss into a storage outage.

Finally, the same shared Tier-0 architecture to GPU data access path was validated. Using Nvidia GPUDirect Storage, aggregate reads from the shared file system into GPU memory reached 132.1 GiB/s across five nodes — effectively the same bandwidth as the maximum sequential read result of the shared storage system.

Together, these results demonstrate the broader value of the approach: the NVMe drives already embedded in GPU infrastructure can become a common storage foundation for shared HPC data, resilient checkpointing, and high-bandwidth GPU workloads, rather than remaining isolated local capacity or requiring a separate Tier-0 appliance.

“GPU and HPC systems already contain substantial high-performance storage resources. Working with Xinnor, we were able to explore how those resources could be used as part of a shared storage architecture rather than remaining tied to individual servers. The results demonstrate a compelling combination of shared access, performance, resilience, and a high-bandwidth path to GPU workloads,” said Dr. Mark Vilensky, scientific computing manager, Weizmann Institute of Science.

“The key idea is simple: if high-performance NVMe drives are already inside the GPU cluster, we should be able to use it as shared Tier-0 storage instead of building another storage island next to it. xiRAID Opus is what makes that possible. It separates storage from the physical node, protects it with software RAID across the nodes, and makes it available in the cluster while preserving the performance characteristics of NVMe drives. The results achieved at the Weizmann Institute of Science show that the same architecture can deliver shared file-system performance, predictable resilience, and direct GPU access — all from the storage resources already present in the compute infrastructure,” added Dmitry Livshits, CEO, Xinnor.

The results achieved provide a practical engineering reference for applying the same architecture to other HPC and AI environments. Deployments can be adapted to different cluster sizes, file systems, availability requirements, and workload profiles while preserving the same principle: use xiRAID Opus to turn distributed node-local NVMe drives into high-performance shared storage.

The newly published case study describes the architecture and key results, while a companion technical article provides detailed benchmark methodology, failure analysis, and implementation information.

Read the case study here and the technical deep dive here.

Read also :
Articles_bottom
AIC