What are you looking for ?
VergeIO
RAIDON

Huawei Connect 2026: Huawei Introduces OceanStor M900 Context Memory Storage to Accelerate AI Inference in Hyperscale Data Centers

Confirming its ambition to cover all AI infrastructure dimensions

Summary:

  • Huawei has introduced OceanStor M900 Context Memory Storage, designed for AI inference in hyperscale data centers. It delivers PB-scale KV cache for SuperPoDs to break through the bottlenecks of on-chip memory and DRAM capacity and unlock the full potential of SuperPoDs
  • OceanStor M900 Context Memory Storage accelerates inference and reduces costs with three key technologies. With the UnifiedBus network, it deploys a KV cache pool for global sharing and enables tiered storage for KV cache, breaking through capacity limits. OceanStor M900 integrates the CPU, network controller unit, and NAND controller unit to boost inference performance with one-hop connections between NPUs and SSDs. With KV-aware adaptive storage technology, it intelligently manages the data lifecycle to extend SSD endurance and lower token costs

At Huawei Connect 2026, David Wang, deputy chairman of the board and rotating chairman, Huawei, officially introduced OceanStor M900 Context Memory Storage during his keynote titled “Advancing the Agentic World, Building a Solid Silicon Foundation”. Designed for AI inference in hyperscale data centers, the product provides SuperPoDs with a fully shared memory space that offers PB-scale capacity and TB/s-level performance. This marks a shift in AI infrastructure from a compute-centric model to deep collaboration among compute, network, and storage. This will help unleash the computing power of SuperPoDs.

2026 has seen the accelerated transition of AI from technological breakthroughs to large-scale implementation. AI applications have evolved from chatbots to agents capable of autonomously completing complex tasks. These agents are widely adopted in critical sectors such as scientific research, healthcare, finance, and public services, marking the beginning of the agentic AI era.

As large models grow to 10 trillion-scale parameters, SuperPoDs are becoming the optimal choice for AI infrastructure. Mainstream large models already support context windows exceeding one million tokens, multi-turn inference and complex tasks have become the norm, and KV cache data generated during inference continues to grow. These trends have pushed on-chip memory and DRAM beyond their limits in capacity and cost-effectiveness. It has become an industry consensus to build a multi-tier storage system that coordinates on-chip memory, DRAM, and SSDs to create a fully shared memory space with massive capacity.

Huawei introduced OceanStor M900 Context Memory Storage to overcome the memory capacity bottlenecks in ultra-long context and multi-turn inference. OceanStor M900 uses the UnifiedBus network to build PB-scale, global multi-tier KV cache with one-hop connections. This fully unleashes the computing power potential of SuperPoDs and accelerates AI inference in hyperscale data centers. OceanStor M900 Context Memory Storage has three key capabilities:

  • Breaking Capacity Boundaries to Empower Large-Scale AI with Massive Memory
    Powered by the high-speed UnifiedBus interconnect network, the KV cache achieves global pooling and sharing with tiered storage. The KV cache of SuperPoDs is expanded from on-chip memory and DRAM to SSDs, enabling a single cluster to deliver 64PB of capacity. The available KV cache capacity per NPU is upgraded from gigabytes to terabytes, allowing more context to be stored, shared, and reused. This significantly boosts the KV cache hit ratio
  • Boosting Inference Performance to Fully Unleash Computing Power
    OceanStor M900 is the industry’s first architecture to integrate the CPU, network controller unit, and NAND controller unit. It provides native KV semantics to enable one-hop connection from the SuperPoD’s NPU to SSDs. This eliminates the need for protocol conversion and CPU forwarding, slashing access latency from milliseconds to 60 microseconds, a 90% reduction. A single cluster delivers 40TB/s of aggregate access bandwidth, 1.5 times higher than peer solutions. In typical AI programming scenarios, this architecture doubles the inference cluster’s token throughput and halves the time to first token (TTFT), converting computing power into productivity
  • Lowering Token Costs to Enable Economical Large-Scale AI Adoption
    OceanStor M900 uses the industry’s first KV-aware adaptive storage technology, which predicts KV cache lifecycles based on data value and intelligently distributes data across storage media. This technology enables up to 24 drive writes per day (DWPD), extending SSD endurance by 16 times and ensuring stability for three years. By reducing media replacement and O&M costs, it lowers the long-term costs of large-scale AI inference infrastructure and enables faster AI adoption

As AI expands into major production systems in all manner of industries, AI infrastructure is evolving from a compute-centric model toward tighter compute-network-storage collaboration. Context memory storage will be essential for continually enhancing the capacity and access efficiency of hyperscale inference KV caches. Huawei will continue to drive hardware-software synergy and system-level innovation, providing open, efficient, and sustainable AI infrastructure to support AI innovation and intelligent transformation in all industries.

Read also :

Comments

The keynote was particularly dense, with several major announcements, including the OceanStor M900 and other key product developments. The M900 represents Huawei's next iteration of its KV cache architecture and is expected to become available in early 2027. Its predecessor, the M800, is already available and is based on the OceanDisk 1800. Both have been displayed at Huawei data storage space.

Huawei's ambition with the M900 is significant: deliver a shared KV cache memory space at petabyte scale for SuperPoDs, with aggregate performance reaching the tens of TB/s. The industry continues to face the well-known "memory wall." Trillion-parameter models, million-token context windows and multi-turn AI agents are pushing on-chip memory and DRAM beyond what they can accommodate at a reasonable cost. The M900 addresses this challenge through a tiered KV cache architecture spanning on-chip memory, DRAM and SSDs.

Across capacity, architecture, performance and endurance, Huawei is claiming some impressive characteristics for the M900:

  • Capacity: Using the UnifiedBus interconnect, KV cache capacity can be pooled and globally shared, reaching up to 64PB per cluster, while KV capacity available to each NPU increases from GB to TB scale
  • Architecture: Huawei describes the M900 as the first design integrating CPU, network controller and NAND controller into a single unit. Native KV semantics provide NPUs with a single-hop path to SSDs
  • Latency and bandwidth: Access latency is reduced from milliseconds to 60µs, while aggregate cluster bandwidth reaches 40TB/s, which Huawei claims is 1.5× higher than competing solutions
  • Inference performance: For typical AI coding workloads, Huawei says token throughput doubles while time to first token is reduced by half
  • Endurance: KV-aware adaptive placement predicts the expected lifetime of cached data, supports up to 24 DWPD and, according to Huawei, can extend SSD lifetime by 16x

Competition is intensifying among a small group of leading players, including DDN, Huawei, Vast Data and Weka, around KV cache architectures. Nvidia is also shaping this market with its G3.5 storage tier, based on the BlueField-4 storage processor, CMX and its associated reference architecture.

Huawei's UnifiedBus is central to its strategy and provides an alternative to Nvidia's Spectrum-X and NVLink technologies. There are important architectural differences between these approaches, but perhaps the most significant is Huawei's vertically integrated model. Huawei controls the compute, networking, interconnect and storage stack, while the other vendors listed here primarily deliver software-defined solutions aligned with Nvidia's infrastructure.

Vast Data runs AIOS natively on BlueField-4 DPUs to provide pod-scale shared KV cache. In our view, this is probably the architecture that comes closest to the M900's single-hop, embedded-controller approach.

Weka has an early-mover advantage with its Augmented Memory Grid. The company claims capacity up to 1,000× that of DRAM while enabling GPUs to access NVMe storage with microsecond-level latency through GPUDirect Storage and RDMA.

DDN provides KV cache acceleration across both Infinia and EXAScaler. The company claims up to 27× faster KV cache loading with sub-millisecond latency.

The next few months should be particularly interesting as these different architectures mature and vendors refine their positioning. SC26, scheduled for November in Chicago, IL, should provide an important stage for the next round of announcements in this rapidly developing segment.

Articles_bottom
AIC