CoreWeave Inc reported the deployment of a multi-rack Nvidia Vera Rubin NVL72 cluster on CoreWeave Cloud, a configuration the company says aggregates hundreds of Rubin GPUs into a single scale-out cluster tailored for agentic AI workloads. The announcement coincided with a 2.9% premarket increase in CoreWeave shares.
The multi-rack Vera Rubin NVL72 architecture is designed to run both training and inference jobs across the aggregated Rubin GPUs. According to the company, a single NVL72 rack pairs 72 Rubin GPUs with 36 Vera CPUs and integrates Nvidia NVLink 6, Nvidia ConnectX-9 SuperNICs, and Nvidia BlueField-4 DPUs into the rack design.
CoreWeave said it is the first AI cloud provider to validate and bring up the Vera Rubin NVL72. To create the single scale-out cluster, the company unifies racks of accelerators using Nvidia Spectrum-X Ethernet networking, enabling workloads to span hundreds of Rubin GPUs in one logical cluster.
Alongside the hardware deployment, CoreWeave introduced two enhancements to its AI Object Storage. The new cross-region write acceleration capability allows data to be written at local latency while the system migrates a copy to a secondary, remote region in the background. The Archive tier provides a lower-cost storage option and, per the announcement, does not levy retrieval, early deletion, or reading fees.
CoreWeave described the AI Object Storage LOTA as delivering reads at local NVMe speeds and reducing latency by 8x versus reading from a traditional storage cluster. The company stated the system can provide up to 7 GB/s of throughput per GPU.
The combined hardware and storage updates are positioned to support high-throughput, low-latency AI workloads that span large numbers of accelerators. The company framed the work as both a capacity and capability expansion for CoreWeave Cloud.
Key points
- CoreWeave deployed a multi-rack Nvidia Vera Rubin NVL72 cluster that aggregates hundreds of Rubin GPUs into a single scale-out cluster - impacting cloud infrastructure and AI compute capacity.
- The company added cross-region write acceleration and an Archive tier to its AI Object Storage - changes relevant to data-center storage and enterprise AI data workflows.
- CoreWeave reported AI Object Storage LOTA delivers reads at local NVMe speeds, claims an 8x latency reduction versus traditional clusters, and up to 7 GB/s throughput per GPU - metrics aimed at performance-sensitive AI workloads.
Risks and uncertainties
- Performance and throughput figures are provided by the company; independent validation of those claims is not included in the announcement - this affects cloud and AI compute procurement decisions.
- Operational complexity of running and managing multi-rack, scale-out Rubin clusters at hundreds of GPUs introduces execution and integration risk for cloud infrastructure and enterprise deployments.
- Details on cost trade-offs for the Archive tier and the practical behavior of cross-region write migration under production loads were not specified in the announcement - implications for storage cost models and data-center operations remain to be seen.