Follow us
Breaking
Tech Services

Nvidia NVLink-C2C interconnect for high-performance scale-up

NVLink-C2C provides a high-speed die-to-die link that achieves 900 GB/s of bidirectional bandwidth in Grace Hopper Superchips. This technology offers significantly higher energy and area efficiency compared to traditional PCIe Gen 5 interconnects.

Share

The NVLink-C2C interconnect operates as a package-level, die-to-die link that tightly couples heterogeneous chiplets, which eliminates the propagation delays, connector parasitics, and re-serialization overhead that board-level links like PCIe unavoidably introduce. This proximity solves the bandwidth and latency limitations found in traditional PCIe 5.0 interconnects. The Grace Hopper Superchip uses this technology to connect a Grace CPU and a Hopper GPU. This connection achieves 900 GB/s of bidirectional bandwidth. The technology provides 25x more energy efficiency and 90x more area efficiency than PCIe Gen 5 on NVIDIA chips.

It works.

The hardware DMA engine on the GPU moves data between device memory and host memory. The engine slices transfers into burst transactions sized to align with cache lines, DRAM bursts, or link MTUs, and then it schedules read and write pipelines to keep both ends full. Each burst receives a traffic class or virtual channel tag to prevent head-of-line blocking for latency-critical traffic. The engine fetches descriptors and translates virtual to physical addresses using the GPU GMMU/TLB. If the system uses host or peer memory, the IOMMU, ATS, or PASID handles address validation.

Breaking the PCIe bottleneck

NVIDIA opened the ecosystem through NVLink Fusion. This allows third-party silicon to join NVLink networks. Hardware vendors license the NVLink-C2C IP and integrate it into their chips. This enables custom CPUs, like Fujitsu’s Arm-based Monaka or Qualcomm silicon, to connect to NVIDIA GPUs. Custom ASICs use a UCIe bridge chiplet to convert between the Universal Chiplet Interconnect Express standard and the NVLink protocol.

The ecosystem expands.

You know that PCIe usually limits how fast different chips talk. NVLink Fusion changes that by letting third-party silicon join the network. For example, AWS Trainium4 uses NVLink 6 via a UCIe bridge. This setup connects 72 Trainium4 ASICs at 3.6 TB/s per chip for a 260 TB/s total scale-up bandwidth. However, every NVLink Fusion deployment must include at least one NVIDIA product like a GPU, CPU, NVLink switch, or ConnectX NIC. To ensure broad device interoperability, NVLink-C2C supports industry-standard protocols such as Arm’s AMBA CHI and Compute Express Link.

Interconnect Bandwidth (Bidirectional) Efficiency vs PCIe Gen 5
NVLink-C2C (Vera) 1.8 TB/s 6x more energy efficient
NVLink-C2C (Grace Hopper) 900 GB/s 25x more energy efficient
PCIe Gen 5 128 GB/s Baseline

Marvell, Groq, and real-world limits

NVIDIA invested $2 billion in Marvell on March 31, 2026. This investment targets Marvell’s role as a dominant custom ASIC design house. Marvell designs the chips that hyperscalers use to build GPU alternatives. Marvell provides custom XPU design, silicon photonics, and optical DSP leadership. These capabilities influence the hardware level of next-generation custom ASICs. Marvell’s clients include AWS, Google, and Microsoft. This investment helps prevent the market from becoming a pure UALink play. Will Marvell eventually abandon NVLink for UALink?

Groq does not use NVLink Fusion. The Groq LPU connects to Vera Rubin via an Ethernet backplane called Oberon ETL256. Groq uses software. NVIDIA’s Dynamo software orchestrates the GPU-LPU workload split.

Real-world performance often falls short of theoretical maximums. Researchers measuring the Grace Hopper Superchip found NVLink-C2C bandwidth at 375 GB/s for host-to-device transfers and 297 GB/s for device-to-host. This totals 672 GB/s, which falls short of the 900 GB/s NVIDIA claim. Researchers also found that 64KB memory pages helped HPC simulations like Qiskit. These larger pages reduced memory management overhead. However, 4KB memory pages provided finer-grained management when access patterns stayed scattered. For some applications, the 4KB size ran up to 2 times faster than 64KB. The study found that system-allocated memory performs better on CPU-side initialization because page faults are both triggered and handled on the CPU-side.

Share

Technewsdaily

Senior tech writer covering AI, gadgets and cybersecurity. Breaking down the news that matters, every day.