HyperAIHyperAI

Command Palette

Search for a command to run...

NVLink Fusion Connects Custom XPUs to NVIDIA AI Factories

NVIDIA has introduced NVLink Fusion, a platform designed to enable hyperscalers and AI-native companies to integrate custom external processors into its established artificial intelligence infrastructure. By bridging third-party silicon with NVIDIA’s scale-up networking, rack-scale architecture, and factory-grade software stack, the initiative aims to accelerate time-to-market for semi-custom AI factories while mitigating the financial and logistical risks associated with building proprietary data centers from the ground up. Modern large language model training and inference demand high-throughput, low-latency interconnects. NVLink Fusion embeds custom accelerators into NVIDIA’s sixth-generation NVLink scale-up domain, supporting 72-processor configurations with end-to-end transfer latency three times lower than standard Ethernet solutions and a packet rate ten times higher. Future roadmaps extend this domain to 1,152 accelerators using co-packaged optics. Additionally, the platform incorporates NVLink chip-to-chip technology to link processors with NVIDIA Vera or third-party CPUs, delivering six times the energy efficiency of legacy PCIe interfaces and streamlining control-compute integration for agentic workloads. Beyond raw performance, NVLink Fusion addresses the operational complexity of AI factory deployment. The platform aligns with NVIDIA’s MGX rack-scale architecture, allowing custom silicon to share standardized power delivery, cooling, networking, and 800-volt direct current infrastructure with existing NVIDIA systems. This commonality enables automated manufacturing lines and multi-vendor supply chains, significantly reducing integration overhead. Industry executives note that the framework allows developers to pace custom silicon development independently while leveraging NVIDIA’s validated factory blueprints for immediate deployment. Infrastructure planning typically precedes silicon finalization, creating scheduling bottlenecks when facilities are locked to single-accelerator designs. NVLink Fusion resolves this by standardizing rack footprints and environmental systems across mixed-accelerator deployments. Hyperscalers can advance data center construction using NVIDIA’s DSX reference architecture and Omniverse digital twin simulations, then dynamically reassign capacity as workload demands and supply chains evolve. The architecture supports fully liquid-cooled, fanless compute and switch trays that maintain full rack operation during maintenance, ensuring high availability. Software integration further unifies the ecosystem. The platform is backed by NVIDIA Communications Library for distributed training, Dynamo and NIXL for compute disaggregation, and Mission Control for cluster telemetry and debugging. By combining custom accelerator innovation with proven infrastructure, supply chain logistics, and operational software, NVLink Fusion establishes a modular blueprint for gigawatt-scale AI factories. The approach enables organizations to offload infrastructure complexity while focusing engineering resources on targeted architectural breakthroughs, ultimately lowering cost-per-token metrics and improving utilization across diverse AI workloads.

Related Links