AI

NVIDIA Opens Its AI Racks to d-Matrix Raptor Chips With NVLink Fusion

By Nino Ray Yeh · September 11, 2026 · 4 min read
NVIDIA NVLink Fusion technology for rack-scale AI infrastructure

The next phase of the AI chip race may be less about replacing NVIDIA and more about plugging directly into its infrastructure.

AI inference startup d-Matrix has announced a multi-year collaboration with NVIDIA that will bring its next-generation Raptor XPUs into NVIDIA’s MGX rack-scale architecture using NVLink Fusion. The first integrated systems are expected in the fourth quarter of 2027.

On the surface, this is a partnership between a much smaller chipmaker and the company that dominates AI computing. Look closer, though, and it points to something bigger: NVIDIA is increasingly positioning its racks, networking and interconnect technology as infrastructure that can accommodate specialised silicon alongside its own GPUs.

d-Matrix is joining the NVIDIA AI factory

Raptor is being designed for AI inference — the work that happens after a model has been trained, when services such as coding assistants, chatbots and voice agents have to generate answers for real users quickly.

According to d-Matrix, Raptor XPUs will connect into NVIDIA’s MGX rack design alongside Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking. Astera Labs is also working with d-Matrix on high-speed connectivity for the system.

NVIDIA describes NVLink Fusion as a way for custom CPUs and accelerators to use its high-bandwidth interconnect and rack-scale ecosystem without having to build an entirely separate infrastructure stack.

That matters because designing an AI chip is only part of the problem. Deploying thousands of accelerators also requires networking, power, cooling, validated rack designs and a supply chain capable of delivering the hardware at scale.

The AI race is moving toward inference

Reuters reports that d-Matrix’s Raptor design is expected to be completed by the end of 2026, with NVIDIA-compatible systems following in 2027. The company is backed by Microsoft and was valued at $2 billion after raising $450 million in 2025.

d-Matrix’s bet is that specialised inference hardware can take on workloads where latency and efficiency matter more than having a general-purpose GPU do everything. Its proposed architecture can split parts of an inference workload between NVIDIA GPUs and Raptor XPUs — for example, leaving compute-heavy prefill work to GPUs while Raptor handles latency-sensitive token generation.

That makes this story especially interesting alongside NVIDIA’s expanding relationship with specialised inference hardware such as Groq. The company that became indispensable to AI training is now building an ecosystem designed to remain central as inference becomes a much larger part of AI infrastructure spending.

NVIDIA does not need every accelerator to be an NVIDIA GPU

There is an apparent contradiction here. Why would NVIDIA make it easier for another accelerator company to enter its racks?

Because NVIDIA’s advantage increasingly extends beyond the GPU itself.

If custom chips still rely on NVLink, MGX, NVIDIA networking and the broader AI-factory ecosystem, NVIDIA can remain deeply embedded in the data centre even when some workloads run on third-party silicon. It turns the company from a supplier of accelerators into something closer to the architecture around the entire AI factory.

That strategy becomes more important as enormous amounts of new capacity are planned. Microsoft is reportedly targeting roughly 38GW of data-centre capacity by 2032, while NVIDIA and its Australian partners are targeting up to 2GW of AI infrastructure capacity by 2027. At that scale, operators have a powerful incentive to use different processors for different jobs rather than treating every AI workload the same way.

What this could mean for AI users

None of this means Raptor has already proved itself at NVIDIA-scale deployment. The first integrated racks are still more than a year away, and performance claims will ultimately need to be tested in real production environments.

But the direction is important. AI services are becoming increasingly sensitive to the cost, power consumption and latency of generating each token. If specialised inference processors can live inside the same infrastructure as NVIDIA GPUs, cloud providers gain another way to tune those economics without rebuilding their data centres around a completely different architecture.

For everyday users, the hardware will remain invisible. The potential result will not: faster AI agents, more responsive coding tools and lower-cost inference could all depend on what happens inside racks like these.

The bigger story is that NVIDIA appears willing to let more kinds of chips into its AI factory — as long as the factory itself continues to run on NVIDIA’s infrastructure.

Sources

d-Matrix announcement · NVIDIA · Reuters

Share this story

Topics

More From The Tech Boom

View all

Share with