d-Matrix Brings Its AI Inference Chips Into NVIDIA Racks
d-Matrix is integrating its Raptor inference accelerators with NVIDIA’s NVLink Fusion and MGX rack infrastructure, combining specialized 3D-stacked memory technology with NVIDIA’s data-center ecosystem.
On this page
d-Matrix is taking a less obvious route to challenge the dominant AI accelerator model: instead of building an entire data-center platform around its own chips, it is plugging its next-generation Raptor inference processors into NVIDIA’s rack-scale infrastructure. Announced on September 10, 2026, the partnership could make specialized inference hardware easier to deploy at the moment when AI workloads are shifting from training models toward serving them quickly to users.
d-Matrix is targeting the part of AI that happens after training
Training teaches an AI model its parameters, while inference is the process of using that trained model to generate an answer. Chatbots, coding assistants and voice agents spend their time doing inference, and those services increasingly care about how quickly the next token appears rather than simply how much raw computing power a chip can provide.
That distinction matters because generating tokens can become heavily constrained by moving model data between memory and compute. d-Matrix has designed its Raptor architecture around that problem. Instead of treating memory as a separate component sitting beside the processor, the company stacks custom dynamic random-access memory (DRAM) directly with the compute logic, shortening the path data has to travel.
Raptor will use NVIDIA infrastructure instead of replacing it
The September announcement puts Raptor inside NVIDIA’s MGX rack architecture through NVLink Fusion. NVIDIA introduced NVLink Fusion as a way for companies to build specialized processors that can connect to NVIDIA’s high-speed accelerator fabric and use its broader rack infrastructure rather than designing an entire deployment platform from scratch.
For d-Matrix, that means Raptor can be combined with NVIDIA components including Vera CPUs, NVLink switches, BlueField-4 data-processing units, ConnectX-9 networking and Spectrum-X Ethernet. The important change is not simply that another AI chip can communicate with NVIDIA hardware. It is that d-Matrix can use an established rack design, networking stack and supply ecosystem while concentrating its own engineering effort on inference silicon.
The interesting part is how Raptor handles memory
d-Matrix's Raptor design is built around a 3D memory architecture. The company has demonstrated working silicon in which a 4-nanometer compute die is bonded directly to a custom DRAM die, creating a much shorter connection between computation and memory than conventional accelerator designs that place high-bandwidth memory around a processor.
At Hot Chips 2026, d-Matrix reported more than 100 terabytes per second of memory bandwidth on a Raptor card with 32GB of 3D-stacked DRAM. Independent technical coverage has confirmed that the underlying measurements came from working silicon, but those figures should not be read as proof that Raptor is a universal replacement for high-bandwidth memory. Its capacity is much smaller than the memory available in some conventional AI accelerators, and its design is specifically aimed at inference workloads.
High bandwidth does not solve every AI bottleneck
The Raptor approach makes sense for workloads where moving data quickly is the limiting factor. During token generation, an accelerator repeatedly accesses model weights and other state needed to produce the next piece of an answer. Faster access can therefore improve responsiveness without requiring every stage of the workload to run on the same type of processor.
There is a catch. A model can outgrow the memory available on a single accelerator, forcing the system to distribute work across multiple chips. Once that happens, communication between processors becomes important again. Raptor's impressive on-package bandwidth cannot eliminate the latency and synchronization costs created by moving data between separate cards or racks.
NVIDIA is turning specialized chips into a rack-scale strategy
This is where NVLink Fusion becomes more significant than the individual Raptor chip. NVIDIA's platform is designed to let custom central processing units (CPUs) and specialized accelerators, often called XPUs, operate alongside NVIDIA GPUs in a common rack-scale system. NVIDIA says its sixth-generation NVLink can provide up to 3TB/s of all-to-all bandwidth per XPU in the NVLink Fusion architecture, while its platform documentation describes the broader goal as heterogeneous computing: using different processors for different parts of an AI workload.
That model fits the way modern inference is increasingly being divided. One processor can handle the initial processing of a prompt, while a specialized accelerator handles the repeated token-generation phase. d-Matrix says its Raptor systems are being designed for precisely these latency-sensitive services, including coding assistants, real-time chatbots and voice agents.
The partnership lowers a major barrier for AI chip startups
Designing a faster accelerator is only one part of building an AI infrastructure company. A chipmaker also needs servers, networking, power delivery, cooling, software, manufacturing partners and a practical way for customers to deploy thousands of processors. NVIDIA's strategy effectively packages much of that infrastructure around the custom silicon.
That can shorten the distance between a promising accelerator and a usable data-center product. It also explains why NVLink Fusion has attracted several custom-silicon and CPU partners. NVIDIA is not giving up its central role in the system; instead, it is attempting to make its infrastructure useful even when every processor in the rack does not carry an NVIDIA GPU.
Raptor is still a future product, not a shipping replacement
The timing is important. d-Matrix says Raptor is expected to tape out before the end of 2026, while initial availability of Raptor XPUs integrated into NVIDIA MGX racks is expected in the fourth quarter of 2027. That leaves considerable time for the design, manufacturing process, software stack and rack integration to prove themselves.
The early silicon results are therefore best understood as evidence that the underlying memory architecture can work, not as evidence that customers can already buy a finished Raptor system. Independent testing of production deployments will matter much more when the hardware reaches customers.
The bigger shift is toward mixed AI infrastructure
The d-Matrix deal points to a broader change in how AI hardware may be built. The future data center does not necessarily have to choose between NVIDIA GPUs and competing accelerators as an all-or-nothing decision. A rack can instead combine general-purpose AI processors with specialized silicon, assigning each stage of an AI workload to the hardware that handles it most efficiently.
For d-Matrix, that creates a realistic path to market for Raptor. For NVIDIA, it expands the reach of its networking and rack architecture without requiring every workload to run entirely on its own accelerators. The real test will come in 2027, when Raptor has to turn impressive memory numbers and a promising architecture into reliable, economical inference at production scale.
Written by


