d-Matrix Joins Nvidia’s AI Inference Push

d-Matrix Joins Nvidia’s AI Inference Push

More

Summary

AI inference startup d-Matrix has announced a strategic partnership with NVIDIA to integrate its Raptor XPU into NVIDIA’s platform via NVLink Fusion, targeting ultra-low latency AI inference workloads. CEO Sid Sheth joins Bloomberg Technology to explain the rationale: NVIDIA’s infrastructure spans every major cloud and data center globally, giving d-Matrix immediate route-to-market without having to rebuild distribution from scratch.

d-Matrix’s core technical differentiator is a memory-centric architecture using 3D stacked DRAM packaged directly with compute — replacing HBM entirely — which Sheth claims punches through the memory wall that bottlenecks AI inference on latency, energy efficiency, and cost. The Raptor XPU will be the world’s first 3D stacked DRAM XPU, and Sheth asserts no competitor will be within a two-year window of matching this capability. The partnership has been in development for six months, with the Raptor XPU targeted to reach market in approximately twelve months.

Demand has accelerated sharply in the past year alongside the rise of agentic coding, where interactive low-latency compute is a hard requirement. Sheth identifies a broad and notable customer pipeline — hyperscalers, frontier AI labs, sovereign AI projects, inference clouds, and high-frequency traders — with formal customer announcements described as imminent. The company has been developing its memory-centric approach for over seven years.


📺 Source: Bloomberg Tech · Published September 10, 2026
🏷️ Format: Interview

1 Item

Channels

1 Item

Companies