September 10, 2026, (Inside AI) — D-Matrix will integrate its inference processors directly into Nvidia server racks using the chip giant's NVLink Fusion interconnect technology. The announcement marks a rare hardware-level collaboration between Nvidia and a startup competing in the inference market.
The Raptor chips from d-Matrix will plug into Nvidia's data-center systems through connectors and specialized memory built for custom AI silicon. Nvidia-compatible racks are expected to ship in 2027, while Raptor's final design stage wraps by the end of this year.
The deal answers the who, what, when, and where clearly. D-Matrix, based in Santa Clara, California, is the startup. Nvidia is the infrastructure partner. The what is a hardware integration for inference workloads. The when is a 2027 commercial timeline. The where is inside Nvidia's own server racks.
The why is equally direct. AI demand is shifting from training models to running them daily, a process called inference. Nvidia dominates training with expensive graphics processors. D-Matrix specializes in inference chips designed for lower cost and higher efficiency at scale.
D-Matrix said the combined systems target fast, low-latency services such as coding assistants, chatbots, and voice agents. Speed is critical in these applications, where users expect sub-second responses.
Financial terms were not disclosed. The startup also said it is partnering with connectivity firm Astera Labs to build custom solutions ensuring fast data flow across the system.
Microsoft has backed d-Matrix since its $110 million financing round in 2023. The startup shipped its first AI chip in November 2024 and was valued at $2 billion when it raised $450 million last year.
Inference Economics Reshape Silicon Choices
Training a large language model requires massive compute and memory bandwidth. Inference, by contrast, rewards chips that deliver tokens quickly while consuming less power per query. That distinction has opened a market for specialized inference silicon.
D-Matrix's bet is that inference will become the dominant cost center for AI operators. Industry analysts have noted that inference spending typically exceeds training spending once models reach production scale. A custom inference chip inside Nvidia racks could reduce total cost per query.
The NVLink Fusion approach is notable because it lets third-party silicon access Nvidia's high-bandwidth interconnect ecosystem. Historically, Nvidia's interconnect technologies were reserved for its own GPUs. Opening that layer to a startup signals a strategic shift.
Nvidia benefits by keeping custom chips inside its data-center footprint rather than losing those workloads to alternative systems. D-Matrix benefits from compatibility with Nvidia's installed base and software ecosystem.
Competing Chips Fight For Inference Share
D-Matrix is not alone in targeting inference. Groq, Cerebras, and SambaNova have all positioned chips for fast token generation. AMD and Intel also sell inference-oriented accelerators. The difference here is physical integration into Nvidia racks.
One official who spoke on condition of anonymity said the collaboration does not mean Nvidia is abandoning its own inference roadmap. Nvidia's Blackwell and future architectures remain central to its data-center strategy. The d-Matrix deal is additive, not a replacement.
Inside AI could not independently verify the performance claims for Raptor. D-Matrix has not published public benchmarks comparing Raptor against Nvidia's inference GPUs or competing custom silicon.
The startup's valuation trajectory suggests investor confidence. A $2 billion valuation after a $450 million round indicates significant capital backing for a company that shipped its first chip less than two years ago.
Astera Labs' role is also worth watching. The connectivity firm specializes in PCIe, CXL, and Ethernet solutions for AI infrastructure. Custom data-flow solutions could address bottlenecks that emerge when mixing heterogeneous silicon in a single rack.
The broader implication is that AI data centers are becoming more modular. Operators want flexibility to mix training and inference hardware without sacrificing performance. Interconnect standards like NVLink Fusion could enable that flexibility.
D-Matrix's previous funding history includes participation from Microsoft, which has its own inference needs across Azure and Copilot services. Microsoft's continued backing suggests potential deployment pathways beyond Nvidia's racks.
The 2027 availability window gives competitors time to respond. Chip design cycles are long, and the inference market is moving fast. Whether Raptor can deliver meaningful cost or latency advantages remains an open question.
What is clear is that the line between Nvidia's ecosystem and custom silicon is blurring. Startups once forced to build entirely separate systems can now plug into the dominant infrastructure. That could accelerate innovation in inference hardware.
The collaboration also reflects the growing importance of low-latency AI services. Coding assistants, voice agents, and chatbots require responses in milliseconds. Every layer of the stack, from silicon to interconnect, must be optimized for that constraint.
D-Matrix did not disclose pricing for Raptor or the combined systems. That omission leaves open questions about total cost of ownership compared to Nvidia's own inference GPUs.
For now, the announcement signals a pragmatic convergence. Nvidia extends its rack ecosystem. D-Matrix gains a path to market. AI operators get another option for inference workloads at scale.