Get the summary TL;DR: d-Matrix has entered production with its Corsair AI inference accelerators, partnering with NVIDIA to deliver high-performance, low-latency heterogeneous compute solutions that overcome the industry's critical memory wall. The Gist Who d-Matrix is an AI chiplet company founded by Sid Sheth (CEO) and Sudeep Bhoja (CTO) 00:00 1 . The company specializes in building high-performance, energy-efficient hardware designed specifically to accelerate AI inference and token generation 03:46 1 . Core Concept As AI models transition toward agentic workflows, low-latency token generation has become crucial 08:11 1 . Traditionally, AI performance has been bottlenecked by the "memory wall"—the physical distance and bandwidth limitations between processing logic and memory substrates 11:35 1 . To solve this, d-Matrix has developed a heterogeneous compute architecture that pairs their memory-on-chip accelerators with traditional GPUs 03:46 1 . By disaggregating workloads, the GPU handles heavy training-style "pre-fill" computations while d-Matrix’s hardware processes fast "decoding," resulting in ultra-low latency and highly efficient token generation 14:24 1 . Key Steps & How It Works Heterogeneous Deployment: In a typical rack configuration, NVIDIA GPUs (such as B300 servers) are paired with Supermicro servers containing d-Matrix Corsair accelerator cards 05:31 1 . Workload Disaggregation: The GPU executes the initial pre-fill stage of the model, while the Corsair cards handle the token decode phase 14:24 1 . Speculative Decoding: Corsair cards run smaller draft models to propose rapid token sequences, which are then quickly verified by the host GPU, turbo-boosting overall performance 13:27 1 . Seamless Integration: The solution fits into standard 19-inch enterprise racks, allowing data centers to upgrade their capabilities without reconstructing entire physical infrastructures 13:27 1 . Key Insights & Takeaways The "Fast Token" Economy: Real-time interactivity (demanded by agentic applications like coding assistants) has created a premium "fast token" market, with pricing currently reaching up to 10x higher than throughput tokens . Turbo Boosting Existing Fleets: Rather than purchasing expensive new-generation GPUs, enterprises can retrofit existing Hopper or Ampere GPU clusters with d-Matrix accelerators to achieve up to a 2x performance improvement 16:08 1 . Breaking the Memory Wall with SRAM: Corsair cards feature the market's densest air-cooled SRAM architecture, reducing the physical distance data must travel and achieving elite energy and cost efficiency without requiring liquid cooling 16:58 1 . The Future is 3D: d-Matrix is actively developing its next-generation architecture to vertically stack DRAM and logic on a single piece of silicon via hybrid bonding, eliminating the memory wall altogether 17:47 1 . Production and Partnership -> 01:55 1 Heterogeneous Compute Architecture -> 03:46 1 The Memory Wall Solution -> 11:35 1 3D Chiplet Roadmap -> 17:47 1 Advice I can take away from this source Based on the interview with the co-founders of d-Matrix, the conversation yields clear, actionable strategies and technological takeaways for enterprises, developers, and cloud providers navigating the rapidly evolving AI inference landscape. Here is the key actionable advice extracted from the discussion: 1. Optimize Inference with Heterogeneous Compute Instead of relying solely on expensive, power-hungry GPU-only setups for AI inference, deploy a heterogeneous compute model. Pair traditional GPUs with dedicated fast accelerators (such as d-Matrix's Corsair) to split the workload efficiently. Let GPUs handle the pre-fill: Leverage the massive compute flops of GPUs to handle the initial pre-fill stage of a workload 15:15 1 . Offload the decode stage: Use specialized memory-efficient hardware to handle the fast token generation (decode) and speculative decoding tasks, which dramatically reduces overall user latency 04:38 1 . 2. Retrofit Existing Hardware Fleets to Save on CapEx Before committing massive capital expenditures to purchase next-generation GPUs (such as upgrading entire fleets to Blackwell), look for opportunities to turbo-boost your existing infrastructure. You can integrate specialized accelerator cards into your current server setups containing older GPUs (like Ampere or Hopper generations) 16:08 1 . This hardware retrofitting can deliver up to a 2x improvement in performance at a fraction of the cost, bypassing the need for a full, disruptive hardware upgrade 16:08 1 . 3. Move Toward 3D Memory and Logic Integration If you are designing AI hardware or selecting future-proof infrastructure, prioritize architectures that solve the "memory wall" by physically bringing memory and compute closer together. Transition away from traditional layouts where memory and logic are separated by large physical distances on the board. Opt for technologies that stack DRAM and logic together in a 3D manner (hybrid bonding) 17:47 1 . This provides massive memory bandwidth at highly reduced energy levels because data travels a much shorter distance on the silicon 12:24 1 . 4. Capitalize on the Premium "Fast Token" Economy For application developers and service providers, there is an immediate opportunity to monetize low-latency, real-time AI experiences (such as agentic coding assistants). Introduce tiered pricing: Consider offering a "fast mode" at a premium price point (historically up to 10x higher than standard throughput tokens) for interactive, machine-to-machine, or highly time-sensitive tasks 06:29 1 . Optimize for TCO: As market competition increases and fast-token prices eventually drop, ensure your underlying infrastructure uses high-density, air-cooled solutions to keep your operating and cooling margins sustainable 07:22 1 . 00:14 1 welcome back around the cube's live stream here in Paris, France. I'm john furrier host of the cube. we're here for the wave summit day two as we're winding down two days of wall-to-wall coverage. the top story has been AI infrastructure and the big part is hardware has been a big part of it. Generate Presentation Presentation d Matrix at RAISE Summit 2026 Your presentation is ready. It covers d-Matrix’s production milestone, its Nvidia partnership, the fast token economy, performance gains, the memory wall problem, and the company’s 3D roadmap. NVIDIA & D-Matrix Partnership Optimizes Model Performance Play Sudeep Bhoja & Sid Sheth, d-Matrix | RAISE Summit 2026 07m 55s