Ming-Chi Kuo: Nvidia Restarts Rubin CPX Project for Inference Prefill Optimization Chip
Renowned analyst Ming-Chi Kuo's survey shows that Nvidia is reviving its AI inference prefill chip "Rubin CPX," with mass production expected in early 2027. The new version uses 168GB HBM4 memory, has a power consumption of 2300 watts, and computing power close to the standard Rubin GPU. This chip adopts an independent modular rack and hierarchical interconnect design, and will be deployed together with Rubin GPU on a 1:1 basis, specifically targeting prefill and KV Cache generation to reduce the cost of long-context inference.
Nvidia has quietly restarted a chip project once believed by the market to have been abandoned, targeting the cost-heavy prefill stage of AI inference.
According to the latest industry research from renowned analyst Ming-Chi Kuo, Nvidia has relaunched the AI inference prefill acceleration GPU project “Rubin CPX”, and has made significant adjustments to the product design. Under current plans, the new Rubin CPX is expected to enter production in Q1 2027.
The project restart sends a clear signal: Nvidia is doubling down on solving the performance and cost bottlenecks of the prefill stage in AI inference. As the context length of large models continues to increase, the prefill stage requires processing a massive amount of input information and generating KV Cache. The efficiency in this phase directly affects the overall cost of AI inference and deployment economics.
Kuo points out that currently, more than 50% of AI inference workloads originate from context input processing and KV Cache construction.
Significant Changes to Chip Specifications, Compute Power Approaches Standard Rubin
The new Rubin CPX has seen significant changes in its core specifications compared to previous designs.
In terms of compute power and power consumption, the new CPX’s performance is close to that of the standard Rubin GPU, with a maximum card power consumption also reaching 2300 watts. Regarding memory, the new CPX uses 168GB of HBM4, a clear upgrade from the earlier 128GB of GDDR7, although still below the standard Rubin GPU’s 288GB of HBM4.
According to Kuo, an 8-card CPX compute tray provides a total HBM4 capacity of about 1.34TB, sufficient to cover most long-context prefill workloads and their corresponding KV Cache needs. This means that Rubin CPX is not simply aiming for higher general-purpose compute power, but has been redesigned around VRAM capacity and prefill efficiency for long-context scenarios.
Moving from Shared Racks to Independent Deployment, More Flexible Configurations
The rack architecture is also a key focus in these adjustments.
In the previous design, CPX was planned to share the rack with Rubin GPU; in the new version, it adopts an independent MGX ETL rack, allowing customers to independently expand CPX compute power as needed.
Specifically, customers can choose deployment scales of 64, 128, 192, or 256 CPX GPUs. Every 64 CPX GPUs form a rack module, comprising 8 compute trays (each with 8 CPX GPUs) plus a switch tray.
The modular design means customers are not required to procure a fixed large-scale Rubin rack, but can flexibly add CPX computing power according to the actual demands of the prefill workload.
Tiered Interconnect Design Reduces Prefill Deployment Costs
In terms of interconnect architecture, Rubin CPX adopts a hierarchical approach, balancing performance and cost.
Within a single compute tray, 8 CPX GPUs use NVLink for scale-up expansion. Each CPX GPU has an NVLink bandwidth of 1 to 1.5TB/s, lower than the standard Rubin GPU’s 3.6TB/s.
At the scale-out layer between trays and within rack modules, CPX utilizes Spectrum-6 all-copper L1 Ethernet links; for cross-rack module connections, each module’s Spectrum-6 switch and OSFP optical fiber links enable interconnection.
Compared to solutions that rely solely on high-bandwidth GPU interconnects, this tiered architecture emphasizes cost optimization tailored for prefill scenarios.
Cooperating with Rubin GPU, Targeting Long-Context Inference
Rubin CPX is not a standalone general-purpose GPU, but serves as a prefill accelerator used in tandem with Vera Rubin NVL72.
According to Kuo, Nvidia recommends a 1:1 deployment ratio of CPX to Rubin GPU: CPX handles the prefill computation and produces the KV Cache, which is then transmitted via Ethernet RDMA to the Rubin GPU for subsequent decoding.
From a product positioning perspective, Rubin CPX’s core purpose is not to replace the standard Rubin GPU, but to further split the AI inference workflow, assigning prefill and decoding tasks to dedicated GPUs.
Kuo characterizes it as “the best value solution for long-context prefill.” As AI model context windows continue expanding, resulting in rising computation and KV Cache demands during prefill, Nvidia’s restart of the CPX project appears to be an effort to further reduce the cost of long-context inference through specialized hardware and heterogeneous deployment strategies.
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like

XRP Delivers 13% Average Gains in September: Will History Repeat?

Solana’s 7% Pullback Isn’t Slowing Demand: Here’s the $150 Setup
United States Treasury yields read Iran strikes as inflation
