Latest news NVIDIA

NVIDIA positions Groq as it once did Mellanox, with LPU set to become the latency weapon in the inference war

📖 Reading time: approx. 4 minutes · 710 words · 4,839 characters

When Jensen Huang compares an acquisition with Mellanox, it is not a casual aside, but rather a strategic assessment. This is precisely what the CEO of NVIDIA did in the Q4 2026 earnings call when he explained that Groq would be integrated “in very much the same way” as the company had expanded its own architecture with Mellanox at the time. Translated, this means that Groq will not be a peripheral product, but a structural building block.

The deal, a non-licensing agreement with a volume of up to $20 billion, is not a tactical acquisition. It is a response to a problem that NVIDIA has not completely solved despite Hopper and Blackwell: latency in inference decoding. Training dominates, no question. But in agentic multi-agent workloads, the bottleneck shifts. It is not the prefill, but the decode that determines response time and scalability. This is where Groq’s LPUs come into play. The architecture relies heavily on on-die SRAM and achieves internal bandwidths in the double-digit terabytes-per-second range. This is not a marketing value, but crucial for deterministic, extremely low-latency token output. While GPUs shine with wide HBM and enormous parallelism, LPUs aim for minimal overhead in the decode path. Two philosophies, one goal: saving time.

Huang hinted that they would “extend” their own architecture. This word is crucial. Mellanox solved the networking problem in the data center at the time, keyword InfiniBand and later NVLink ecosystem. The result was a co-design approach in which compute and interconnect were no longer thought of separately. Groq is now supposed to do the same for decode. There are two possible integration paths. First, rack-scale hybrids with dedicated LPU nodes. Analysts are circulating the theory of a possible “LPX rack” with 256 LPU units. The connection could be made via a plesiosynchronous chip-to-chip protocol, supplemented on the GPU side by NVLink Fusion to offload the KV cache during prefill. This would be a clear functional separation, with GPUs for attention and prefill, and LPUs for decode.

Second, the more radical option of integrating LPUs directly into future GPUs such as Feynman via hybrid bonding. Technically appealing, but significantly more complex in terms of packaging, yield, and thermal tuning. In the short term, rack scale seems more realistic. NVIDIA traditionally thinks systemically, not monolithically. The comparison with Mellanox is therefore more than symbolic. Mellanox was the lever that enabled NVIDIA to transform itself from a GPU provider to a data center architect. Groq could now be the lever that brings inference completely under its own control. Whoever controls decode controls agentic workloads. And whoever controls these controls the application layer, where, according to Huang, compute and revenue are now growing in lockstep.

One can view this soberly. Groq is not a replacement for GPUs, but an accelerator for a very specific problem. But that is precisely where the strategic sophistication lies. NVIDIA is not building a one-size-fits-all monster, but a modular ecosystem. Mellanox for networking, GPUs for training and prefill, LPUs for latency-critical decode. This is not expansion, it is architectural consolidation.

Whether an LPX rack or a closer GPU-LPU coupling will actually be presented at GTC 2026 remains to be seen. What is clear, however, is that NVIDIA is preparing the next step in the inference race. And this time, it’s not about more FLOPS, but fewer microseconds.

Source Key message Link
Wccftech Report on Jensen’s statements in the Q4 2026 earnings call, classification of Groq as an “accelerator” for low latency decode, and comparison of its strategic role with Mellanox, including the integration scenarios mentioned in the article. https://wccftech.com/nvidia-says-groq-acquisition-will-play-a-role-similar-to-mellanox/
The Motley Fool Transcript excerpt in which NVIDIA mentions an agreement with Groq for low latency inference, describes the architecture extension analogous to Mellanox, and refers to a more detailed presentation at GTC. https://www.fool.com/earnings/call-transcripts/2026/02/25/nvidia-nvda-q4-2026-earnings-call-transcript/
Groq Official announcement of a non-exclusive agreement between Groq and NVIDIA in the context of inference, including a description of the agreement framework from Groq’s perspective. https://groq.com/newsroom/groq-and-nvidia-enter-non-exclusive-inference-technology-licensing-agreement-to-accelerate-ai-inference-at-global-scale
Reuters Report on the structure of the Groq NVIDIA deal as licensing and personnel entry, including classification as an alternative to traditional acquisitions and reference to the rumored size of the deal. https://www.reuters.com/business/nvidia-buy-ai-chip-startup-groq-about-20-billion-cnbc-reports-2025-12-24/
Preisvergleich

Ähnliche Artikel zu NVIDIA positions Groq as it once did Mellanox, with LPU set to become the latency weapon in the inference war

Geizhals.de