Nvidia Is Turning Its $20 Billion Groq Bet Into AI Infrastructure

Nvidia now says its Groq 3 LPX rack-scale system has entered full production, moving technology from its $20 billion acquisition of chip startup Groq toward commercial deployment as demand grows for faster artificial intelligence inference.

Nvidia Is Turning Its $20 Billion Groq Bet Into AI Infrastructure

The Groq systems will be deployed alongside Nvidia’s Vera central processing units and Rubin graphics processing units at cloud provider Nebius (NBIS.O), with the racks expected to be operational later this year, Nvidia senior director Dion Harris told reporters.

The push to bring Groq technology to customers highlights the growing importance of low-latency inference, which allows AI models to generate responses more quickly. That capability is increasingly important for AI agents, particularly applications such as coding in which users expect near-instant responses.

“For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive” service agreements, Harris said.

Nvidia acquired assets from Groq for $20 billion in December, making it the chipmaker’s largest acquisition. The deal gave Nvidia access to Groq’s technology for accelerating the inference stage of AI workloads.

Inference refers to the process of running a trained AI model to generate an answer or prediction. In many large language models, inference includes a “decode” phase in which the system generates output tokens one after another. The speed at which those tokens are produced can determine how responsive an AI application feels to a user.

Groq’s architecture is designed specifically to accelerate this process. Its chips include 500 megabytes of high-speed static random-access memory, or SRAM, on the chip itself. Keeping frequently used data close to the processing circuitry can reduce the time spent moving information between the processor and external memory, a common bottleneck in AI workloads.

Nvidia said it packages 256 Groq 3 chips into each LPX rack. The company said the system can produce 3,400 tokens per second, citing a benchmark from Artificial Analysis.

Groq chips are manufactured by Samsung Electronics (005930.KS), while Taiwan Semiconductor Manufacturing Co. (2330.TW) manufactures Nvidia’s GPUs.

The technology operates alongside, rather than replaces, GPUs, which remain the main processors used for training AI models and a broad range of inference workloads.

“This isn’t about replacing GPUs,” Harris said. “It’s about using the right price, right processor for the right part of the workload.”

Nvidia is betting that specialized processors can complement its GPUs as AI companies increasingly seek to reduce the cost and latency of serving models.

The market for specialized inference hardware is becoming more competitive. Advanced Micro Devices (AMD.O) said earlier this year it would integrate its rack-scale systems with chips from Cerebras Systems, which has also focused on low-latency AI inference.

OpenAI’s recently announced Ultrafast mode, meanwhile, promises speeds of up to 750 tokens per second and is “powered by Cerebras.”

Nvidia is also ramping up production of its Vera Rubin systems, which began production earlier this year. At the unveiling of the Vera Rubin and Groq 3 LPX systems in March, Chief Executive Jensen Huang projected $1 trillion in cumulative sales from Nvidia’s Blackwell and Vera Rubin platforms through 2027.

Huang said at the time that Nvidia would allocate about a quarter of data center capacity intended for coding applications to Groq processors.

“The rest of my data center is all 100% Vera Rubin,” Huang said.

Nvidia is scheduled to report quarterly results on Wednesday, as investors assess whether spending by cloud providers on AI infrastructure remains strong and how the company’s latest systems are being adopted.

Stay ahead of the Stories shaping our world. Subscribe to Impact Newswire and join our 
WhatsApp Channel for updates on global tech, business, and innovation—all in one place.

Dive deeper into the future with the Cause Effect 4.0 Podcast, where we explore the ideas, trends, and technologies driving the global AI conversation.

Got a story to share? Contact Us to reach a global audience with Impact Newswire.


Discover more from Impact AI News

Subscribe to get the latest posts sent to your email.

Scroll to Top

Discover more from Impact AI News

Subscribe now to keep reading and get access to the full archive.

Continue reading