Nvidia integrates Groq 3 LPX chips to accelerate AI responses in its infrastructure
Read more
Olhar Digital
olhardigital.com.br

Nvidia integrates Groq 3 LPX chips to accelerate AI responses in its infrastructure

Nvidia has started production of the Groq 3 LPX rack and plans to put it into operation this year at the Nebius neocloud. This technology was incorporated into the company after acquiring Groq assets for $20 billion, representing Nvidia's largest purchase to date.

The main objective of this implementation is to focus on low-latency inference, which allows reducing the time between a user sending a request and obtaining an answer from an artificial intelligence system. This functionality is particularly important in programming contexts, where any delay can be immediately noticed.

In the Groq 3 LPX racks installed at Nebius, the equipment will be positioned alongside Vera processors and Rubin GPUs. Dion Harris, Senior Director at Nvidia, confirmed that these devices will be available this year.

The architecture developed by Groq incorporates 500 MB of high-speed SRAM memory directly onto the chip. Thus, the technology aims to mitigate bottlenecks caused by memory access during the execution of AI models.

Each rack consists of 256 Groq 3 chips. According to data provided by Nvidia, the system achieves a rate of up to 3,400 tokens per second, a figure obtained through a performance test conducted by Artificial Analysis.

Specific function of Groq chips

Groq technology plays a more specialized role. Its chips are primarily directed towards the 'decode' phase of inference, which is responsible for generating the tokens that constitute the final response of a model.

GPUs, on the other hand, maintain their use in both training and inference, offering greater versatility to handle various models and technologies. For Harris, there is no need to eliminate one technology in favor of another; it is about using the appropriate processor with the correct cost for each part of the workload demand.

The proposed strategy is therefore to distribute tasks according to the inherent characteristics of each type of processor.

In the market, Nvidia competes with other companies seeking to accelerate AI response generation. AMD announced plans to integrate its rack-scale systems with Cerebras chips, a company that recently went public. OpenAI also participated in this competition, presenting its Ultrafast mode, which promises 750 tokens per second and is 'powered by Cerebras.'

Expansion of Vera Rubin systems

In parallel, Nvidia is increasing shipments of the Vera Rubin systems, whose production had previously begun. In March, during the launch event for the Vera Rubin and Groq 3 LPX systems, Jensen Huang declared that he would dedicate a quarter of the data center space for programming applications to Groq chips.

Huang emphasized at the time: 'The rest of my data center is 100% Vera Rubin.' With the introduction of Groq, Nvidia gains a dedicated option to speed up a specific stage of inference. GPUs remain the foundation of its systems, while the new racks take on responsibilities where response time is a determining factor.

Popular