Aiming at the low-latency AI inference market, NVIDIA Corporation (NVDA.US) Groq 3 LPX has entered full mass production. The first batch of systems will be deployed at NEBIUS (NBIS.US).

date
00:07 25/08/2026
avatar
GMT Eight
NVIDIA announced on Monday that the Groq 3 LPX rack-mounted system has entered full-scale production, marking the formal commercialization of this low-latency AI inference technology after the company spent approximately $20 billion to acquire Groq-related assets last year.
NVIDIA Corporation (NVDA.US) announced on Monday that the Groq 3 LPX rack-level system has entered full-scale production, marking the commercial rollout of this low-latency AI inference technology following the company's acquisition of Groq-related assets for approximately $20 billion last year. The first systems will be deployed at AI cloud service provider NEBIUS (NBIS.US) and are expected to go live later this year. Dion Harris, NVIDIA Corporation's Senior Director, told the media that the Groq 3 LPX will be deployed alongside NVIDIA Corporation's Vera CPU and Rubin GPU in Nebius's data centers. The rapid production of Groq products reflects a shift in AI applications from model training to actual deployment, where low-latency AI inference is becoming a key market for NVIDIA Corporation. Particularly in applications such as AI agents and AI programming, models need to continuously generate content at faster speeds, reducing user wait times. Last December, NVIDIA Corporation spent approximately $20 billion to acquire the assets of AI chip startup Groq, making it the largest acquisition in the company's history. A notable feature of the Groq chip architecture is the integration of 500MB of high-speed SRAM within the chip, which reduces data transfer bottlenecks caused by traditional memory access and enhances response times during the AI model inference process. Unlike NVIDIA Corporation's main GPUs, which are produced by Taiwan Semiconductor Manufacturing Co., Ltd. Sponsored ADR (TSM.US), Groq chips are manufactured by Samsung. NVIDIA Corporation currently integrates 256 Groq 3 chips into a single LPX rack. According to benchmark data cited by NVIDIA Corporation from Artificial Analysis, the Groq 3 LPX system can achieve performance of approximately 3,400 tokens generated per second. In the context of generative AI, a token can be understood as the fundamental unit for a model to process and generate text. A higher token generation speed per second means that AI can respond more quickly to user requests, which is particularly important for latency-sensitive applications like AI programming and real-time agents. Harris stated that for cloud computing companies providing AI inference services, lower latency means they can offer higher-priced, premium services to customers with demanding response time requirements. However, Groq chips are not intended to replace NVIDIA Corporation's traditional GPUs. GPUs remain the core of current AI computing infrastructure, capable of both AI model training and inference tasks, while also offering greater versatility to adapt to different models and technical architectures. Low-latency chips like Groq focus more on specific aspects of the AI inference process, especially during the "decoding" stage when models generate content. Harris remarked, "This is not about replacing GPUs, but rather using the most suitable processors for different parts of the workload based on price and performance." This suggests that NVIDIA Corporation is attempting to establish a more segmented AI computing architecture: the Vera CPU and Rubin GPU will continue to handle broader AI computing tasks, while Groq chips will be optimized for latency-sensitive inference workloads. As the AI industry gradually transitions from large-scale model training to a phase of rapid growth in inference demand, competition in the low-latency inference market is intensifying. AMD (AMD.US) announced earlier this year that it would integrate its rack-level AI systems with Cerebras (CBRS.US) chips, also focusing on low-latency AI inference. The significance of this field is further increasing with the proliferation of applications like AI programming. Reports indicate that OpenAI's newly announced Ultrafast mode promises to reach 750 tokens per second and is supported by Cerebras' underlying computing power. In comparison, benchmark data cited by NVIDIA Corporation shows that the Groq 3 LPX can achieve 3,400 tokens per second. However, different systems may have varying testing conditions and application scenarios, so these figures should not be simply viewed as direct performance comparisons. Meanwhile, NVIDIA Corporation is accelerating the shipment of the Vera Rubin system, which entered production earlier this year. NVIDIA Corporation CEO Jensen Huang estimated in March during the launch of the Vera Rubin and Groq 3 LPX that cumulative sales could reach $1 trillion by 2027 from the current Blackwell chips to the new generation Vera Rubin systems. Huang also revealed that in a data center space dedicated to AI programming applications, he plans to allocate about one-quarter to Groq chips, with the remaining part entirely deploying Vera Rubin systems. This allocation ratio further underscores NVIDIA Corporation's emphasis on the low-latency inference market. As AI agents, AI programming, and real-time generative AI applications develop rapidly, inference speed is becoming an important metric for competition among cloud computing enterprises and AI developers. The formal entry of Groq 3 LPX into full-scale production also signifies that NVIDIA Corporation is expanding from being a GPU-centric AI chip supplier to providing dedicated computing architectures tailored for different AI workloads. The market will also be closely watching NVIDIA Corporation's latest performance. The company will release its earnings report this Wednesday, where the AI inference business and Groq's commercialization progress are expected to be new focal points for investors, in addition to the demand for Blackwell and Vera Rubin.