NVIDIA’s So-Called “Hot Chips” Are Actually “Hot Platforms”

Sep 03, 2024

Leave a message

 

NVIDIA is focusing on system-level and data center-level engineering projects aimed at creating advanced systems and platforms capable of handling complex generative AI challenges.

 

Earlier this month, NVIDIA encountered rare bad news when reports surfaced that the company's highly anticipated "Blackwell" GPU accelerators might be delayed by as much as three months due to design flaws. However, a NVIDIA spokesperson stated that everything is proceeding as planned. Some suppliers indicated that nothing has changed, while others noted some normal delays.

 

Industry insiders expect that when NVIDIA reports its Q2 FY2025 financial results next Wednesday, users will gain more insights into the status of Blackwell.

 

It is reported that Blackwell chips-B100, B200, and GB200-will be a highlight of this year's Hot Chips conference, to be held next week at Stanford University in California. NVIDIA will introduce its architecture, detailing some new innovations, outlining the use of AI in chip design, and discussing liquid cooling research in data centers used to run these growing AI workloads. According to NVIDIA's Director of Accelerated Computing Products, Dave Salvator, the company will also showcase Blackwell chips already operating in one of its data centers.

 

Blackwell chips

▲ Blackwell chips

 

Much of what NVIDIA is discussing about Blackwell is already known, such as the Blackwell Ultra GPU being launched next year, and the next-generation Rubin GPU and Vera CPU starting to roll out in 2026. However, Salvator emphasized that when talking about Blackwell, it is crucial to view it as a platform rather than a single chip. Salvator made this point in a briefing for journalists and analysts this week as part of the preparations for Hot Chips.

 

"When you think about NVIDIA and the platforms we are building, the GPU, networking, and even our CPU are just the beginning," he said. "We are doing system-level and data center-level engineering to build these systems and platforms that can truly go out and tackle those really tough generative AI challenges. We've seen the scale of models grow over time, and most generative AI applications need to run in real-time, with the demands for inference increasing dramatically over the past few years. Real-time large language model inference requires multiple GPUs, and in the near future, it will require multiple server nodes."

 

ANNOUNCING NVIDIA BLACKWELLPLATFORM FOR TRILLION-PARAMETER SCALE GENERATIE AI

 

This includes not just Blackwell GPUs and Grace CPUs, but also NVLink Switch chips, Bluefield-3 DPUs, ConnextX-7 and ConnectX-8 NICs, Spectrum-4 Ethernet switches, and Quantum-3 InfiniBand switches. Salvator also provided different insights for NVLink Switch (below), compute, Spectrum-X800, and Quantum-X800.

 

NVIDIA introduced the much-anticipated Blackwell architecture at its GTC 2024 conference in March this year, with hyperscale vendors and OEMs quickly signing on. The company is targeting the rapidly expanding generative AI field, where large language models (LLMs) are becoming even more massive. Meta's Llama 3.1, launched in June, is a testament to this trend, featuring a model with 4.05 trillion parameters. Salvator noted that as LLMs grow larger, the demand for real-time inference persists, necessitating more computation and lower latency, which calls for a platform approach.

 

'As with most other LLMs, the services powered by this model are expected to run in real-time. To achieve this, you need multiple GPUs. The challenge is how to strike a huge balance between the high performance of the GPUs, the high utilization of the GPUs, and providing a good user experience for the end-users consuming these AI-driven services," he said.

 

 

The Need for Speed

 

With Blackwell, NVIDIA has doubled the bandwidth of each switch, increasing it from 900 GB/s to 1.8 TB/s. The company's Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) technology brings more computing into the systems that actually reside within the switches. It allows us to offload some tasks from the GPU to help accelerate performance and also helps smooth network traffic over the NVLink fabric. These are innovations that we continue to drive at the platform level.

 

The multi-node GB200 NVL72 is a liquid-cooled chassis that connects 72 Blackwell GPUs and 36 Grace CPUs in a rack-scale design. NVIDIA claims that it provides higher inference performance for trillion-parameter LLMs like GPT-MoE-1.8T, effectively functioning as a single GPU. Its performance is 30 times that of the HGX H100 system, with training speed four times faster than the H100.

 

NVIDIA has also added native support for FP4, using the company's Quasar Quantization System, which delivers the same precision as FP16 while reducing bandwidth usage by 75%. The Quasar Quantization System is software that leverages Blackwell's Transformer Engine to ensure accuracy. Salvator demonstrated this by comparing generative AI images created using FP4 and FP16, with little to no discernible difference between the two.

 

Using FP4, models can use less memory and perform even better than FP8 in the Hopper GPU.

 

 

Liquid Cooling Systems

 

In terms of liquid cooling, NVIDIA will introduce a warm water direct chip-to-chip method, which can reduce data center power consumption by 28%.

 

Salvator said, "What's interesting about this method is some of its benefits, which include increased cooling efficiency, lower operating costs, extended server life, and the potential to repurpose captured heat for other uses. It definitely helps improve cooling efficiency. One of the ways this is achieved, as the name suggests, is that this system doesn't actually use chillers. If you think about how a refrigerator works, it works quite well. But it also requires electricity. By adopting this warm water solution, we don't have to use chillers, which saves us some energy and reduces operating costs."

 

Another topic is how NVIDIA is leveraging AI to design its AI chips using Verilog, a hardware description language that has been used for forty years to describe circuits in code. NVIDIA is advancing this effort through an autonomous Verilog agent called VerilogCoder.

 

AI chips

 

He said, "Our researchers have developed a large language model that can accelerate the creation of Verilog code that describes our systems. We will use it in future product generations to help build these codes. It can do a lot of things. It can help speed up the design and verification process. It can accelerate the manual operations of the design and fundamentally automate many tasks."

 

 

 

Send Inquiry