Skip to main navigation Skip to main content Skip to page footer

MEGWARE Blog

NVLink, PCIe Switch, or PCIe Direct?

  • About 3 min read
Interconnect Benchmark Report

How important is the connection between multiple GPUs for LLM inference?

Choosing the right GPU is only part of the equation. Just as important is how the GPUs are connected to each other. At the MEGWARE Benchmark Center, we compared three different GPU interconnect architectures in a server equipped with four NVIDIA H200 NVL GPUs.

When planning AI servers, the focus is usually on the GPUs: compute performance, HBM memory, and the number of GPUs. But especially when running large language models, there is another factor that can be crucial:

How are the GPUs connected to each other?

As soon as a model is distributed across multiple GPUs, the GPUs need to continuously exchange data during computation. With tensor parallelism (each layer's weight matrices are split across GPUs, so every GPU computes only part of each operation), this includes so-called all-reduce operations — making the interconnect an integral part of the system’s actual computing performance.

Comparing Three GPU Topologies

In its latest Benchmark Report, MEGWARE therefore examined three different approaches:

  • NVLink – direct connection between all four GPUs
  • PCIe Switch – four GPUs connected via a shared PCIe switch
  • PCIe Direct – each GPU connected directly to the CPU via its own PCIe connection

Testing was conducted on a system equipped with four NVIDIA H200 NVL GPUs and several current large language models (Qwen3-32B, Llama-3.3-70B, Qwen3.8-Flash-Next-FP8). The focus was not only on theoretical bandwidth, but primarily on a practical question:

How strongly does the choice of interconnect affect LLM response times under real-world load?

The Difference Becomes Clear Under Load

With a single request, the differences between the various configurations may appear relatively small.

But what happens when many requests are processed simultaneously?

This is where GPU-to-GPU communication becomes increasingly important. In our Benchmark Report, we therefore examined, among other scenarios, a sustained high level of concurrency with 64 simultaneous requests. The results show that the choice of interconnect can have a significant impact on how quickly users receive the first response.

There is a particularly interesting difference between a direct PCIe connection and a configuration in which the GPUs communicate through a shared switch.

Exactly how large this difference is, and at what point it becomes relevant for production LLM applications, is shown in the full Benchmark Report.

The answer is less straightforward than it might initially seem.

When NVLink is available, it is tempting to assume that the faster GPU-to-GPU connection is automatically the best choice. But how large is the actual performance difference?

And if NVLink cannot be used:

Is a direct PCIe connection sufficient — or does a shared PCIe switch become a bottleneck?

We also investigated the question of CPU topology: Does it make a difference whether four GPUs are connected to a single CPU socket or distributed across two CPU sockets?

Find the answers in the full Benchmark Report.

Particularly Relevant for Modern AI Workloads

The study also goes beyond short chat prompts.

Current applications such as coding assistants, retrieval-augmented generation, and autonomous agents increasingly work with very large contexts. Long source-code files, documents, or extensive tool histories can significantly change the requirements placed on an inference system.

We therefore also investigated how the different interconnects perform with substantially longer contexts.

Does a prompt that is 16 times longer actually change the performance relationship between the different GPU interconnects?

The Benchmark Report provides concrete measurement results