GPU-Interconnect Benchmarks for LLM-Inference: NVLink vs PCIe Switch vs PCIe Direct
Performance comparison of three GPU interconnect architectures for four NVIDIA H200 NVL GPUs, based on real-world LLM inference workloads measured at the MEGWARE Benchmark Center.
How do the three architectures differ in GPU-to-GPU communication and All-Reduce operations?
Response Time & Throughput
How does the interconnect affect LLM response time and throughput as the workload increases?
Long-Context Inference
How do the different interconnects perform with long input contexts?
Download the Interconnect Comparison Whitepaper
The MEGWARE Benchmark Center whitepaper documents the tested GPU interconnect architectures, the LLMs used, detailed benchmark results, and the underlying test methodology.
Complete Benchmark Results Detailed results for NVLink, PCIe Switch, and PCIe Direct.
Real-World LLM Workloads Tests with Qwen3-32B, Llama-3.3-70B, and Qwen3.8-Flash-Next-FP8.
Transparent Test Methodology Detailed information on hardware, software, workloads, and test conditions.