Skip to main navigation Skip to main content Skip to page footer

AMD Instinct™ MI350P PCIe now available in the MEGWARE Benchmark Center

| Partner

The MEGWARE Benchmark Center has been extended with AMD's latest CDNA™ 4 accelerator: the AMD Instinct™ MI350P PCIe card is available now for customer test workloads.

144 GB of HBM3E in a standard PCIe form factor

With 144 GB of HBM3E per card, the MI350P lets you run models on a single GPU that previously required two or four accelerators. Because it is a standard dual-slot PCIe 5.0 add-in card, it drops into existing server platforms - no OAM baseboard, no new rack design. That makes it an interesting option for customers who want to run generative and agentic AI inside an installed base rather than build a new island.

Key specifications

ArchitectureAMD CDNA™ 4 (TSMC 3nm), 128 CUs, 512 Matrix Cores 
Memory144 GB HBM3E, 4 TB/s, 128 MB LLC, full-chip ECC 
Form FactorPCIe 5.0 x16, dual-slot, full height, 267 mm, passive cooling 
Power600 W TBP, configurable down to 450 W 
Peak performance4.6 PFLOPS MXFP4/MXFP6, 2.3 PFLOPS FP8/INT8, 1.15 PFLOPS FP16/BF16, 36 TFLOPS FP64 
VirtualizationSR-IOV, RAS support 

New in CDNA 4 is native MXFP4 and MXFP6 support - a genuine step up over MI300X for quantized inference, and the reason very large MoE models now fit into a single card.

What you can run on it

As a rule of thumb, 144 GB of HBM leaves room for roughly:

  • 60B parameters at BF16
  • 120B at FP8
  • 240B at FP4

including KV cache and activations.

Examples that run on a single MI350P under ROCm:

  1. Mistral Small 4
    • 119B total, 6.5B active
    • ~119 GB at FP8
    • Low-rank MLA keeps the KV cache small
    • European weights
  2. Qwen3.5 122B-A3B and Nemotron 3 Super 120B-A12B
    • ~120 GB at FP8
    • ~60 GB at FP4
  3. Qwen3.8-Flash-Next
    • ~180B total, 6B active
    • ~93 GB at FP4
    • 262K context
  4. gpt-oss-120b
    • 65 GB in MXFP4
    • Apache 2.0. 
    • Native MXFP4 weights map directly onto CDNA 4's FP4 path
  5. Qwen3-Coder-Next 80B-A3B
    • ~80 GB at FP8
    • coding and agent workloads
  6. FLUX.1-dev, SDXL and Whisper large-v3
    • image generation and ASR
  7. LoRA- and QLoRA-Finetuning
    • of 70B-class models

Larger frontier models such as DeepSeek-R1/V3 (671B) scale across multiple cards.

Software-Stack

ROCm™ 7 with vLLM, SGLang, llama.cpp, PyTorch, JAX, Triton, ONNX Runtime and TensorFlow. Pre-built containers are available via AMD Infinity Hub and the AMD Enterprise AI Reference Stack, so most existing CUDA-based pipelines port with minimal effort via HIP.

Test it now!

Benchmarking your own models and applications in the MEGWARE Benchmark Center is free of charge for an initial analysis. Our HPC and AI engineers support you with setup, tuning and interpretation of the results - and advise on the right configuration for production.

More information about AMD InstinctTM MI350P