↓ Skip to main content

Proxmox GPU Passthrough Guide

·6 mins
Table of Contents
After running Ollama on CPU-only in my R720 and getting 0.46 tokens/s (roughly 18 minutes per response), I picked up a Tesla P40 to fix that. The P40 is a datacenter card, no display output, 24GB GDDR5, runs at ~€100–150 used. To use it inside a VM on Proxmox, you need to pass it through via VFIO.

What you need: a CPU with IOMMU support (Intel VT-d or AMD-Vi), a motherboard BIOS that exposes it, Proxmox VE 8.x, and ideally a second GPU for the Proxmox host so the host still has display output. Intel integrated graphics on the host while passing a discrete GPU to the VM is the cleanest setup.

Step 1: Enable IOMMU
#

In BIOS/UEFI: Intel → enable VT-d. AMD → enable AMD-Vi or IOMMU.

On AMD the kernel turns the IOMMU on by itself, and so does Intel from kernel 6.8 (Proxmox VE 8.2) onwards. On an older Intel kernel, add intel_iommu=on to the kernel command line. Add iommu=pt on either vendor: devices that aren’t passed through skip the IOMMU translation, which saves overhead. Where the command line lives depends on the bootloader, and proxmox-boot-tool status tells you which one you have:

In /etc/default/grub:

GRUB_CMDLINE_LINUX_DEFAULT="quiet intel_iommu=on iommu=pt"
update-grub && reboot

Append to the single line in /etc/kernel/cmdline:

intel_iommu=on iommu=pt
proxmox-boot-tool refresh && reboot

Verify after reboot:

dmesg | grep -e DMAR -e IOMMU

Intel shows DMAR: IOMMU enabled, AMD shows AMD-Vi lines. Nothing means IOMMU isn’t active; go back to the BIOS.

Also check interrupt remapping, since passthrough won’t work without it:

dmesg | grep 'remapping'

You should see AMD-Vi: Interrupt remapping enabled or DMAR-IR: Enabled IRQ remapping. If not, and your hardware doesn’t support it, you can allow unsafe interrupts as a workaround:

echo "options vfio_iommu_type1 allow_unsafe_interrupts=1" > /etc/modprobe.d/iommu_unsafe_interrupts.conf

Check the IOMMU groups
#

for d in /sys/kernel/iommu_groups/*/devices/*; do
    n=${d#*/iommu_groups/*}; n=${n%%/*}
    printf 'IOMMU Group %s ' "$n"
    lspci -nns "${d##*/}"
done

Your GPU and its audio device need to be in the same IOMMU group, ideally alone in it. If they share a group with a SATA controller or USB hub, see ACS override in the troubleshooting section.

Step 2: Load the VFIO modules
#

Add to /etc/modules:

vfio
vfio_iommu_type1
vfio_pci

Older guides also list vfio_virqfd. Since kernel 6.2 (every Proxmox VE 8 kernel) it’s part of vfio and the separate module no longer exists.

update-initramfs -u -k all && reboot

Step 3: Bind the GPU to VFIO
#

Find your GPU’s PCI IDs:

lspci -nn | grep -iE "nvidia|amd|vga|3d"

Output example:

01:00.0 VGA compatible controller [0300]: NVIDIA GeForce RTX 3080 [10de:2206]
01:00.1 Audio device [0403]: NVIDIA HD Audio [10de:1aef]

Bind both the GPU and its audio device to VFIO:

echo "options vfio-pci ids=10de:2206,10de:1aef disable_vga=1" > /etc/modprobe.d/vfio.conf
update-initramfs -u -k all && reboot

Verify:

lspci -nnk -d 10de:2206

Look for Kernel driver in use: vfio-pci. If you see nvidia or nouveau instead, blacklist them, one module per line, since blacklist doesn’t take wildcards:

cat >> /etc/modprobe.d/blacklist.conf <<'EOF'
blacklist nouveau
blacklist nvidia
blacklist nvidiafb
EOF
update-initramfs -u -k all && reboot

Step 4: Create the VM
#

In the Proxmox web UI, these settings matter:

  • Machine: q35, required for PCIe passthrough
  • BIOS: OVMF (UEFI); add an EFI disk when prompted
  • CPU type: host; don’t use kvm64 or the default
  • Memory: disable ballooning, set a fixed amount; a VM with a passed-through device pins all of its memory anyway
  • OS: for Windows, add the VirtIO drivers ISO as a second CD-ROM

Attaching the card
#

Hardware → Add → PCI Device, then check:

  • All Functions: passes GPU + audio together
  • Primary GPU: only if this GPU handles the VM’s display output
  • PCI-Express: presents it as a PCIe device, which q35 needs
  • ROM-Bar: on by default; leave it on
Warning

If “Primary GPU” is checked, set vga: none in the VM config, otherwise the VM may boot to the virtual display instead of the GPU.

Step 5: NVIDIA Error 43
#

Note

NVIDIA drivers 465 and newer no longer trigger Error 43 in VMs. On a recent driver, skip this step.

Older consumer NVIDIA drivers detect the hypervisor and refuse to initialize. Proxmox can hide it without raw QEMU arguments, in /etc/pve/qemu-server/<VMID>.conf:

cpu: host,hidden=1,hv-vendor-id=NV43FIX

hidden=1 removes the KVM signature from CPUID, and hv-vendor-id replaces the Hyper-V vendor string with any string up to 12 characters.

For Windows VMs where GeForce Experience or other GPU software crashes the VM, have KVM ignore the unknown MSRs it touches:

echo "options kvm ignore_msrs=1 report_ignored_msrs=0" > /etc/modprobe.d/kvm.conf

Step 6: What’s different about a Tesla P40
#

The P40 is a datacenter card: 24GB GDDR5, no display output, pure compute, and the best value per GB for homelab AI.

24 GB

GDDR5

the whole point of the card

250 W

TDP

full-length, full-height

0

fans

passive, needs chassis airflow

Compared with passing through a gaming GPU:

  • No display output: leave “Primary GPU” unchecked; the VM keeps its virtual display for management
  • No Error 43 workaround: datacenter drivers load cleanly in VMs
  • Passive cooling: in a server chassis the fans handle it; in a desktop case, point active airflow at the card
  • Power: check the slot, the length and the PSU before buying
Warning

The P40 maps a 32GB BAR, larger than OVMF’s default 64-bit MMIO window, and without room for it the driver fails with a BAR0 error. OVMF sizes the window from the CPU’s physical address bits, so CPU type host (already set in Step 4) fixes it. With any other CPU type, pass the host’s address width through: qm set <VMID> --cpu x86-64-v2-AES,phys-bits=host.

Install the driver and CUDA inside the VM (Ubuntu 22.04):

wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update && sudo apt install -y cuda-toolkit-12-4 nvidia-driver-550-server

The P40 is Pascal. CUDA 13 dropped Pascal support and the 580 driver branch is the last to carry it, so keep this VM on CUDA 12.x and a driver of 580 or older. An unpinned apt upgrade to the newest toolkit is what breaks it.

Reboot and verify:

nvidia-smi
# Tesla P40 | 24576MiB | ...

Troubleshooting
#

Black screen after the VM boots. Four causes, in the order worth checking:

  1. The GPU isn’t bound to vfio-pci; confirm Kernel driver in use: vfio-pci with lspci -nnk
  2. The machine type isn’t q35
  3. “Primary GPU” is checked but vga: none isn’t set, or the reverse
  4. Wrong PCIe slot: some boards only expose IOMMU correctly on specific slots

Error 43 is still there. Check the cpu: line survived; a leftover args: -cpu ... line from an older guide overrides it. Failing that, change the hv-vendor-id string.

The GPU hangs after the VM shuts down. Some cards can’t complete a PCIe reset, so they’re stuck until the whole host reboots. It’s mostly an AMD problem (Polaris, Vega and Navi 10), and vendor-reset implements the reset sequences those cards need:

apt install proxmox-headers-$(uname -r) build-essential git
git clone https://github.com/gnif/vendor-reset
cd vendor-reset && make && make install
echo "vendor-reset" >> /etc/modules
update-initramfs -u -k all && reboot

On kernel 5.15 and newer the module alone isn’t enough; each card also has to be told to use it, before the VM starts: echo device_specific > /sys/bus/pci/devices/0000:01:00.0/reset_method.

The IOMMU groups are wrong.

Warning

ACS override weakens IOMMU isolation. Use only if you understand the security implications.

If your GPU shares a group with a SATA controller or USB hub, ACS override splits them apart. Add it to the kernel command line from Step 1, and refresh the bootloader the same way:

pcie_acs_override=downstream,multifunction

References
#