Accept GPU assets down to the PCIe and NUMA topology
NUMA (Non-Uniform Memory Access): an architecture in which path cost depends on which memory and PCIe devices a CPU accesses.
Matching only the GPU count is not enough. The GPU UUIDs and serials, PCIe generation/width, NUMA locality, NVLink/NVSwitch, PSU, and cooling state must all match the design topology.
A path that crosses a CPU socket and a PCIe switch affects GPU-to-NIC and GPU-to-storage traffic. If a link negotiates to a lower generation or width, functionality may still work while performance can drop sharply.
Receiving evidence must link beyond the carton serial to the identity read by the firmware and the OS. Compare the inventory and corrected error deltas before and after burn-in to find early-life defects.
In the training example the GPU and NIC are attached to CPU socket 0 while the data processing thread runs on socket 1. The fact that every device is visible does not let you say the data path is short. Record the GPU, NIC and CPU affinity and the PCIe link width first, then compare a representative load in the same placement. Move a card between slots only through a maintenance procedure that has confirmed the supported cabling and power configuration.
- Why does this happen?
- A path that crosses the CPU socket and a PCIe switch affects GPU-to-NIC and GPU-to-storage traffic. If a link negotiates at a lower generation or width, functionality may still work while performance can drop sharply.
- When is it a problem?
- If you see a missing GPU, a PCIe downgrade, or an Xid event, there are insufficient grounds to proceed.
- Common beginner misconceptions
- Before opening a chassis or reseating a GPU, you need procedures for power isolation, discharge, ESD, and heavy lifting. This lab is read-only.
- How to verify it yourself
- Link the chassis and GPU UUIDs and serials to the asset list. Compare the current and maximum speed and width in lspci.