NUMA-Aware Scheduling: Pinning AI Processes for Maximum Throughput
A DGX Station running four GPUs on a dual-socket motherboard. The vibration-analysis AI model is loaded on GPU-2, which sits on Socket 1's PCIe root complex. But the data-loading process runs on Socket 0's cores — because the Linux scheduler put it there. Every tensor transfer from CPU memory to GPU memory now crosses the QPI interconnect between sockets, adding ~60ns per access and cutting inference throughput by 19%. The model is fine. The hardware is fine. The scheduling is wrong. NUMA-aware process pinning fixes this without touching a single line of model code, without buying new hardware, and without changing the inference framework. It is the single highest-ROI infrastructure optimization that most AI deployments never make. Sign up free to audit your NUMA topology and pinning configuration.
NUMA PINNING · CPU AFFINITY · MEMORY BINDING · IRQ ROUTING
Same Hardware. Same Model. 19% More Throughput. The Only Change Is Where the Process Runs.
NUMA-aware scheduling pins AI processes to the CPU cores and memory banks on the same socket as the GPU they serve. Local memory access runs at ~80ns. Cross-socket remote access runs at ~140ns — nearly 2× the latency on every tensor transfer, every weight load, every data preprocessing batch. On a 4-GPU DGX server running 24/7 inference, that 2× penalty compounds into 15-30% lost throughput that no model optimization, no framework upgrade, and no hardware purchase can recover. Only correct scheduling fixes it.
Five NUMA Rules That Recover 15-30% of Your AI Throughput
Each rule addresses a specific source of cross-socket traffic that the default Linux scheduler ignores. Apply them in order — each one stacks on the previous. Most deployments see measurable improvement after Rule 1 alone. Sign up free to get a NUMA audit report for your deployment.
01MAP YOUR GPU-TO-SOCKET TOPOLOGYFoundation
Before pinning anything, discover which GPU sits on which NUMA node. Every GPU connects to a specific PCIe root complex on a specific CPU socket. If you do not know this mapping, every subsequent rule is a guess.
nvidia-smi topo -m shows GPU-to-NUMA affinity. lscpu --extended shows CPU core-to-NUMA mapping. numactl -H shows memory per NUMA node. Cross-reference all three to build the topology map.
02PIN AI PROCESSES TO THE GPU'S NUMA NODE+15-19%
Bind each AI inference or training process to the CPU cores and memory banks on the same socket as its GPU. This ensures every tensor transfer, weight load, and data buffer allocation uses local memory at ~80ns — not remote memory at ~140ns.
numactl --cpunodebind=1 --membind=1 ./inference_gpu2 pins the process to Socket 1's cores and memory (assuming GPU 2 is on Socket 1). For Python: os.sched_setaffinity() for CPU + libnuma for memory binding together.
03DISABLE AUTO NUMA BALANCING+3-5%
The Linux kernel's automatic NUMA balancing migrates pages and tasks between NUMA nodes to optimize locality. For general workloads this is helpful. For deliberately pinned AI processes, it is destructive — it moves memory you already placed correctly, adding migration overhead and breaking your pinning.
echo 0 > /proc/sys/kernel/numa_balancing disables it for the running session. For persistence: sysctl -w kernel.numa_balancing=0 in /etc/sysctl.conf. Every HPC vendor and NVIDIA's own DGX documentation recommends this for deterministic AI workloads.
04PIN IRQ AFFINITY FOR I/O COMPLETIONS+2-4%
NVMe storage and network interrupts (IRQs) should complete on the same socket feeding the GPU. If NVMe sits on Socket 0 but its interrupt handler runs on Socket 1, every I/O completion crosses sockets — adding latency to data loading that sits on the critical path of inference throughput.
Set IRQ affinity via /proc/irq/<irq_num>/smp_affinity_list to restrict each NVMe and NIC interrupt to cores on the same NUMA node as the device's PCIe slot. Use cat /proc/interrupts to identify which IRQs belong to which devices.
05SET I/O SCHEDULER AND HUGEPAGES+1-3%
Switch NVMe I/O scheduler to none (lowest overhead for fast storage). Allocate hugepages (2MB or 1GB) on the correct NUMA node to reduce TLB misses during large tensor allocations. Both reduce per-access overhead on the hot path.
echo none > /sys/block/nvme0n1/queue/scheduler for I/O scheduler. numactl --membind=1 hugeadm --pool-pages-min 2MB:1024 for per-node hugepages. Verify with numastat -m that hugepages are allocated on the intended node.
15-30%
Total throughput recovery · all 5 rules stacked
0 lines
Of model code changed
$0
Hardware cost · scheduling only
p99
Tail latency improvement · not just avg
The five rules compound. Rule 1 is discovery. Rule 2 alone delivers the majority of the gain (15-19%). Rules 3-5 recover the remaining 5-10%. The total effect on p95/p99 tail latency is often larger than the average throughput gain — because cross-socket access creates variance, not just overhead. Book a free demo to see NUMA optimization applied to your AI deployment.
"Our DGX Station runs 4 GPUs doing vibration-analysis inference for 200 assets. We were getting 340 inferences per second. After NUMA pinning: 412 inferences per second. Same hardware, same model, same framework."
THE PROBLEM
Manufacturing plant running vibration-analysis AI on a DGX Station with 4 GPUs across 2 sockets. The deployment team installed the inference server using default settings — no NUMA awareness. The Linux scheduler spread data-loading threads across both sockets regardless of which GPU the inference ran on. GPU-2 and GPU-3 (on Socket 1) consistently underperformed GPU-0 and GPU-1 (on Socket 0) by 22%. The default scheduler happened to place more data-loading threads on Socket 0, meaning GPUs on Socket 1 paid the cross-socket penalty on every batch. Total throughput: 340 inferences/second across all 4 GPUs.
THE FIX · 30 MINUTES
Rule 1 · Topology Map
nvidia-smi topo -m confirmed GPU 0-1 on NUMA node 0, GPU 2-3 on NUMA node 1. lscpu --extended confirmed cores 0-23 on node 0, cores 24-47 on node 1. Mapping complete in 5 minutes.
Rule 2 · Process Pinning
Inference processes wrapped with numactl --cpunodebind=0 --membind=0 for GPU 0-1 and --cpunodebind=1 --membind=1 for GPU 2-3. Data-loading threads pinned to the same node as their GPU. Applied via systemd service file.
Rules 3-5 · Hardening
Auto NUMA balancing disabled. NVMe IRQs pinned to matching socket. Hugepages allocated per-node. Total configuration time including validation: 30 minutes. Zero downtime — applied during a scheduled maintenance window.
THE RESULT
340 → 412 inferences/sec (+21%). GPU 2-3 throughput equalized with GPU 0-1. 30-minute fix. Zero hardware cost. Zero model changes. 21% more capacity from existing investment.
SCENARIO 02
"Our vision quality-inspection AI had random p99 latency spikes — 4× the median. Occasionally a good bottle would get rejected because the inference took too long and the bottle passed the reject point. Root cause: cross-socket memory access jitter."
THE PROBLEM
Beverage bottling plant running vision-based quality inspection on a 2-socket server with 2 GPUs. Median inference latency: 8ms (well within the 30ms reject-window budget). But p99 latency: 34ms — exceeding the reject window. Approximately 1 in 100 inspections arrived too late, causing either a missed defect or a false reject (good bottle past the actuator). Root cause: the Linux scheduler's automatic NUMA balancing periodically migrated inference memory pages to the "wrong" socket, causing latency spikes during the migration. The spikes were invisible in average metrics but devastating at p99.
THE FIX · 15 MINUTES
Rule 2 · Process Pinning
Vision inference process pinned to Socket 0 cores + memory (where the GPU's PCIe root sits). Data-loading thread pinned to the same node. Camera input buffer allocated on Socket 0 memory via explicit --membind.
Rule 3 · Auto Balancing Off
kernel.numa_balancing=0 applied. The kernel stopped migrating pages between sockets. The periodic latency spikes — caused by page migration — disappeared entirely.
Validation
p99 latency dropped from 34ms to 11ms — comfortably inside the 30ms reject window. Zero late inspections. Zero false rejects from timing. Applied in 15 minutes during a line changeover.
THE RESULT
p99 latency 34ms → 11ms. Late inspections eliminated. 15-minute fix. False-reject rate dropped 40%. The "random" spikes were NUMA page migration all along.
Does NUMA matter for single-socket servers like the RTX PRO 6000?
Single-socket systems have only one NUMA node — all memory access is "local" by definition. NUMA pinning is unnecessary because there is no remote socket to accidentally schedule on. However, CPU affinity still matters: pinning inference processes to specific core ranges prevents context-switching overhead and improves cache locality. The RTX PRO 6000 tower is single-socket with one GPU — NUMA is not an issue. The DGX Station GB300 with multiple sockets is where NUMA optimization is critical.
Can NUMA pinning hurt performance in some cases?
Yes — in two situations. First, if one NUMA node has CPUs but no GPU, pinning all work to the GPU's node starves CPU resources. Second, if the GPU's NUMA node has insufficient memory for the model and data, forcing all allocations to that node causes swapping. The solution: always map the topology first (Rule 1), verify memory capacity per node with numactl -H, and ensure the GPU's node has enough free memory for the model weights + data buffers + overhead. For large models that exceed single-node memory, interleaved allocation across nodes is sometimes the pragmatic choice.
How do I verify NUMA pinning is working correctly?
Three checks. First: numastat -p <pid> shows memory allocation per NUMA node for a running process — local_node should be close to 100%, other_node close to 0%. Second: perf stat -e node-load-misses,node-store-misses on the process — these counters show cross-node (remote) memory accesses. Third: measure p99 latency before and after pinning — the tail-latency improvement is often more dramatic than the average throughput gain. If node-load-misses drops by 90%+ and p99 latency drops by 50%+, the pinning is working.
How does this interact with container orchestration (Docker, Kubernetes)?
Kubernetes supports NUMA-aware resource allocation via the Topology Manager (policy: single-numa-node or restricted). When enabled, the kubelet ensures that CPU cores, memory, and GPU devices assigned to a pod are all on the same NUMA node. For Docker without Kubernetes, use --cpuset-cpus and --cpuset-mems flags to pin the container to specific cores and NUMA nodes. The key: if your orchestrator is not NUMA-aware by default, it will scatter pod resources across NUMA nodes — losing 15-30% throughput silently.
Does OxMaint handle NUMA optimization automatically?
Yes. The OxMaint deployment team runs a NUMA topology audit during the initial site deployment (Weeks 3-4 of the pilot). The inference services are configured with proper numactl bindings in the systemd service files, auto NUMA balancing is disabled, IRQ affinity is set for NVMe and NIC devices, and hugepages are allocated per-node. The configuration is documented, version-controlled, and validated against performance benchmarks before handover. For DGX and multi-socket deployments, this is a standard part of the deployment checklist — not an optional add-on.
NUMA Optimization · 15-30% Recovery · Zero Hardware Cost
Same Hardware. Same Model. 15-30% More Throughput. The Fix Takes 30 Minutes.
Book a 30-minute call with our infrastructure engineers. We will map your GPU-to-NUMA topology, identify cross-socket traffic, and show you exactly how much throughput you are leaving on the table. Perpetual license, source code included, $0/mo.