Hugepages — 2 MiB and 1 GiB pages on GPU/HPC nodes
Why hugepages matter for RDMA and pinned-memory workloads, how to allocate them at boot vs runtime, and how to validate they're actually being used.
help for the full list, or solutions for copy-paste fix recipes.Linux defaults to 4 KiB pages. A modern 8-GPU node with 1 TiB of RAM holds 256 million page table entries to address all of it. The TLB does not have 256 million slots. Every TLB miss costs you a page-table walk, and at HPC working-set sizes you spend a measurable percentage of wallclock in TLB walks.
Hugepages collapse that. A 2 MiB page covers 512× more memory per TLB entry; a 1 GiB page covers 262144× more. For workloads that pin large buffers (RDMA, GPUDirect, in-memory databases, big tensor allocators), this is free performance — provided you allocate them at the right time and your application actually uses them.
4 KiB vs 2 MiB vs 1 GiB
| Page size | Allocated where | Use case | Cost |
|---|---|---|---|
| 4 KiB | Anywhere, demand-paged | Default. Anything fine-grained. | TLB pressure on large working sets |
| 2 MiB | Reserved pool or THP | RDMA buffers, GPU pinned host mem, JVM heaps | Less fragmentation tolerance |
| 1 GiB | Reserved at boot only | Big DBs, very large RDMA/pinned buffers, DPDK | Must be reserved before fragmentation |
1 GiB pages have one operationally-painful constraint: they must be allocated at boot from contiguous physical memory. After the system has been up for a while and memory is fragmented, the kernel cannot assemble a 1 GiB physically-contiguous region. So 1 GiB hugepages are a GRUB cmdline decision, not a runtime decision.
2 MiB pages can be allocated at runtime via sysfs, though it's still better to do it early — and Transparent Hugepages (THP) can hand out 2 MiB pages opportunistically without the application asking.
Reserving 1 GiB pages at boot (GRUB)
/etc/default/grub:
GRUB_CMDLINE_LINUX_DEFAULT="... default_hugepagesz=1G hugepagesz=1G hugepages=64"
This reserves 64 × 1 GiB = 64 GiB of 1 GiB hugepages, system-wide, at boot. default_hugepagesz=1G makes 1 GiB the size used when an application calls mmap(... MAP_HUGETLB ...) without an explicit size.
Apply:
update-grub # Debian/Ubuntu
# or
grub2-mkconfig -o /boot/grub2/grub.cfg # RHEL/Rocky
reboot
Validate after reboot:
grep -i huge /proc/meminfo
# AnonHugePages: 104448 kB
# ShmemHugePages: 0 kB
# FileHugePages: 0 kB
# HugePages_Total: 64
# HugePages_Free: 64
# HugePages_Rsvd: 0
# HugePages_Surp: 0
# Hugepagesize: 1048576 kB <-- 1 GiB
# Hugetlb: 67108864 kB
HugePages_Total: 64 confirms the reservation; Hugepagesize: 1048576 kB confirms 1 GiB granularity.
To request both sizes (1 GiB pool plus a 2 MiB pool):
GRUB_CMDLINE_LINUX_DEFAULT="... default_hugepagesz=1G hugepagesz=1G hugepages=32 hugepagesz=2M hugepages=8192"
That's 32 × 1 GiB + 8192 × 2 MiB = 32 GiB + 16 GiB reserved. Verify with:
ls /sys/kernel/mm/hugepages/
# hugepages-1048576kB hugepages-2048kB
cat /sys/kernel/mm/hugepages/hugepages-1048576kB/nr_hugepages
# 32
cat /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages
# 8192
Reserving 2 MiB pages at runtime
Per-NUMA-node, written via sysfs:
# Reserve 2048 × 2 MiB = 4 GiB on node 0
echo 2048 > /sys/devices/system/node/node0/hugepages/hugepages-2048kB/nr_hugepages
# And the same on node 1 — for a multi-socket GPU node, you almost always want
# hugepages on BOTH nodes so neither socket pays remote-memory cost.
echo 2048 > /sys/devices/system/node/node1/hugepages/hugepages-2048kB/nr_hugepages
Confirm:
cat /sys/devices/system/node/node{0,1}/hugepages/hugepages-2048kB/nr_hugepages
# 2048
# 2048
grep ^Huge /proc/meminfo
# HugePages_Total: 4096
# HugePages_Free: 4096
# Hugepagesize: 2048 kB
If HugePages_Total ends up lower than what you requested, the kernel could not find enough contiguous memory — fragmentation has already happened. Either reboot and put it on GRUB cmdline, or do a echo 1 > /proc/sys/vm/compact_memory first and try again.
Persist the runtime reservation across reboots: drop it in /etc/sysctl.d/:
# /etc/sysctl.d/98-hugepages.conf
vm.nr_hugepages = 4096
(This sets the total across all NUMA nodes; the kernel splits it evenly. For per-NUMA control, write the sysfs files in a systemd unit instead.)
Transparent Hugepages
THP is the kernel's opportunistic hugepage allocator — it tries to back anonymous mappings with 2 MiB pages without the application asking. Three modes:
cat /sys/kernel/mm/transparent_hugepage/enabled
# always [madvise] never
- always — kernel uses THP for everything. Easy wins, but
khugepagedpauses can show up as latency spikes for jitter-sensitive workloads. - madvise — only for mappings the application explicitly asks for via
madvise(MADV_HUGEPAGE). The PyTorch / CUDA stack, jemalloc, and tcmalloc all do this. This is the default on most distros and the right choice for GPU nodes. - never — disabled. Choose this only if you've measured
khugepagedcausing problems for your specific workload (HFT, RT).
Set it:
echo madvise > /sys/kernel/mm/transparent_hugepage/enabled
Persist via tuned profile, GRUB (transparent_hugepage=madvise), or /etc/rc.local.
AnonHugePages in /proc/meminfo shows how much THP is currently in use.
When hugepages help (and when they don't)
| Workload | 2 MiB / THP | 1 GiB reserved |
|---|---|---|
GPU pinned host memory (cudaHostAlloc) | Yes | Sometimes |
| GPUDirect RDMA buffers | Yes | Yes (large) |
| NCCL host-staging buffers | Yes | No |
| Slurm jobs allocating multi-GB host arrays | Yes | No |
| In-memory databases (Redis, KDB, Aerospike) | Yes | Yes |
| DPDK / userspace networking | No | Yes (req) |
| Generic web/microservice workloads | Marginal | No |
The 1 GiB pool is wasted on workloads that don't make multi-GB allocations — the pages are reserved, meaning they're unavailable to anything that doesn't explicitly request hugepages. Reserving 64 GiB of 1 GiB pages on a node that doesn't use them is 64 GiB of unusable RAM.
Validation: is my application actually using hugepages?
# Per-process anonymous hugepage usage
grep -i huge /proc/<pid>/status
# AnonHugePages: 12288 kB
# ShmemHugePages: 0 kB
# FileHugePages: 0 kB
# HugetlbPages: 33554432 kB <-- explicit hugetlb mappings
Or pmap:
pmap -X <pid> | grep -E 'KernelPageSize|Anon|Huge'
Look for KernelPageSize=2048 kB or 1048576 kB on big mappings; that's the proof a specific mmap region is backed by hugepages.
For Kubernetes pods that requested hugepages-1Gi:
kubectl exec -it <pod> -- grep ^Huge /proc/meminfo
# HugePages_Total: 8 <-- reserved for this pod's cgroup
# HugePages_Free: 8
The kubelet only sees and exposes a hugepage resource if the host has a non-zero pool, and the kubelet was started with --feature-gates=HugePages=true (default since 1.14). You also need the kubelet config to know about the page size:
# /var/lib/rancher/rke2/agent/etc/kubelet/config.yaml or wherever
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
HugePages: true
A pod requests them like:
resources:
limits:
hugepages-1Gi: 8Gi
memory: 16Gi
Common failure modes
mmapreturnsENOMEMwith hugepages. The pool is fully consumed; some other process on the node already grabbed the pages. Check/proc/meminfo:HugePages_Free.HugePages_Total: 0despite GRUB cmdline saying otherwise. The cmdline didn't take. Checkcat /proc/cmdline— if your params aren't there,update-grubwasn't run or the wrong default file was edited.- High
khugepagedCPU usage with THP=always. Switch tomadvise. - Pinned memory allocations on GPU nodes are slow. CUDA pinned memory (
cudaHostAlloc) benefits hugely from THP=madvise + a 2 MiB pool. Without it, the kernel pins individual 4 KiB pages and the bookkeeping dominates.
See also
- GRUB cmdline reference — where
default_hugepageszandhugepageszgo - NUMA pinning — hugepages and NUMA balance interact (allocate per-node)
- Sysctl tuning —
vm.nr_hugepages,vm.nr_overcommit_hugepages
External:
Documentation/admin-guide/mm/hugetlbpage.rst(Linux kernel)Documentation/admin-guide/mm/transhuge.rst- Kubernetes hugepages docs: kubernetes.io/docs/tasks/manage-hugepages/scheduling-hugepages/