cgroup v2 memory.max — How Do You Cap Process Memory?
cgroup v2 memory.max is the hard memory limit for a cgroup and its descendants. When usage hits the limit and cannot be reduced, the OOM killer runs inside that cgroup. This post covers three axes only: where to write · OOM symptoms · vs containers. No pricing, affiliates, or invented benches.
Grounded in Control Group v2 — Memory (memory.max, memory.events, memory.high, memory.oom.group).
Where do you write it?
One-line answer: Under the unified hierarchy, write a byte value into that cgroup’s memory.max file. The parent must have enabled memory in cgroup.subtree_control so children get the memory interface files.
Manual checklist:
- Confirm cgroup2 (often
/sys/fs/cgroup;/proc/self/cgroupshows0::/...). - On the parent,
echo +memory > …/cgroup.subtree_controlif not already enabled. mkdira child cgroup.echo <bytes> > memory.max. Default ismax(no limit).echo <PID> > cgroup.procs(one PID per write).
CG=/sys/fs/cgroup/demo-job
mkdir -p "$CG"
echo 268435456 > "$CG/memory.max" # 256 MiB in bytes
echo $$ > "$CG/cgroup.procs"
cat "$CG/memory.current" "$CG/memory.max"
From the docs:
- Amounts are in bytes; non-
PAGE_SIZE-aligned values may round up on read. memory.maxis the hard limit. If usage cannot be reduced, the OOM killer is invoked in the cgroup. Usage may briefly exceed the limit.- Prefer
memory.high(throttle, never invokes OOM) when an external agent should react first—Usage Guidelines treat high as the main control knob. - Migrating processes often is discouraged; charged memory does not move with the process (Memory Ownership).
What are the OOM symptoms?
One-line answer: Near or past the hard limit, memory.events counters max / oom / oom_kill increase, and victim processes die via the kernel OOM path. Crossing memory.high alone means throttle/reclaim—not OOM.
Key fields in memory.events (non-root, hierarchical):
| Key | Meaning (docs) |
|---|---|
max | Times usage was about to go over the max boundary; failed reclaim leads to OOM state |
oom | Times the limit was hit and allocation was about to fail (skipped when OOM is not an option) |
oom_kill | Processes in this cgroup killed by any OOM killer |
oom_group_kill | Group OOM occurrences |
high | Times the cgroup was throttled into direct reclaim past the high boundary |
cat /sys/fs/cgroup/demo-job/memory.events
cat /sys/fs/cgroup/demo-job/memory.current
cat /sys/fs/cgroup/demo-job/memory.events.local
Also watch:
- dmesg / journal OOM kill lines naming the cgroup/process.
- Allocation failures that return
-ENOMEMwithout invoking the OOM killer (documented for some paths). memory.oom.group=1treats the cgroup (and descendants) as all-or-nothing. Tasks withoom_score_adj=-1000are exempt. An OOM in this cgroup does not kill tasks outside it.
echo 1 > /sys/fs/cgroup/demo-job/memory.oom.group
Do not confuse with memory.high: heavy reclaim and throttle, no OOM, and under extremes the high limit may still be breached.
How does this differ from containers?
One-line answer: Docker, Podman, and Kubernetes memory limits ultimately write the same cgroup v2 memory controller. The difference is who creates the directory and writes memory.max (runtime/orchestrator vs you or systemd).
| Axis | Manual memory.max | Container runtime |
|---|---|---|
| Writer | Admin / script / systemd | dockerd / conmon / kubelet, … |
| Path | Direct /sys/fs/cgroup/... | Runtime-created leaf cgroup |
| Observe | memory.current / events | Same files + docker stats wrappers |
| Scope | PIDs you move in | Container PID-1 tree |
| Extras | You set memory.swap.max, … | Mapped from --memory-swap / Limits |
Practical reading:
- A container with
--memory/resources.limits.memoryshould show a matchingmemory.maxon the host (check units/rounding). - Bare processes and batch jobs get the same hard cap via systemd
MemoryMax=or the manual steps above—no container required. - File names differ from v1 (
memory.limit_in_bytes); on unified v2 the hard limit ismemory.max. - Debug on the host cgroup path (
memory.events) to separate cgroup-limit OOM from global pressure.
cat /sys/fs/cgroup/…/memory.max
cat /sys/fs/cgroup/…/memory.events
Frequently asked questions
Is memory.high enough?
Usage Guidelines push high as the main knob; max is the hard stop that can OOM. With a monitor agent, high-first is natural.
Does lowering the limit kill immediately?
If usage is already above the new limit, reclaim and later charges can open the OOM path. Opening memory.max with O_NONBLOCK can bypass synchronous reclaim and oom-kill (for admin writers)—per the docs.
Does memory follow a migrated process?
No. Charge stays with the cgroup that instantiated the memory. Frequent migration is discouraged.
Why is there no memory.max on the root?
By convention the root is exempt from resource-control interface files. Write on non-root cgroups.
What should you remember?
memory.max is the memory controller’s hard limit; failure to reclaim leads to in-cgroup OOM. Write the cgroup’s memory.max, watch memory.events, and remember containers are the same files written by the runtime. Soft control is memory.high. Source: cgroup-v2 Memory. No affiliates.
Where are the official sources?
- Control Group v2 — The Linux Kernel documentation — Memory:
memory.max,memory.high,memory.events,memory.oom.group, Ownership - cgroup-v2.rst (source) — same document source