cgroup v2 memory.max — How Do You Cap Process Memory?

cgroup v2 memory.max is the hard memory limit for a cgroup and its descendants. When usage hits the limit and cannot be reduced, the OOM killer runs inside that cgroup. This post covers three axes only: where to write · OOM symptoms · vs containers. No pricing, affiliates, or invented benches.

Grounded in Control Group v2 — Memory (memory.max, memory.events, memory.high, memory.oom.group).

Where do you write it?

One-line answer: Under the unified hierarchy, write a byte value into that cgroup’s memory.max file. The parent must have enabled memory in cgroup.subtree_control so children get the memory interface files.

Manual checklist:

  1. Confirm cgroup2 (often /sys/fs/cgroup; /proc/self/cgroup shows 0::/...).
  2. On the parent, echo +memory > …/cgroup.subtree_control if not already enabled.
  3. mkdir a child cgroup.
  4. echo <bytes> > memory.max. Default is max (no limit).
  5. echo <PID> > cgroup.procs (one PID per write).
CG=/sys/fs/cgroup/demo-job
mkdir -p "$CG"
echo 268435456 > "$CG/memory.max"   # 256 MiB in bytes
echo $$ > "$CG/cgroup.procs"
cat "$CG/memory.current" "$CG/memory.max"

From the docs:

  • Amounts are in bytes; non-PAGE_SIZE-aligned values may round up on read.
  • memory.max is the hard limit. If usage cannot be reduced, the OOM killer is invoked in the cgroup. Usage may briefly exceed the limit.
  • Prefer memory.high (throttle, never invokes OOM) when an external agent should react first—Usage Guidelines treat high as the main control knob.
  • Migrating processes often is discouraged; charged memory does not move with the process (Memory Ownership).

What are the OOM symptoms?

One-line answer: Near or past the hard limit, memory.events counters max / oom / oom_kill increase, and victim processes die via the kernel OOM path. Crossing memory.high alone means throttle/reclaim—not OOM.

Key fields in memory.events (non-root, hierarchical):

KeyMeaning (docs)
maxTimes usage was about to go over the max boundary; failed reclaim leads to OOM state
oomTimes the limit was hit and allocation was about to fail (skipped when OOM is not an option)
oom_killProcesses in this cgroup killed by any OOM killer
oom_group_killGroup OOM occurrences
highTimes the cgroup was throttled into direct reclaim past the high boundary
cat /sys/fs/cgroup/demo-job/memory.events
cat /sys/fs/cgroup/demo-job/memory.current
cat /sys/fs/cgroup/demo-job/memory.events.local

Also watch:

  1. dmesg / journal OOM kill lines naming the cgroup/process.
  2. Allocation failures that return -ENOMEM without invoking the OOM killer (documented for some paths).
  3. memory.oom.group=1 treats the cgroup (and descendants) as all-or-nothing. Tasks with oom_score_adj=-1000 are exempt. An OOM in this cgroup does not kill tasks outside it.
echo 1 > /sys/fs/cgroup/demo-job/memory.oom.group

Do not confuse with memory.high: heavy reclaim and throttle, no OOM, and under extremes the high limit may still be breached.

How does this differ from containers?

One-line answer: Docker, Podman, and Kubernetes memory limits ultimately write the same cgroup v2 memory controller. The difference is who creates the directory and writes memory.max (runtime/orchestrator vs you or systemd).

AxisManual memory.maxContainer runtime
WriterAdmin / script / systemddockerd / conmon / kubelet, …
PathDirect /sys/fs/cgroup/...Runtime-created leaf cgroup
Observememory.current / eventsSame files + docker stats wrappers
ScopePIDs you move inContainer PID-1 tree
ExtrasYou set memory.swap.max, …Mapped from --memory-swap / Limits

Practical reading:

  1. A container with --memory / resources.limits.memory should show a matching memory.max on the host (check units/rounding).
  2. Bare processes and batch jobs get the same hard cap via systemd MemoryMax= or the manual steps above—no container required.
  3. File names differ from v1 (memory.limit_in_bytes); on unified v2 the hard limit is memory.max.
  4. Debug on the host cgroup path (memory.events) to separate cgroup-limit OOM from global pressure.
cat /sys/fs/cgroup/…/memory.max
cat /sys/fs/cgroup/…/memory.events

Frequently asked questions

Is memory.high enough?
Usage Guidelines push high as the main knob; max is the hard stop that can OOM. With a monitor agent, high-first is natural.

Does lowering the limit kill immediately?
If usage is already above the new limit, reclaim and later charges can open the OOM path. Opening memory.max with O_NONBLOCK can bypass synchronous reclaim and oom-kill (for admin writers)—per the docs.

Does memory follow a migrated process?
No. Charge stays with the cgroup that instantiated the memory. Frequent migration is discouraged.

Why is there no memory.max on the root?
By convention the root is exempt from resource-control interface files. Write on non-root cgroups.

What should you remember?

memory.max is the memory controller’s hard limit; failure to reclaim leads to in-cgroup OOM. Write the cgroup’s memory.max, watch memory.events, and remember containers are the same files written by the runtime. Soft control is memory.high. Source: cgroup-v2 Memory. No affiliates.

Where are the official sources?