Tune Linux virtual memory safely with /proc/sys/vm
You will finish with a safe way to inspect the kernel's virtual-memory controls, make one reversible test change, and recognise the settings that can affect allocation, swapping, out-of-memory handling or security. The examples match Linux 6.8.0-139-generic and the installed Linux man-pages 6.7 package. Allow about fifteen minutes, plus time to observe the workload after a change.
The route
Jump straight to the step you need, or tick off Done means at the end.
You need a shell and a readable /proc filesystem. Reading values is normally unprivileged. Writing them requires root or the relevant kernel capability, so use sudo only for the individual write that you have decided to make. These files are live kernel state, not ordinary configuration files.
1. Inspect the current values
Start with a read-only snapshot of the controls most likely to be discussed during troubleshooting:
$ for name in overcommit_memory overcommit_ratio overcommit_kbytes swappiness admin_reserve_kbytes user_reserve_kbytes oom_dump_tasks oom_kill_allocating_task panic_on_oom unprivileged_userfaultfd; do
> printf '%s=' "$name"
> cat "/proc/sys/vm/$name"
> done
overcommit_memory=0
overcommit_ratio=50
overcommit_kbytes=0
swappiness=10
admin_reserve_kbytes=8192
user_reserve_kbytes=131072
oom_dump_tasks=1
oom_kill_allocating_task=0
panic_on_oom=0
unprivileged_userfaultfd=0
Your values may differ. The installed sysctl command offers the same read operation with dotted names:
$ sysctl vm.swappiness vm.overcommit_memory vm.overcommit_ratio vm.overcommit_kbytes
vm.swappiness = 10
vm.overcommit_memory = 0
vm.overcommit_ratio = 50
vm.overcommit_kbytes = 0
Checkpoint: save the values you intend to change. A setting can be valid and still be wrong for a particular workload, so record the baseline before experimenting.
2. Understand the allocation policy before changing it
vm.overcommit_memory selects the kernel's virtual-memory accounting mode. The default is 0, heuristic overcommit. Mode 1 always overcommits, while mode 2 checks allocations against a commit limit. Mode 0 can allow a later out-of-memory kill, and mode 1 is not a general performance switch: it tells the kernel to accept allocations until memory actually runs out.
Mode 2 calculates the commit limit from RAM, memory reserved for huge pages, swap and either vm.overcommit_ratio or vm.overcommit_kbytes. A non-zero overcommit_kbytes takes precedence and writing either limit setting clears the other. The formula is therefore part of the change review, not a detail to guess after the host starts rejecting allocations.
Do not change vm.panic_on_oom casually. Its values can select OOM killing or a kernel panic, and values 1 and 2 are intended for clustering failover policies. Likewise, vm.oom_kill_allocating_task changes which process the OOM killer selects, while vm.oom_dump_tasks can add a task dump when the kernel kills a memory-hogging process.
3. Test a modest swappiness change
vm.swappiness controls how aggressively the kernel swaps memory pages. Higher values increase aggressiveness and lower values decrease it. The manpage documents a default of 60; this machine currently uses 10. That difference is a host policy choice, not proof that 10 is universally better.
Changing a sysctl is live and normally lasts until reboot unless a system configuration manager reapplies it. First capture the original value, then make a small test change:
$ old_swappiness=$(cat /proc/sys/vm/swappiness)
$ printf 'baseline swappiness: %s\n' "$old_swappiness"
baseline swappiness: 10
$ sudo sysctl -w vm.swappiness=20
vm.swappiness = 20
$ sysctl -n vm.swappiness
20
Use a value you have chosen for this host, not the example blindly. Observe swap activity, memory pressure and application latency with your normal monitoring. The command only changes the kernel's current value; it does not rewrite a persistent configuration file.
4. Restore a test change
Restore the captured value when the test is complete or its effect is unhelpful:
$ sudo sysctl -w "vm.swappiness=$old_swappiness"
vm.swappiness = 10
$ test "$(sysctl -n vm.swappiness)" = "$old_swappiness" && echo restored
restored
The shell variable exists only in the shell where it was assigned. If you opened a new shell, use the value you wrote down instead. A reboot also removes a transient write, but do not reboot a production system merely to undo a tunable.
5. Drop caches only for a specific test
vm.drop_caches is an action trigger, not a cache-size setting. Writing 1 asks the kernel to drop clean page cache, 2 asks it to drop clean dentries and inodes, and 3 asks for both. This discards useful cached data and can make the system slower. It does not free dirty objects, and it is not a routine memory-cleaning command.
Before a reproducible filesystem benchmark, synchronise pending writes, then run the narrowly chosen action with elevated privilege:
$ sync
$ sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'
$ grep -E '^(Cached|SReclaimable|MemFree):' /proc/meminfo
The write is nondestructive in the sense described by proc_sys_vm(5): it does not delete files or alter their contents. It does remove cache state, so the performance cost is real. There is no value to restore afterwards; the normal recovery is simply to let the workload warm the caches again. Never use this as a substitute for finding a process that is consuming anonymous memory.
6. Treat reservations and security flags as policy
When strict overcommit mode is enabled, vm.admin_reserve_kbytes and vm.user_reserve_kbytes help leave room to log in, inspect the host and terminate a memory hog. Setting the user reservation to zero can leave a user able to consume almost all free memory, after which even starting a recovery command may fail. On a 64-bit system using mode 2, the manpage gives 131072 KiB as an example administrative reserve, but choose a value based on the actual recovery tools you need.
vm.unprivileged_userfaultfd controls whether unprivileged processes may use userfaultfd(2). The documented default is 1, but this machine reports 0, which means the capability CAP_SYS_PTRACE is required. Do not flip this setting to make an application work without reviewing the security consequence and the application's need.
Some files are conditional. Memory-failure controls appear only with CONFIG_MEMORY_FAILURE, compaction requires CONFIG_COMPACTION, and hardware support also affects machine-check recovery. A missing file is therefore not automatically a permissions error or a broken sysctl command.
Done means
- You recorded the current values before changing a virtual-memory control.
- You can explain the difference between heuristic, always and strict overcommit.
- Any swappiness experiment was observed and restored, or deliberately documented as a live policy.
- You used
syncand dropped caches only for a defined benchmark or memory-management test. - You left OOM panic, reservation and userfaultfd settings unchanged unless a reviewed host policy requires them.