Home / Alt manpages / proc_sys_fs(5)

  • proc_sys_fs(5)
  • File format
  • linux

Inspect and Safely Tune /proc/sys/fs on Linux

You will finish with a small, repeatable workflow for reading the filesystem-related kernel settings under /proc/sys/fs, deciding which value matters, and changing one setting at runtime when there is a clear reason. The examples match the installed proc_sys_fs(5) from the Debian manpages package, version 6.7-2, on a Linux 6.8 host.

Allow about fifteen minutes. You need a shell and read access to /proc/sys/fs. Reading is normally unprivileged. Writing these files changes live kernel behaviour and normally requires root, so keep a recovery value before changing anything. This guide does not make a persistent boot or service configuration change.

1. Confirm the kernel and the local contract

Start by checking the kernel and the manpage package. These values are ordinary read-only information:

$ uname -r
6.8.0-139-generic
$ dpkg-query -W -f='${Package} ${Version}\n' manpages
manpages 6.7-2
$ man 5 proc_sys_fs

The installed manual is the contract for this machine's documentation set. Kernel versions can expose different files, and some entries are directories rather than single values. Do not assume that every name in an article or a script exists on another host.

Checkpoint: verify that the directory is mounted and list the entries without changing them:

$ test -d /proc/sys/fs && echo '/proc/sys/fs is available'
/proc/sys/fs is available
$ find /proc/sys/fs -maxdepth 1 -mindepth 1 -printf '%f\n' | sort

A missing directory usually means that procfs is not mounted or that you are in a restricted environment such as a container. Fix that environment issue before designing a tuning change. Do not create replacement files under /proc.

2. Read the values that explain capacity

Several files describe system-wide or namespace-wide ceilings. Read them together so that one number is not mistaken for another:

$ for name in file-max file-nr nr_open mount-max; do
>     printf '%s: ' "$name"
>     cat "/proc/sys/fs/$name"
> done
file-max: 9223372036854775807
file-nr: 13856 0 9223372036854775807
nr_open: 1048576
mount-max: 100000

Your output will differ. file-max is the system-wide limit on open file descriptions. file-nr reports allocated handles, free handles and the same maximum. On Linux 2.6 and later, the middle value is normally zero because freed handles are released rather than kept on an allocation list. A high historical free count is not, by itself, evidence that the current limit is being exhausted.

nr_open is the ceiling for a process's RLIMIT_NOFILE, while a process can still have a lower limit. Check a process's actual soft and hard limits separately:

$ ulimit -Sn
1024
$ ulimit -Hn
1048576

mount-max limits the number of mounts in one mount namespace. It is not a count of mounted filesystems across every namespace on the host. If a workload reports too many open files or mounts, collect the relevant error and current values first. Raising a limit without measuring the workload can hide a leak and allow it to consume more kernel resources.

3. Inspect the security switches before touching them

The protected_* files control safeguards around common filesystem races. Read their current values and keep the output as your baseline:

$ for name in protected_hardlinks protected_symlinks protected_regular protected_fifos; do
>     printf '%s: ' "$name"
>     cat "/proc/sys/fs/$name"
> done
protected_hardlinks: 1
protected_symlinks: 1
protected_regular: 2
protected_fifos: 1

For protected_hardlinks and protected_symlinks, zero removes the restriction and one enables it. The hard-link protection limits links to files the caller owns or can suitably access. The symlink protection restricts following links in sticky world-writable directories unless the ownership conditions are safe.

protected_regular and protected_fifos accept zero, one or two. Zero is unrestricted. One protects creation attempts in world-writable sticky directories. Two applies the same kind of restriction to group-writable sticky directories as well. These settings are intended to stop a program expecting to create its own file from being redirected to an attacker-controlled regular file or FIFO.

Disabling these protections is security-sensitive and can reintroduce old time-of-check, time-of-use attacks. Do not set a value to zero merely to make a legacy script pass. Find the exact failing operation, check its ownership and directory permissions, and change the application or deployment if possible.

4. Choose a reversible runtime change

Only make a change when you have a measured capacity problem, a documented compatibility need, and a rollback value. This example raises file-max temporarily. It is an example of the write shape, not a recommended number:

$ old_file_max=$(cat /proc/sys/fs/file-max)
$ printf 'current file-max: %s\n' "$old_file_max"
current file-max: 9223372036854775807
$ sudo sh -c 'printf "%s\n" 100000 > /proc/sys/fs/file-max'
$ cat /proc/sys/fs/file-max
100000

The shell running the redirection must have permission to open the proc file. That is why sudo echo 100000 > ... is a trap: sudo elevates echo, but the calling shell performs the redirection first. Use sudo sh -c or an equivalent privileged tool, and review the exact value before pressing Enter.

Warning: lowering a live limit can disrupt unrelated processes. A value that is too low can cause file-opening operations to fail with ENFILE. Never use a guessed low value on a production host, and do not change a security switch during an incident without recording who approved the risk.

5. Verify and undo the change

Check the file immediately and compare it with the value you intended. A successful write proves only that the kernel accepted the value:

$ test "$(cat /proc/sys/fs/file-max)" = 100000 && echo 'file-max changed'
file-max changed
$ cat /proc/sys/fs/file-nr
13856 0 100000

Watch the application metric or error that justified the change. For file-max, inspect open-file pressure and service logs rather than treating the new number as proof that the underlying issue is fixed. For a security switch, test the intended operation as the same user and from the same directory context that previously failed.

Restore the saved value when the test is complete:

$ sudo sh -c 'printf "%s\n" 9223372036854775807 > /proc/sys/fs/file-max'
$ cat /proc/sys/fs/file-max
9223372036854775807

Use your captured old_file_max value, not the host-specific value shown here. If the shell variable still exists, this avoids retyping it:

$ sudo sh -c 'printf "%s\n" "$1" > /proc/sys/fs/file-max' sh "$old_file_max"

These writes are runtime-only. A reboot, or a later sysctl management action, may restore a distribution default or another configured value. If a change must persist, handle that as a separate change: document the reason, review the security and resource impact, and use your system's normal sysctl configuration process. Do not assume that writing /proc/sys/fs/file-max creates persistent policy.

6. Check the less obvious entries when needed

The same directory also contains counters and controls for asynchronous I/O, directory caching, inotify, epoll, pipes, POSIX message queues, file leases, quotas and core dumps. Read the relevant manual section before changing one. For example, aio-nr is a running total, while aio-max-nr is its limit; increasing the limit does not preallocate kernel structures.

suid_dumpable deserves extra care. Mode zero is the default, mode one can expose privileged process memory to unprivileged readers through core dumps, and mode two limits protected dumps to root. Core dump behaviour also depends on /proc/sys/kernel/core_pattern. Treat a request to enable broad core dumping as a security and data-protection change, not as a harmless diagnostic toggle.

If an entry is a directory, such as inotify or mqueue, inspect its children explicitly:

$ find /proc/sys/fs/inotify -maxdepth 1 -type f -printf '%f\n' | sort
max_queued_events
max_user_instances
max_user_watches
$ cat /proc/sys/fs/inotify/max_user_watches

Keep the check read-only until you know which resource is constrained and which process consumes it. The kernel documentation and the installed proc_sys_fs(5) page are the right places to resolve version-specific meaning.

Done means

  • You confirmed the local kernel and manpages version.
  • You distinguished file-max, file-nr, nr_open and mount-max.
  • You recorded the current value before any privileged write.
  • You treated the protected-link, protected-file and core-dump switches as security controls.
  • You verified any runtime change and have an exact undo value.
  • You know that a procfs write is temporary until a separate persistence change is reviewed.