A process that keeps growing in RSS with no leak visible in the code is exactly what trace-cmd mem was built to catch. You will finish with a report that ranks allocator call sites by bytes allocated but not requested, and you will know which of its columns are running totals and which are peaks.
The installed command is trace-cmd 3.2.0 from package version 3.2-1ubuntu2. Allow 10 to 20 minutes for a short capture and review. You need the trace-cmd package, a writable directory for the trace file, and a workload that exercises the code you want to investigate.
Recording kernel events normally needs elevated privileges. The report itself is read-only, but the capture can affect tracing state and creates a large file.
Confirm the installed version and choose an output path before you start. Keep the trace outside a shared directory if it may contain sensitive process names or kernel activity.
$ trace-cmd --version
trace-cmd version 3.2.0 (not-a-git-repo)
$ mkdir -p /tmp/trace-cmd-mem-work
$ test -w /tmp/trace-cmd-mem-work && echo writable
writable
Checkpoint: if the version is not 3.2.0, keep the version in your notes. The report columns described here come from the installed trace-cmd-mem(1) page, and other releases can differ.
kmalloc, kmalloc_node, kfree, kmem_cache_alloc, kmem_cache_alloc_node, and kmem_cache_alloc_free.Run a deliberately short capture around a command you can repeat. This example records to a new file and traces a five-second sleep:
$ sudo trace-cmd record -e kmem -o /tmp/trace-cmd-mem-work/trace.dat -- sleep 5
... trace-cmd recording messages ...
The command creates one trace file after the recording finishes. Do not add a long-running service to this first test: kernel tracing consumes buffer space and disk space, and a broad allocation capture can be noisy.
If the event group is unavailable, inspect the kernel's event list without changing it:
$ sudo trace-cmd list -e kmem
Checkpoint: verify that the output file exists and is non-empty:
$ test -s /tmp/trace-cmd-mem-work/trace.dat && echo trace-ready
trace-ready
If recording fails with a permission error, use the privilege required by your host's tracing policy. If it fails because an event is missing, do not pretend the resulting report is complete: investigate the available kernel events first.
Run mem against the captured file. Supply the input file with -i, as shown here:
$ trace-cmd mem -i /tmp/trace-cmd-mem-work/trace.dat
Function Waste Alloc req TotAlloc TotReq MaxAlloc MaxReq MaxWaste
-------- ----- ----- --- -------- ------ --------- ------ --------
example_function 768 2304 1536 2304 1536 2304 1536 768
...
The actual functions and numbers depend entirely on the workload. The list is sorted by descending Waste, so start at the top. Waste is allocated bytes minus requested bytes that were not freed. A zero waste value is still useful: it can identify an allocator path whose recorded allocation and request sizes match.
Tip: without -i, the command reads trace.dat in the current directory. That default is easy to miss when the capture was written elsewhere, so use an explicit path in scripts and investigation notes.
The report has three kinds of measurement, and mixing them up is the easiest way to misread a leak:
Alloc and req are the final amounts still represented after the run.TotAlloc and TotReq are the totals seen across the run.MaxAlloc and MaxReq are the highest amounts observed.MaxWaste is the largest waste value observed, while Waste is the value at the end.Example: a row with a high TotAlloc but zero Alloc describes repeated activity that was eventually freed, not a current leak. A high final Waste points at memory still unfreed in the trace's accounting, but it is not by itself proof of a permanent leak: the capture may have stopped before cleanup. Repeat the workload and extend the capture only after the short run is understandable.
Keep the original trace.dat until you have recorded the report and checked the workload boundaries. It is a binary trace, not a configuration file. Compressing a copy is safe if storage is tight:
$ cp --preserve=all /tmp/trace-cmd-mem-work/trace.dat /tmp/trace-cmd-mem-work/trace.dat.copy
$ gzip /tmp/trace-cmd-mem-work/trace.dat.copy
$ trace-cmd mem -i /tmp/trace-cmd-mem-work/trace.dat
Warning: the capture command normally stops tracing when its workload exits. If you interrupted or separately started tracing, check the tracing state before changing it. Do not run broad reset commands merely to tidy up: a reset can disrupt another person's investigation or service diagnostics.
trace-cmd mem -i completes against a non-empty trace file.