Collect Cluster sos Reports Safely with sos collect
You will finish with a repeatable command for collecting sos reports from selected cluster nodes into one local archive. This guide uses sos 4.10.2 from the installed sosreport package. Allow 15 to 30 minutes for preparation and collection, plus the time needed for each remote report.
The route
Jump straight to the step you need, or tick off Done means at the end.
The current interface is the sos collect subcommand. The installed manpages still have the names sos-collect(1) and sos-collector(1), but both document the same interface and say that the standalone sos-collector command is deprecated and will be removed in sos 4.9. Use sos collect in new notes and scripts.
1. Check the installed interface
Run these read-only checks as your ordinary account:
$ command -v sos
/usr/bin/sos
$ sos collect --help
usage: sos collect [options]
$ sos --version
sos 4.10.2
The exact help text can vary between package builds. The useful checkpoint is that the command exists and identifies itself as sos 4.10.2. If your version is older, read its local help before copying the examples. The installed manual describes the same command as sos collect, not as a separate executable that must be present in $PATH.
2. Confirm access before collecting
sos collect connects to the nodes and runs sos report there. The manual assumes SSH key authentication, unless you explicitly choose a password prompt. Before starting a collection, check that the account and port work for every target:
$ ssh -p 22 [email protected] 'hostname'
node-a.example.test
$ ssh -p 22 [email protected] 'hostname'
node-b.example.test
Replace the example names with real hosts. The default remote user is root. If you use a non-root account, pass --ssh-user USER and normally also --become so the remote collection can obtain the required privileges. The manual says that sos collect prompts for a sudo password for a non-root user. With passwordless sudo, --nopasswd-sudo tells it not to expect one.
Checkpoint: do not begin until SSH reaches the intended machines and the account has the access your support process permits. Collection archives can contain logs, configuration and other sensitive host data.
3. Select nodes explicitly for a simple cluster
For a cluster profile that does not need automatic enumeration, set the cluster type to none and give a comma-separated node list. This makes the target set visible in the command:
$ sos collect \
--cluster-type=none \
--nodes=node-a.example.test,node-b.example.test,node-c.example.test \
--ssh-user=SUPPORT_USER \
--become \
--label=incident-2026-09-27
With --cluster-type=none, cluster-specific checks are disabled and the nodes supplied with --nodes are used. This is suitable for a deliberately chosen list. A node expression passed to --nodes is a whitelist for discovered names, not a blacklist: it selects matching nodes and cannot be used to exclude arbitrary names from an already selected set.
The command writes a combined archive in a temporary directory. The default temporary directory is created under /tmp and removed when sos collect finishes. The final archive is reported by the command. Do not assume that a quiet shell prompt means that no data was collected.
4. Control concurrency and collection time
The default is four concurrent node collections and four report threads per node. Those defaults are reasonable for a small test, but they can put pressure on a busy cluster. Lower them when the incident host or network is already constrained:
$ sos collect \
--cluster-type=none \
--nodes=node-a.example.test,node-b.example.test \
--jobs=2 \
--threads=2 \
--timeout=300 \
--batch
--jobs limits how many nodes are collected concurrently. --threads controls concurrent plugin work in each remote report. --timeout is the per-node report-generation timeout in seconds. The manual's rough runtime estimate is timeout multiplied by the number of nodes divided by the concurrency, so allow extra time for SSH setup and retries. --batch skips prompts; use it only when authentication and all other required input are already non-interactive.
5. Keep the archive bounded
sos report limits most collected files and command output to 25 MiB by default, with journal collection limited to 100 MiB. Keep that default unless the investigation needs more data. Setting --log-size=0 removes those limits and can greatly increase archive size and memory use, particularly for large journals.
$ sos collect \
--cluster-type=none \
--nodes=node-a.example.test,node-b.example.test \
--log-size=50 \
--skip-files='/var/log/example/*'
--log-size is measured in MiB. --skip-files accepts comma-separated paths or shell-style wildcard matches. Use it only when you understand which evidence is being omitted. You can similarly use --skip-commands for commands that hang, or --skip-plugins for a troublesome plugin. Skipping data changes the diagnostic value of the result, so record those exclusions with the case.
6. Handle sensitive output deliberately
Inspect the archive before sending it to a third party. It can contain credentials in configuration files, hostnames, addresses, logs and application data. Restrict access to the directory holding the result, and transfer it only through your approved support channel.
If policy requires encryption, use a GPG recipient already present in the keyring of the account running sos collect:
$ sos collect \
--cluster-type=none \
--nodes=node-a.example.test,node-b.example.test \
--encrypt-key=RECIPIENT_KEY_ID
This encrypts the final archive on the machine where sos collect runs. The individual reports are collected unencrypted on the nodes. The manual also warns that encryption temporarily needs about twice the archive space because the encrypted file is written separately. If encryption fails, the original unencrypted archive is preserved, so treat that failure as a security checkpoint: do not upload the preserved file until you have resolved the problem or applied your organisation's handling rules.
--encrypt-pass uses a supplied passphrase instead of a public key, but putting secrets directly in a shell command can expose them through shell history or process inspection. Prefer the supported interactive or environment-based secret handling for your deployment, and never paste a real password into a shared terminal transcript.
7. Verify success and recover from a failed run
At the end, check the exit status and locate the reported archive:
$ printf 'exit status: %s\n' "$?"
exit status: 0
$ file /path/to/reported-archive.tar.*
$ tar -tf /path/to/reported-archive.tar.* | head
A zero status means the command completed successfully. It does not prove that every node contributed a report. Read the command output for connection, version-check and per-node errors, then check the archive contents. A run that says no reports were collected is not a usable result, even if the shell command returned to the prompt.
If one node fails, test it separately with SSH, correct its name, port, credentials or remote sos installation, and rerun with the same node list. Keep the failed archive and logs until the incident record is complete. sos collect does not provide an undo operation: it gathers data and normally removes its temporary directory after completion, but it does not change cluster configuration. You can safely remove an unwanted local archive only after confirming that it is no longer required by your retention or incident process.
Done means
- The installed command is
sos collect, and its version was checked. - SSH and the selected remote account were tested against every intended node.
- The node list, concurrency and timeout match the incident's limits.
- Any skipped files, commands or plugins are recorded and understood.
- The archive was checked for contributing nodes, protected as sensitive data and encrypted when policy requires it.