Home / Alt manpages / oqmgr(8postfix)

  • oqmgr(8postfix)
  • Postfix admin command
  • linux

Tune and Troubleshoot the Postfix Queue Manager

You will finish with a safe way to inspect oqmgr's queues, relate a backlog to the queue manager's controls, and apply a small configuration change without guessing what the daemon does. The examples match Postfix 3.8.6 from the Ubuntu package installed here, but the queue model is documented by the installed oqmgr(8postfix) manual.

Allow about twenty minutes. You need a shell account that can read Postfix configuration and logs. Changing main.cf requires root or equivalent privilege and can affect mail delivery, so use a maintenance window if you change limits or retry timing. Most inspection commands below are unprivileged.

1. Confirm which queue manager you have

oqmgr is not normally a command that you start by hand. It is a Postfix daemon started by the master process manager. Starting a second copy manually can break the service's assumptions about sockets, privileges and queue access.

Check the package and the installed defaults first:

$ postconf mail_version queue_directory config_directory
mail_version = 3.8.6
queue_directory = /var/spool/postfix
config_directory = /etc/postfix
$ dpkg-query -W -f='${Package} ${Version}\n' postfix
postfix 3.8.6-1ubuntu0.1

The executable is usually beneath Postfix's private binary directory rather than on the ordinary command path:

$ ls -l /usr/lib/postfix/sbin/oqmgr
-rwxr-xr-x 1 root root ... /usr/lib/postfix/sbin/oqmgr

Checkpoint: if mail_version or the queue directory differs, keep the values returned by your host. Do not paste the paths above into a different installation.

2. Read the queues as separate states

The queue directory is not one undifferentiated pile of mail. oqmgr uses incoming for newly accepted messages, active for messages currently opened for delivery, and deferred for messages whose first delivery attempt failed temporarily. hold contains messages deliberately kept from delivery, while unreadable queue files are moved to corrupt for inspection.

Use the Postfix queue display rather than editing queue files:

$ postqueue -p
-Queue ID-  --Size-- ----Arrival Time---- -Sender/Recipient-------
... output depends on the host ...

An empty queue normally produces a message such as Mail queue is empty. Queue IDs, recipients and errors are host-specific. The command is a snapshot, so run it twice a few minutes apart when deciding whether a backlog is growing.

For a coarse view of the on-disk directories, read their entries without changing them:

$ sudo find /var/spool/postfix/incoming /var/spool/postfix/active /var/spool/postfix/deferred \
    -type f -printf '%h\n' 2>/dev/null | sort | uniq -c

This uses sudo only because queue files are normally protected. Do not open, rename or delete individual files to clear a backlog. Postfix maintains related queue and status files, and manual removal can lose mail or leave misleading status.

3. Match the symptom to oqmgr's strategy

oqmgr limits how many messages enter active. This leaky-bucket behaviour protects the queue manager's memory, but it means a busy system can show a large deferred queue even while delivery is working normally. It also takes one message from incoming and one from deferred when active space is available, so new mail is not supposed to be permanently starved by an older backlog.

Delivery to one destination is controlled separately. oqmgr starts with an initial concurrency, adjusts it from connection and handshake results, and applies a per-destination maximum. A destination that repeatedly fails can be treated as unavailable for a time. Round-robin selection prevents one destination from consuming every delivery opportunity.

Deferred delivery also has exponential backoff. The default minimum delay is 300 seconds, the maximum is 4000 seconds, and deferred scans are normally 300 seconds apart on this installation. These are delays between attempts, not a promise that every message is retried at exactly those times.

Checkpoint: inspect the effective values before changing anything:

$ postconf qmgr_message_active_limit qmgr_message_recipient_limit \
    minimal_backoff_time maximal_backoff_time queue_run_delay \
    default_destination_concurrency_limit initial_destination_concurrency
qmgr_message_active_limit = 20000
qmgr_message_recipient_limit = 20000
minimal_backoff_time = 300s
maximal_backoff_time = 4000s
queue_run_delay = 300s
default_destination_concurrency_limit = 20
initial_destination_concurrency = 5

postconf shows the effective value, including defaults and transport-specific expansion. Use postconf -n when you need to see only settings explicitly present in main.cf.

4. Change one control and keep an undo value

Only change a setting when logs and queue measurements show a specific problem. For example, increasing concurrency can increase pressure on a remote destination and may make connection failures worse. Lowering retry delays can create repeated traffic to a failing service. Treat both as operational changes, not generic performance improvements.

Save the current setting, make one small change, then inspect the result:

$ old_active_limit=$(postconf -h qmgr_message_active_limit)
$ printf 'old active limit: %s\n' "$old_active_limit"
old active limit: 20000
$ sudo postconf -e 'qmgr_message_active_limit = 10000'
$ postconf -h qmgr_message_active_limit
10000

The command above changes persistent configuration but does not reload the running queue manager. Keep the recorded value until the change has been assessed.

Undo it by restoring the captured value, then reload:

$ sudo postconf -e "qmgr_message_active_limit = $old_active_limit"
$ sudo postfix reload
$ postconf -h qmgr_message_active_limit
20000

Do not put an untrusted value into the quoted postconf -e command. If you edited main.cf by hand instead, restore the previous line and use the same reload step. A reload asks Postfix to re-read configuration; it does not delete queued mail.

5. Force a scan only when you mean to

Normally oqmgr wakes on timers and trigger events. A queue flush request, such as postqueue -f, asks Postfix to attempt delivery sooner. It does not repair a rejected recipient, remove a deferred message or bypass a remote server's temporary failure.

Use it after correcting a known temporary cause, such as restoring a required transport, and expect it to create delivery work:

$ sudo postqueue -f
$ printf 'flush request status: %s\n' "$?"
flush request status: 0

Do not repeatedly flush a large queue as a substitute for diagnosis. Repeated requests can add load while DNS, network, authentication or remote policy is still failing. Check the queue and mail logs after one request.

6. Read logs and separate failures

oqmgr reports transactions and problems through the system logging service, such as syslogd or postlogd. Search the mail log for a queue ID from postqueue -p, then follow that ID across accepted, deferred and delivered records:

$ sudo journalctl -u postfix --since '15 minutes ago' --no-pager
$ sudo journalctl -u postfix --since '15 minutes ago' --no-pager | grep 'QUEUE_ID'

Replace QUEUE_ID with an actual ID and quote it if you store it in a shell variable. A deferred status usually identifies the failing destination and the temporary reason. A corrupt queue file is a different class of problem and belongs in the corrupt queue for inspection, not in a retry loop.

When a destination is slow, check the destination-specific concurrency and rate-delay settings before raising global limits. When new mail is delayed during a burst, check disk contention and the active queue limit as well as network delivery. The manual records a known trade-off: a single queue manager competes for disk access with front-end processes such as cleanup, so inbound bursts can reduce outbound delivery rates.

Done means

  • You confirmed the installed Postfix version and queue directory.
  • You can distinguish incoming, active, deferred, hold and corrupt queue states.
  • You inspected effective oqmgr controls before changing them.
  • Any configuration change has a recorded previous value and a tested undo command.
  • You reload after changing main.cf, and flush only for a specific reason.
  • You use queue IDs and logs to identify the failing destination instead of deleting queue files.