Home / Alt manpages / pg_resetwal(1)

  • pg_resetwal(1)
  • User command
  • linux

Recover a Corrupt PostgreSQL Cluster with pg_resetwal

You will use PostgreSQL 16.15's pg_resetwal to rebuild a cluster's WAL and control information only when corruption prevents the server starting. The recovery ends with a logical dump, a fresh initdb cluster and a restore. It is not a routine repair command.

Allow at least 30 minutes for the command work, plus however long the dump and restore take. You need the PostgreSQL 16 utilities, the cluster's exact data directory, read and write access as the user that installed the server, and enough storage for a dump and replacement cluster. Take a filesystem or storage snapshot first if your recovery procedure supports one.

1. Confirm the binary and data directory

Check the installed version and locate the executable. On this host the command is installed under the PostgreSQL 16 binary directory:

$ /usr/lib/postgresql/16/bin/pg_resetwal --version
pg_resetwal (PostgreSQL) 16.15 (Ubuntu 16.15-0ubuntu0.24.04.1)
$ command -v /usr/lib/postgresql/16/bin/pg_resetwal
/usr/lib/postgresql/16/bin/pg_resetwal

Replace /var/lib/postgresql/16/main below with the real cluster directory. Do not rely on PGDATA: pg_resetwal requires the directory on its command line. Check that it is a PostgreSQL data directory before doing anything destructive:

$ test -f /var/lib/postgresql/16/main/PG_VERSION && cat /var/lib/postgresql/16/main/PG_VERSION
16
$ test -f /var/lib/postgresql/16/main/global/pg_control && echo 'pg_control present'
pg_control present

If either check fails, stop. A typo here can make you inspect or modify the wrong directory.

2. Make certain the server is stopped

Warning

Never run pg_resetwal against a running server. Stop the relevant service using the method your system administrator normally uses, then verify that no PostgreSQL server process remains:

$ sudo systemctl stop postgresql@16-main
$ pgrep -a -u postgres '[p]ostgres' || echo 'no postgres processes found'
no postgres processes found

The stop command needs elevated privileges on a normal service installation. The process check does not. If the service has a different unit name, identify that unit first. A stale postmaster.pid can remain after a crash, but do not remove it until you have proved that no server process is alive. Removing the lock file while a server is running risks data corruption.

3. Preview the proposed reset

Run a dry run as the PostgreSQL cluster owner. It prints reconstructed values and exits without changing the cluster:

$ sudo -u postgres /usr/lib/postgresql/16/bin/pg_resetwal \
    --dry-run --pgdata=/var/lib/postgresql/16/main
Current pg_control values:

The exact report depends on the damaged cluster. The meaningful checks are that the data directory is the one you intended, the command can read it, and the dry run exits successfully. Capture the report in your incident record. An unsuccessful dry run is a reason to investigate the backup, permissions or corruption, not to add --force automatically.

Checkpoint: verify that the directory is still unchanged:

$ sudo -u postgres /usr/lib/postgresql/16/bin/pg_resetwal \
    --dry-run --pgdata=/var/lib/postgresql/16/main > /tmp/pg-resetwal-preview.txt
$ test -s /tmp/pg-resetwal-preview.txt && echo 'preview captured'
preview captured

4. Apply the reset only as a last resort

Warning

This changes the cluster's WAL and control information and can leave partially committed transactions inconsistent. Preserve your snapshot or a copy of the data directory before proceeding. Keep the output and the exact command.

$ sudo -u postgres /usr/lib/postgresql/16/bin/pg_resetwal \
    --pgdata=/var/lib/postgresql/16/main
Write-ahead log reset

Diagnostic wording varies by build. A zero exit status is the useful success signal. There is no undo option for an applied reset. Recovery is by restoring the original cluster from a known-good backup or snapshot, or by completing the dump and rebuild procedure below.

Use --force only when the utility says it cannot determine valid pg_control data and you have no safer route. It substitutes plausible values. If you force the reset, do not perform ordinary writes after the server starts. The manual values that may need expert review include the next OID, transaction ID and epoch, multitransaction ID and offset, and next WAL file.

5. Start only to extract the data

Start the service and inspect its logs. If it starts, treat the cluster as damaged even if clients can connect:

$ sudo systemctl start postgresql@16-main
$ sudo -u postgres pg_dumpall --file=/var/lib/postgresql/16/main-recovery.sql
$ test -s /var/lib/postgresql/16/main-recovery.sql && echo 'dump captured'
dump captured

Run the dump as the database owner and use a destination with enough free space. If the server will not start, stop retrying and return to a backup, snapshot or specialist recovery plan. Do not run application migrations, updates or other data-modifying SQL before the dump. Any apparent success can hide damage from the reset.

6. Rebuild and restore a clean cluster

Plan the replacement with your normal service and backup procedure. The broad sequence is to stop PostgreSQL, move the damaged directory to a clearly named quarantine location, initialise a new cluster with the same major version, start it, and restore the dump. Those operations alter service state and filesystem contents, so use your distribution's documented paths and a tested change window rather than copying this as a blind command.

For a plain PostgreSQL utility sequence, the important version boundary is that pg_resetwal works only with the same major server version. Do not use the PostgreSQL 16 binary on a PostgreSQL 15 or 17 data directory. After the restore, check application tables, sequences, constraints and replication state, then take a new backup before returning traffic.

7. Know the options that are easy to misuse

Numeric overrides accept hexadecimal values with a 0x prefix, but they are specialist controls, not guesses. --next-wal-file takes a WAL segment file name, not an LSN, and should normally be higher than every file in pg_wal. The utility normally chooses that value itself. --wal-segsize accepts a power of two from 1 to 1024 megabytes and may require an explicit WAL starting file to avoid filename overlap.

Do not confuse a missing archive segment with evidence that a larger WAL value is safe. If you need --epoch, --next-transaction-id, --oldest-transaction-id, --multixact-ids, --multixact-offset, --next-oid or --commit-timestamp-ids, derive it from the directory contents and recovery evidence, or involve a PostgreSQL specialist. The dry run does not validate a value you invent.

Done means

  • The PostgreSQL 16.15 binary and the exact data directory were confirmed.
  • No PostgreSQL server process was running when the reset was previewed or applied.
  • A dry-run report and pre-change backup or snapshot were preserved.
  • --force and manual control-value overrides were used only with evidence and review.
  • The recovered cluster was dumped before normal writes resumed.
  • A same-major-version clean cluster was rebuilt, restored, checked and backed up.