Home / Alt manpages / galera_recovery(1)

  • galera_recovery(1)
  • User command
  • linux

Recover MariaDB's Galera Position with galera_recovery

You will learn what the installed galera_recovery helper does, how MariaDB's systemd unit uses its output, and how to inspect a failed recovery without inventing a cluster position. The examples match MariaDB Server 10.11.14 from the mariadb-server package on this machine. The installed manual page is dated 15 May 2020 and identifies the interface as MariaDB 10.11.

Allow about fifteen minutes for inspection. You need shell access and permission to read the MariaDB unit and its logs. The normal checks are read-only. Starting or restarting MariaDB is service-disrupting and may affect replication, so this guide does not ask you to do that blindly.

1. Confirm the helper and package version

Start with ordinary, read-only checks. These do not need sudo if the files are readable on your host:

$ command -v galera_recovery
/usr/bin/galera_recovery
$ dpkg-query -W -f='${Package} ${Version}\n' mariadb-server
mariadb-server 1:10.11.14-0ubuntu0.24.04.1
$ man galera_recovery

The manual page is deliberately short. It describes the purpose as recovery after a non-graceful shutdown and points to the MariaDB Knowledge Base. The executable is a shell helper installed by the server package, so its real contract is also visible in the local systemd unit.

Checkpoint

Do not substitute a similarly named script from another MariaDB installation. Keep /usr/bin/galera_recovery and the server package version together when diagnosing a host.

2. Understand what the helper returns

When Galera is enabled, the helper starts mariadbd with --wsrep_recover. It reads the temporary recovery log, extracts the last recovered global transaction position, and writes one option to standard output in this form:

--wsrep_start_position=RECOVERED_POSITION

That output is not a report for a person to copy into a SQL session. It is an argument intended for the next mariadbd start. Diagnostic messages go to standard error so standard output remains suitable for systemd to capture.

If the server is configured without wsrep, the helper skips position recovery and writes an empty line. An empty result therefore does not mean that a new Galera position was found. It means the helper did not need to provide one for this start.

3. Inspect the systemd hand-off

Read the unit that consumes the output. This is still an inspection step and does not restart the database:

$ systemctl cat mariadb.service | sed -n '/galera_recovery/,+12p'
ExecStartPre=/bin/sh -c "[ ! -e /usr/bin/galera_recovery ] && VAR= || \\
 VAR=`/usr/bin/galera_recovery`; [ $? -eq 0 ] \\
 && systemctl set-environment _WSREP_START_POSITION=$VAR || exit 1"
$ systemctl show mariadb.service -p ExecStart
ExecStart={ path=/usr/sbin/mariadbd ; ... $_WSREP_START_POSITION ; ... }

The unit first clears the old _WSREP_START_POSITION environment value. It then runs the helper, stores its standard output in that environment value, and passes the value to mariadbd. After the service starts, a post-start command clears the environment value again. This prevents a recovered position from being silently reused by a later, unrelated start.

The unit also treats a missing helper as an empty value. Do not read that fallback as proof that recovery happened. It only prevents a missing optional executable from blocking the unit at that particular check.

Checkpoint

Verify that the unit calls the same absolute path shown by command -v. If a local override replaces the pre-start command, inspect that override before trusting any recovery conclusion:

$ systemctl cat mariadb.service
$ systemctl show mariadb.service -p DropInPaths

4. Run a safe configuration-path check

Do not run galera_recovery on a production data directory merely to see what it prints. Without a controlled test, the helper may invoke mariadbd --wsrep_recover against the configured database files. That is part of the recovery operation, not a harmless version query.

You can check the non-wsrep branch without starting the server by supplying the script's documented skip setting. On this installed script, that setting is parsed while the script discovers MariaDB defaults:

$ galera_recovery --skip-wsrep-on >/tmp/galera-recovery.out 2>/tmp/galera-recovery.err
$ printf 'status=%s\n' "$?"
status=0
$ wc -c /tmp/galera-recovery.out
1
$ cat /tmp/galera-recovery.out

The one byte is the terminating newline for an empty result. A warning from my_print_defaults about an option that is not part of your local defaults is a configuration warning, not a recovered position. Remove these temporary files after inspection if they contain no information you need:

$ rm -f /tmp/galera-recovery.out /tmp/galera-recovery.err

This check does not prove that a real Galera recovery will succeed. It only confirms the installed helper's empty-output path and its exit status in the current environment.

5. Read a failed service start in the right order

If MariaDB failed to start after an unclean shutdown, first capture the unit result and recent journal entries. These commands require no elevated privilege when your journal policy permits access; use sudo only if the host denies the read:

$ systemctl status mariadb.service --no-pager
$ journalctl -u mariadb.service -b --no-pager -n 120

Look for the helper's messages, which begin with WSREP:. A successful position recovery reports the recovered position on standard error and emits the corresponding --wsrep_start_position=... option on standard output. A failed server recovery reports that it could not start mysqld for wsrep recovery. A missing recovered position is reported separately from the non-wsrep skip case.

Do not convert every failure into a new cluster bootstrap. Starting a new primary component can create a split-brain or discard the correct cluster history. Check the other members, the intended primary component, and your MariaDB or Galera recovery runbook before using any cluster-start operation.

If the journal is too brief, inspect the MariaDB error log location configured by your [mysqld] settings. Do not assume a fixed path: the helper creates its own temporary recovery log under /tmp, removes it on exit, and only includes its contents in an error message when recovery fails.

6. Restart only after the evidence is understood

A restart changes service state and may interrupt clients. Before doing it, confirm that the configured data directory is the intended one, that the mysql service account can access it, and that any cluster operator knows the node is being brought back. Then use the normal service manager command:

$ sudo systemctl restart mariadb.service
$ systemctl is-active mariadb.service
active
$ systemctl show mariadb.service -p Environment

The environment output should not retain a stale _WSREP_START_POSITION after startup because the unit clears it in ExecStartPost. If the restart fails, preserve the evidence: run systemctl status and journalctl again before attempting another restart. Repeated restarts can obscure the first useful error and prolong an outage.

If you created a systemd drop-in while testing, undo that specific change before retrying the service. Do not edit /usr/lib/systemd/system/mariadb.service; package upgrades can replace it. A drop-in can be removed with sudo rm -- /etc/systemd/system/mariadb.service.d/NAME.conf, followed by sudo systemctl daemon-reload. Only remove a file you created for this incident, after checking its exact path.

Done means

  • galera_recovery resolves to the installed MariaDB server helper and its package version is recorded.
  • You understand that recovered position text is a systemd-to-mariadbd argument, not a SQL value.
  • The unit clears stale state, captures standard output, and clears the environment after startup.
  • An empty output is treated as the non-wsrep path, not as proof of a recovered cluster position.
  • Failures are checked in the service status and journal before any disruptive retry or cluster bootstrap decision.