Safely Roll Back a Docker Swarm Service

The last deploy broke something, and docker service rollback is Swarm's built-in way to undo it. It reverts the most recent configuration change and waits for the tasks to converge. The examples use Docker CLI 29.8.1, installed here as package version 5:29.8.1-1~ubuntu.24.04~noble.

Allow about fifteen minutes for the command and its checks, plus however long your service needs to replace tasks. You need access to a Docker Swarm manager and a service that has had a configuration update. The rollback command is a cluster operation: it changes running workloads and may briefly disrupt traffic.

Checkpoint: stop here if you do not know which service should change, whether the previous configuration is safe, or how to restore your application if that configuration turns out to be bad too.

1. Confirm the command and your Swarm role

Check the installed command before touching the service. This is an ordinary, read-only command and normally needs no elevated privileges:

$ docker version --format '{{.Client.Version}}'
29.8.1
$ docker service rollback --help
Usage:  docker service rollback [OPTIONS] SERVICE

Revert changes to a service's configuration

The local command accepts only two options: --detach (or -d) to return immediately, and --quiet (or -q) to suppress progress output. The service name is a required final argument.

Rollback is a Swarm manager operation, so check the node state before continuing:

$ docker info --format '{{.Swarm.LocalNodeState}} {{.Swarm.ControlAvailable}}'
active true

The values must describe an active manager. On a worker, or a host with no active Swarm, move to a manager rather than trying sudo: elevated privileges do not turn a worker into a manager. If Docker access itself is denied, ask the system administrator to grant proper access instead of loosening socket permissions as a quick fix.

2. Identify the exact service

List services and choose the exact name. This remains read-only:

$ docker service ls
ID            NAME           MODE         REPLICAS   IMAGE
SERVICE_ID    SERVICE_NAME   replicated   3/3        IMAGE:TAG

The IDs, names and replica counts above are examples; use the values from your own cluster. Do not select by a partial name from memory. Set a shell variable after checking it:

$ SERVICE_NAME='SERVICE_NAME'
$ docker service inspect --pretty "$SERVICE_NAME"
ID:             SERVICE_ID
Name:           SERVICE_NAME
Service Mode:   Replicated
 Replicas:      3

Inspect output is host-specific. Review the image, published ports, replica count, constraints and update or rollback policy. Keep the service name quoted, especially when it comes from a script or another operator's input.

Checkpoint: you have one confirmed service name and a manager prompt. If the service is healthy and the rollback reason is unclear, stop and investigate rather than using rollback as a diagnostic button.

3. Record the current state before changing it

Capture the current service definition and tasks so you can compare the result and give another operator a recovery reference:

$ docker service inspect "$SERVICE_NAME" > service-before-rollback.json
$ docker service ps "$SERVICE_NAME" > tasks-before-rollback.txt
$ test -s service-before-rollback.json && echo 'service definition captured'
service definition captured

These files can contain deployment details such as image names, environment settings or secret references. Store them with the same care as operational records, and remove them later only under your normal retention policy. The redirection creates or replaces local files, so pick a controlled directory rather than a shared location for sensitive output.

Recovery: this is your recovery boundary. docker service rollback gives you no general undo switch in the installed interface. If the previous configuration is unsuitable, use your documented deployment command or a reviewed docker service update that restores the known-good settings. Do not assume that running rollback twice returns you to where you started.

4. Warn the service owner and start the rollback

Warning: this step changes the service specification and can replace running tasks. Confirm the maintenance window, alert routing and application owner before running it. Do not add sudo unless your Docker installation specifically requires it and your operator policy allows it.

Run the normal attached form so the CLI waits for the service to converge:

$ docker service rollback "$SERVICE_NAME"
SERVICE_NAME

Successful progress output varies with the CLI and the service. The command should return zero once the service has converged, or report a failure instead of silently claiming success. A non-zero exit is a failed attempt, not a reason to repeat it immediately: record the error and inspect the service and tasks first.

Use --detach only when you deliberately want the command to return before convergence:

$ docker service rollback --detach "$SERVICE_NAME"
SERVICE_NAME

Detached mode suits an operator who will monitor the rollout separately, but its successful return does not mean the tasks are healthy. --quiet only suppresses progress output; it does not make the operation safer. Do not combine both merely to make automation look clean unless that automation performs its own checks.

5. Verify the rolled-back tasks

Whether you waited or detached, check the service and its tasks:

$ docker service ls
$ docker service ps "$SERVICE_NAME"
ID            NAME                 IMAGE:TAG   NODE       DESIRED STATE   CURRENT STATE
TASK_ID       SERVICE_NAME.1       IMAGE:TAG   NODE_NAME  Running         Running ... ago

Look for the expected replica count, the intended image, and tasks whose desired and current states are healthy. The timestamps, IDs and exact columns vary. A task showing Rejected, repeated restarts or an unexpected image needs investigation: do not declare success from the rollback command's exit status alone.

Compare the resulting definition with the captured file when you need to prove which fields changed:

$ docker service inspect --pretty "$SERVICE_NAME"
$ docker service ps "$SERVICE_NAME" --no-trunc

Then test the application through its normal health check or client path. Swarm can report running tasks while the application is still failing readiness checks, serving the wrong configuration, or unable to reach a dependency.

6. Handle a failed or unsafe rollback

If the command fails before or during convergence, preserve the error text and inspect the service again. Check you are still on a manager, that the service exists, and that nodes can pull the referenced image. Do not delete the service as a first response: removal is a separate, destructive operation that loses the service object entirely.

If the previous specification is also broken, pause further changes and use the recorded definition plus your deployment source of truth to choose a known-good configuration. A rollback can restore the earlier service settings, but it cannot repair a bad image, an unavailable registry, a broken dependency or an invalid application configuration. Escalate to the service owner when the recovery path is not documented.

Once the service is healthy, keep the before-and-after records until the incident or change review is complete. Remove temporary copies only after confirming the authoritative deployment record exists and contains no accidentally exposed sensitive data.

Done means