Blog / Security

  • backups
  • encryption
  • security
  • key-management
  • authenticated-encryption
  • linux

Why Encrypting Your Backups Is Not the Same as Encrypting Them Well

Here is a backup script many of us have written at some point: tar piped into openssl enc, a passphrase in a file, a cron entry. It ticks the box marked "encrypted". It also has several problems the tick box cannot see.

tar czf - /srv/data \
  | openssl enc -aes-256-cbc -pbkdf2 -pass file:/root/backup.pass \
  > /mnt/offsite/data-$(date +%F).tar.gz.enc

This post is about the gap between "the bytes are ciphertext" and "the backup still does its job under attack and under failure". Five gaps, roughly in order of how often they bite.

The key sits next to the data

The commonest failure is boring: the key lives wherever the backup can be reached from. A passphrase file in /root on the machine being backed up is fine against a stolen offsite disk. It is useless against an attacker who owns the machine, because they now have the key and write access to the backup target.

The inverse is worse. The key exists only on the machine that died. The backup is perfectly encrypted and perfectly unreadable, and you find out mid-disaster.

  • Keep a copy of the key somewhere that does not share a failure with the source (a password manager, a printed sheet in a safe, a hardware token).
  • Decide who may read the key, separately from who may write backups.
  • Rehearse a restore on a clean machine with only the stored key. If it needs anything else, that thing is part of your key.

Nothing notices if the ciphertext is changed

AES-CBC gives confidentiality and nothing else. openssl enc does not authenticate what it produces, so a flipped bit in storage, or one flipped on purpose, goes straight through decryption.

Quick detour, because this is neat in a slightly horrid way. In CBC, flipping a bit in ciphertext block N garbles plaintext block N completely, but flips exactly the same bit in block N+1. An attacker who cannot read your data can still make precise edits to the block after the one they sacrificed.

For a gzip stream the damage is mostly "decompression fails halfway", which is bad enough. The fix is authenticated encryption (AES-GCM, ChaCha20-Poly1305) so any modification fails loudly. I covered the same trap from the Rust side in the post on AES-CBC needing a MAC.

Tools like age, restic and Borg authenticate their data. With age, the stream is split into authenticated chunks, and a truncated file is rejected rather than silently decrypting to a shorter archive.

A valid old backup is still valid

Authentication proves the data was produced by someone holding the key. It does not prove it is the latest data. An attacker with write access can swap today's backup for last year's, and every check passes.

This is a rollback attack, and plain file-per-day encryption has no defence. Mitigations are mostly operational:

  • Make the backup target append-only, so the client cannot delete or replace old data.
  • Record the expected latest snapshot ID or checksum somewhere the attacker cannot write.
  • Monitor for backups that stopped growing or whose dates are wrong.

The structure leaks even when the contents do not

Ciphertext still has a size. A single encrypted tarball tells an observer roughly how much data you hold and how that changed between nights. File-per-file schemes also reveal how many files exist and how big each is, which can be surprisingly identifying.

Deduplicating tools add a subtler leak. They split data into chunks at content-defined boundaries. If those boundaries were the same for everyone, chunk sizes could fingerprint a known file, such as a leaked document, in your repository. Restic and Borg both derive the chunker parameters from per-repository secret material, which is the reason this is a non-issue for them and a real consideration if you write your own.

Whether any of this matters depends on your threat model. If the adversary is a thief with a disk, sizes are irrelevant. If it is someone with long-term access to the storage provider, they are not.

The passphrase is the whole security level

The cipher is 256-bit; your passphrase is probably not. A key derivation function slows guessing, but it only slows it. -pbkdf2 in openssl enc uses a default iteration count that was reasonable once and is modest by current standards, and it is not memory-hard, so GPUs guess quickly.

Better options, in order of preference:

  1. A random key generated by the tool, stored as in the first section. No guessing at all.
  2. A long random passphrase from a password manager, run through a memory-hard KDF such as scrypt or Argon2 (restic and Borg do this; age uses scrypt in passphrase mode).
  3. A human-memorable passphrase, with eyes open about what that means.

What a better version looks like

You do not need to write cryptography to fix most of this. A tool that already authenticates, derives keys properly and handles chunking is the lazy and correct answer.

age -r "$(cat /etc/backup/recipient.pub)" -o data.tar.age < data.tar

Encrypting to a public key with age means the machine being backed up never holds the decryption key, which removes the worst of the first problem. It does nothing for rollback or size leakage, which is why a repository tool with an append-only mode covers more ground.

Whichever you pick, the test that matters is the dull one: restore from the offsite copy, on another machine, using only what you wrote down. Then flip one byte in the ciphertext and confirm the restore refuses rather than shrugs.