Home / Alt manpages / pdfseparate(1)

  • pdfseparate(1)
  • User command
  • linux

Split a PDF into Verified One-Page Files with pdfseparate

You will finish with one PDF file per selected page, named from a pattern you control, and a quick check that each output is a single-page PDF. The examples use pdfseparate from Poppler 24.02.0, supplied by Ubuntu package poppler-utils version 24.02.0-1ubuntu9.9.

Allow about ten minutes. You need a readable, unencrypted PDF and a writable destination directory. The normal commands are unprivileged. Do not use sudo to extract a document unless its input or destination genuinely requires elevated access.

1. Check the installed command

Confirm the binary and its local version before building a script around it:

$ command -v pdfseparate
/usr/bin/pdfseparate
$ pdfseparate -v
pdfseparate version 24.02.0
Copyright 2005-2024 The Poppler Developers - http://poppler.freedesktop.org
Copyright 1996-2011, 2022 Glyph & Cog, LLC
$ dpkg-query -W -f='${Package} ${Version}\n' poppler-utils
poppler-utils 24.02.0-1ubuntu9.9

Your copyright lines and package revision may differ. The useful checkpoint is that the command exists and reports the Poppler version you intend to use.

2. Inspect the source and choose a safe destination

Before extraction, check that the input is a PDF and find its page count. pdfinfo is a separate Poppler utility, so this check does not alter the document:

$ pdfinfo /path/to/input.pdf | grep -E '^(Pages|Encrypted):'
Pages:           19
Encrypted:       no

The exact metadata varies. Stop if the file is encrypted: the pdfseparate manual says the PDF should not be encrypted, and this guide does not cover supplying a password or removing encryption.

Make a new output directory or use an empty existing one. This avoids confusing newly extracted pages with older files and makes it easier to spot an accidental overwrite:

$ mkdir -p /path/to/extracted-pages
$ find /path/to/extracted-pages -maxdepth 1 -type f -print

Do not place the output pattern in the same path as important originals unless you have checked every name it can produce. pdfseparate writes files directly, so an existing matching file can be replaced.

3. Extract every page with a numbered pattern

The command takes two positional arguments: the source PDF and a destination pattern. The pattern must contain %d, or a printf-compatible variant, because pdfseparate replaces it with the page number:

$ pdfseparate /path/to/input.pdf /path/to/extracted-pages/page-%d.pdf

For a 19-page source, the expected names are page-1.pdf through page-19.pdf. There is normally no progress report for each page. A return to the shell prompt means the command completed; verify the result rather than relying on the absence of an error.

Checkpoint: count the outputs and inspect their names:

$ find /path/to/extracted-pages -maxdepth 1 -type f -name 'page-*.pdf' | sort -V | wc -l
19
$ find /path/to/extracted-pages -maxdepth 1 -type f -name 'page-*.pdf' -printf '%f\n' | sort -V | sed -n '1,3p;$p'
page-1.pdf
page-2.pdf
page-3.pdf
page-19.pdf

Replace the count with the Pages: value from your own pdfinfo output. The count command is a useful check, but it does not prove that every output is valid or that no stale file was already present.

4. Extract only a page range

Use -f for the first page and -l for the last page, both inclusive. This example extracts pages 4 to 6:

$ pdfseparate -f 4 -l 6 /path/to/input.pdf /path/to/extracted-pages/section-%d.pdf
$ find /path/to/extracted-pages -maxdepth 1 -type f -name 'section-*.pdf' -printf '%f\n' | sort -V
section-4.pdf
section-5.pdf
section-6.pdf

Page numbers are PDF page positions, starting at 1. The output number remains the source page number, which is helpful when files are later put back in order. If you omit -f, extraction starts at page 1. If you omit -l, it ends at the last page.

Do not infer a range from a printed document's labels. A PDF can have front matter, Roman-numbered pages or labels that do not match its internal page positions. Use pdfinfo, a viewer or a careful visual check to establish the positions first.

5. Verify each extracted file

Each output should be an independent one-page PDF. Check a representative file with pdfinfo:

$ pdfinfo /path/to/extracted-pages/section-4.pdf | grep -E '^(Pages|Page size):'
Pages:           1
Page size:       609.714 x 789.041 pts

Then check all files in a batch. This loop prints a filename and its page count; every count should be 1:

$ for file in /path/to/extracted-pages/section-*.pdf; do
>     printf '%s: ' "$file"
>     pdfinfo "$file" | awk -F: '/^Pages:/ {gsub(/[[:space:]]/, "", $2); print $2}'
> done
/path/to/extracted-pages/section-4.pdf: 1
/path/to/extracted-pages/section-5.pdf: 1
/path/to/extracted-pages/section-6.pdf: 1

Open a sample if the visual content matters. A successful command and a one-page count do not tell you whether a page is the one you intended to send or archive.

6. Recover from mistakes safely

pdfseparate does not maintain a job record or an undo operation. If you selected the wrong range, leave the original PDF untouched and extract again into a fresh directory or with a new prefix. If a failed run left an incomplete output, remove only that known output after checking its path:

$ test -f /path/to/extracted-pages/section-6.pdf && printf '%s\n' 'review before removing'
$ rm -- /path/to/extracted-pages/section-6.pdf

The rm command is irreversible. Do not use a wildcard until you have listed the exact files it would match. If you need to replace an existing extraction, rename the old directory as a backup and create a new one, rather than mixing outputs from two runs.

Common failures are straightforward: an input error usually means the path or read permission is wrong; an encrypted source is outside this workflow; and a destination pattern without %d does not provide the required page numbering. Check the spelling, permissions and pattern before reaching for elevated privileges.

Done means

  • The installed Poppler version and source PDF were checked.
  • The source was readable and unencrypted, with its page count recorded.
  • The destination pattern contains %d and does not collide with valuable files.
  • Every requested source page produced one numbered, one-page PDF.
  • A sample or all outputs were checked with pdfinfo, and the original PDF remains untouched.