Extract PDF Images Safely with pdfimages
You will inspect the images embedded in a PDF, then extract selected images into a separate directory without changing the original document. The examples use pdfimages from Poppler 24.02.0, provided here by poppler-utils version 24.02.0-1ubuntu9.9. Allow about ten minutes for a small PDF, plus the time needed to inspect the extracted files.
The route
Jump straight to the step you need, or tick off Done means at the end.
This guide assumes a readable PDF and a shell. Extraction normally needs no elevated privileges. Do not use sudo unless the input or destination is genuinely inaccessible to your user. The command can write many files, so choose an empty or newly created destination directory before you start.
1. Check the installed tool
Confirm that the command resolves to the binary you expect and record its version. These are read-only checks:
$ command -v pdfimages
/usr/bin/pdfimages
$ dpkg-query -W -f='${Package} ${Version}\n' poppler-utils
poppler-utils 24.02.0-1ubuntu9.9
$ pdfimages -v 2>&1 | head -n 2
pdfimages version 24.02.0
Copyright 2005-2024 The Poppler Developers - http://poppler.freedesktop.org
The manual installed with this package describes the command as version 3.03, but the executable reports Poppler 24.02.0. Use the executable's output when recording what actually ran on your machine.
Checkpoint
You have a readable PDF path and a destination name that does not contain files you need to preserve.
2. List embedded images before extracting them
Use -list for a read-only inventory. With this option, do not supply an image-root:
$ pdfimages -list /path/to/document.pdf
The output contains one row for each image or mask and reports fields such as page, image number, type, embedded width and height, colour space, bits per component, encoding, resolution and size. The values are properties of the embedded image. In particular, width and height are not the dimensions at which a PDF viewer happens to render it.
Several rows can belong to one visible picture. Transparency is represented by a separate image and mask, and the manual says that a mask or soft mask immediately follows its associated image. Do not assume that image number 1 is the first photograph, or that every row is a standalone picture suitable for sharing.
To keep a long inventory available for later comparison, redirect it to a text file in your working directory:
$ pdfimages -list /path/to/document.pdf > image-inventory.txt
$ test -s image-inventory.txt && echo "inventory written"
inventory written
Redirection replaces an existing file. If that matters, choose a new name first or use a temporary file and inspect it before replacing an old report.
3. Extract into a new directory
Create a destination with a name that makes its purpose clear. This changes the filesystem by creating the directory, so check the path before running it:
$ mkdir -p pdfimages-output
$ test -d pdfimages-output && echo "ready: pdfimages-output"
ready: pdfimages-output
$ pdfimages /path/to/document.pdf pdfimages-output/image
$ find pdfimages-output -maxdepth 1 -type f -printf '%f\n' | sort
The second argument is an output root, not a directory. Each extracted file uses a name such as image-000.jpg or image-001.ppm, depending on the image and selected options. The exact list is document-specific. Keep the original PDF untouched and treat every generated file as disposable until checked.
By default, non-monochrome images are written as PPM and monochrome images as PBM. Those formats are useful for lossless pixel data but can be much larger than the source PDF. If you need ordinary image files for review, request PNG or TIFF output explicitly.
4. Choose an output format deliberately
Use -png to make PNG the default for images that are not being copied in a native encoded format:
$ mkdir -p pdfimages-png
$ pdfimages -png /path/to/document.pdf pdfimages-png/image
$ file pdfimages-png/image-*
Use -tiff when TIFF is the better fit for your next tool:
$ mkdir -p pdfimages-tiff
$ pdfimages -tiff /path/to/document.pdf pdfimages-tiff/image
$ file pdfimages-tiff/image-*
The output of file depends on the PDF. A wildcard that matches no files is a shell error in some shells, so inspect the directory with find if the extraction produced no files.
If the PDF contains JPEG, JPEG2000, JBIG2 or CCITT data, the native options -j, -jp2, -jbig2 and -ccitt preserve those encodings instead of converting them to the normal default. The combined -all option enables all four native modes and selects TIFF for CMYK images and PNG for other images. Native output is not automatically safer or easier to edit: JBIG2 and CCITT extraction can create companion files and parameters.
5. Restrict pages and make filenames useful
Limit the scan with -f for the first page and -l for the last page. Page numbers are inclusive:
$ mkdir -p pdfimages-pages-3-5
$ pdfimages -f 3 -l 5 -png /path/to/document.pdf pdfimages-pages-3-5/image
$ find pdfimages-pages-3-5 -maxdepth 1 -type f -printf '%f\n' | sort
These options reduce work and help avoid unrelated images in a large document. They do not mean image numbers 3 through 5. Image numbering still follows the images encountered during the scan.
Add -p when page numbers in the filenames will help you trace an image back to its source:
$ mkdir -p pdfimages-page-names
$ pdfimages -p -png /path/to/document.pdf pdfimages-page-names/image
$ find pdfimages-page-names -maxdepth 1 -type f -printf '%f\n' | sort
Read the resulting names rather than assuming a particular zero-padding scheme. The installed command's documented contract is that page numbers are included; the exact names and extensions remain dependent on the input and format.
6. Handle passwords and failures carefully
A PDF may require a user password before it can be read. Supply it with -upw only when you have a legitimate password:
$ pdfimages -upw 'REPLACE_WITH_PDF_PASSWORD' -list /path/to/protected.pdf
Passwords placed directly in a shell command can enter shell history and process listings. Prefer a controlled, private shell session and remove the command from history according to your shell's normal procedure. Do not paste a real password into a shared terminal, ticket or script.
-opw supplies an owner password and the manual says it bypasses PDF security restrictions. That is a security-sensitive operation. Only use it where you are authorised to do so, and do not treat it as a way around access controls on somebody else's document.
Check the exit status when using the command in a script:
$ pdfimages -list /path/to/document.pdf
$ status=$?
$ printf 'pdfimages exit status: %s\n' "$status"
pdfimages exit status: 0
The documented statuses are 0 for no error, 1 when the PDF cannot be opened, 2 when an output file cannot be opened, 3 for a PDF-permission error and 99 for another error. A successful exit status confirms that the command completed; it does not confirm that every extracted image is the one you wanted. Check the inventory and open representative outputs with a trusted image tool.
7. Avoid accidental overwrites and recover cleanly
Do not point the output root at an existing collection until you know how the installed command will behave for that document. The safest recovery is to stop, inspect the generated names, and move the whole new directory aside rather than deleting files immediately:
$ mv pdfimages-output pdfimages-output-review
$ find pdfimages-output-review -maxdepth 1 -type f -print
This undo step only changes the directory name. If you no longer need the extracted files, review the exact directory first, then remove it explicitly with rm -r pdfimages-output-review. That deletion is irreversible, so do not include it in an unattended command. The source PDF is not modified by any of the extraction commands in this guide.
Done means
- You confirmed the installed Poppler version and used a readable PDF.
- You ran
-listbefore extraction when the document's image structure mattered. - Images were written to a new, clearly named directory with an intentional format.
- Page limits and
-pwere used when they made tracing results easier. - You checked the exit status and inspected representative files.
- The original PDF remains untouched, and any generated directory can be moved aside or removed after review.