Home / Alt manpages / pdftotext(1)

  • pdftotext(1)
  • User command
  • linux

Extract Reliable Text from PDFs with pdftotext

You will finish with plain text from a PDF, a repeatable way to extract only the pages you need, and a quick method for deciding whether the result is trustworthy. The examples use pdftotext from Poppler 24.02.0, provided here by the poppler-utils package.

Allow about ten minutes. You need a readable PDF and a shell. The normal conversion is unprivileged: do not use sudo unless file permissions genuinely require it. The command extracts text that already exists in the PDF; it does not perform OCR on a scanned page.

1. Check the installed command

Confirm which executable and package version you are using. These are read-only checks:

$ command -v pdftotext
/usr/bin/pdftotext
$ dpkg-query -W -f='${Package} ${Version}\n' poppler-utils
poppler-utils 24.02.0-1ubuntu9.9
$ pdftotext -v 2>&1 | head -n 1
pdftotext version 24.02.0

Your package revision may differ. The installed program's help and manpage are the contract for this machine. In particular, Poppler releases can add options, so do not copy a flag from a newer system without checking pdftotext -h first.

2. Convert a PDF to a new text file

Give the input PDF and an explicit output path. The output file is created or replaced, so choose a new name if the existing text matters:

$ pdftotext /path/to/input.pdf /path/to/input.txt
$ test -s /path/to/input.txt && echo 'text file is non-empty'
text file is non-empty

With no second argument, pdftotext derives the destination by replacing .pdf with .txt. An explicit destination is easier to review in scripts and avoids surprises when the source has an unusual name.

Checkpoint: inspect the beginning without opening a potentially large file in an editor:

$ sed -n '1,30p' /path/to/input.txt

A successful exit status means the conversion completed. It does not prove that columns, reading order or font encodings were reconstructed correctly.

3. Send the result to standard output

Use - as the text destination when the next shell command should consume the extracted text:

$ pdftotext /path/to/input.pdf - | sed -n '1,40p'

The same convention works for the input. This lets you read a PDF from standard input, although a file path is usually clearer when troubleshooting:

$ cat /path/to/input.pdf | pdftotext - - | sed -n '1,20p'

A pipeline can hide which command failed. Preserve the exit status when a script needs to act on failure. In Bash, enable pipefail before the pipeline, or use a temporary output file and test its status directly.

4. Choose reading order or physical layout

The default tries to undo the PDF's physical arrangement, including columns and hyphenation, and emits reading order. That is often best for paragraphs, but a table or form may need the page's spacing instead:

$ pdftotext -layout /path/to/report.pdf /tmp/report-layout.txt
$ sed -n '1,40p' /tmp/report-layout.txt

Compare both modes before building a parser:

$ pdftotext /path/to/report.pdf /tmp/report-reading.txt
$ diff -u /tmp/report-reading.txt /tmp/report-layout.txt | sed -n '1,80p'

-layout preserves spacing as well as the converter can, not as a guarantee that the PDF's visual table becomes a valid data table. For fixed-pitch or tabular text, -fixed NUMBER forces physical layout with the supplied character width in points. Measure the value from the document rather than guessing, then inspect the result.

Do not start with -raw as a general fix. It keeps content-stream order, and the local manpage describes it as a hack that is no longer recommended. It can be useful for investigating a troublesome file, but it is usually less readable.

5. Limit the page range and line endings

Extract a page interval with -f for the first page and -l for the last page. Page numbers are document page numbers, not zero-based indexes:

$ pdftotext -f 3 -l 5 /path/to/manual.pdf /tmp/manual-pages-3-to-5.txt
$ wc -l -w /tmp/manual-pages-3-to-5.txt
      86     512 /tmp/manual-pages-3-to-5.txt

The counts will depend on the PDF. Check the first and last extracted headings if the range matters. If you are feeding a cross-platform text workflow, make the line endings explicit:

$ pdftotext -eol unix /path/to/manual.pdf /tmp/manual-unix.txt
$ file /tmp/manual-unix.txt

The accepted values are unix, dos and mac. The default is suitable for many Linux tools, but an explicit value makes an automated job easier to audit.

6. Handle encoding, page breaks and metadata output

Text output defaults to UTF-8. Check the available encoding names before selecting a legacy encoding:

$ pdftotext -listenc | sed -n '1,12p'
Available encodings are:
ASCII7
Big5
Big5ascii
EUC-CN
EUC-JP
EUC-KR

The full list can be longer than this sample. Use -enc ENCODING only when the receiving system needs something other than UTF-8. A wrong choice can turn readable characters into replacement symbols.

By default, page breaks are represented by form-feed characters. Use -nopgbrk when a line-oriented consumer must not see those separators:

$ pdftotext -nopgbrk /path/to/input.pdf /tmp/no-page-breaks.txt
$ grep -n $'\f' /tmp/no-page-breaks.txt || echo 'no form-feed page breaks found'

For geometry rather than prose, -bbox and -bbox-layout produce XHTML containing bounding boxes. They are different output formats, not improved plain-text extraction. Use them when coordinates are part of the task, and inspect the generated file as XHTML.

7. Diagnose bad or protected PDFs

If the output is empty or garbled, first check whether the PDF contains an extractable text layer. A scanned document may contain only images, in which case pdftotext cannot recover words. OCR is a separate workflow; do not mistake an empty result for a successful transcription.

Fonts with damaged encodings can also defeat extraction. Try -layout and inspect another page, but keep the original PDF unchanged. The manpage explicitly identifies badly mangled font encodings as a case where OCR may be needed.

For an encrypted document, the command may need a password. -upw PASSWORD supplies the user password and -opw PASSWORD supplies the owner password. Do not put a real password in shell history, process listings or a shared script. Prefer the least exposed method available in your environment, and do not treat an owner password as permission to bypass someone else's controls.

Use the exit status to distinguish common failures:

$ pdftotext /path/to/missing.pdf /tmp/missing.txt
$ printf 'exit status: %s\n' "$?"
exit status: 1

The documented statuses are 0 for success, 1 when the PDF cannot be opened, 2 when the output cannot be opened, 3 for PDF permission errors, and 99 for other errors. If you used a temporary output path, remove only the incomplete temporary file after a failed run. Do not delete the source PDF as a cleanup shortcut.

Done means

  • You confirmed the installed Poppler version and checked its local help.
  • The PDF was converted to a deliberately chosen output path without elevated privileges.
  • You inspected the extracted text and selected default or layout mode based on the document.
  • Page ranges, line endings and page-break handling match the receiving workflow.
  • You treated empty output, damaged fonts and encryption as extraction limits to investigate, not as evidence that the text is correct.
  • The original PDF remains untouched, and failed conversions cannot overwrite the known-good result.