Convert PDFs to HTML with pdftohtml Without Losing the Output Files
You will finish with a repeatable way to turn a PDF into HTML, limit the conversion to a page range, and check which files were written. This guide uses pdftohtml from Poppler 24.02.0, supplied by Ubuntu's poppler-utils package.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about fifteen minutes. You need a readable PDF and enough free space for the HTML and extracted images. The examples write into a new working directory, so they do not alter the source PDF. No command here needs sudo; use elevated privileges only if your chosen input or destination is genuinely inaccessible to your account.
1. Check the installed command
Start with a read-only version check. This catches a common distraction: the options and output naming can vary between PDF tools, so do not substitute a similarly named converter.
$ command -v pdftohtml
/usr/bin/pdftohtml
$ pdftohtml -v
pdftohtml version 24.02.0
Copyright 2005-2024 The Poppler Developers - http://poppler.freedesktop.org
The package version is useful when you are comparing output on another machine:
$ dpkg-query -W -f='${Package} ${Version}\n' poppler-utils
poppler-utils 24.02.0-1ubuntu9.9
Your package revision may differ even when the Poppler program version is the same.
2. Make an isolated destination
Work in a directory that contains only this conversion. The command normally creates several files, and its positional HTML name is used as a naming stem rather than as a guaranteed single output file.
$ mkdir pdf-html-test
$ cd pdf-html-test
$ cp /path/to/input.pdf source.pdf
Replace /path/to/input.pdf with a real path. Keep the original PDF until you have checked the result. Before converting, confirm that it is readable:
$ file source.pdf
source.pdf: PDF document, version 1.x, ...
$ test -r source.pdf && echo 'input is readable'
input is readable
Checkpoint: you should now be in a throwaway directory with a readable file named source.pdf. If file cannot identify it as a PDF, fix the input path or obtain a complete copy before investigating conversion flags.
3. Convert a page range to ordinary HTML
For a first pass, convert pages 1 through 3 and use a clear output stem. The -f and -l options select the first and last page, respectively.
$ pdftohtml -f 1 -l 3 source.pdf converted.html
Page-1
Page-2
Page-3
The progress lines are normal. On Poppler 24.02.0, this command creates a frames file, one HTML file for each converted page, and extracted images where the document needs them. The exact names depend on the supplied stem and the document:
$ find . -maxdepth 1 -type f -printf '%f %s bytes\n' | sort
converted-html.html ... bytes
converted001.png ... bytes
converted002.png ... bytes
converted003.png ... bytes
converteds.html ... bytes
source.pdf ... bytes
$ file converted*.html
converted-html.html: HTML document, UTF-8 Unicode text
converteds.html: HTML document, UTF-8 Unicode text
Do not assume that a requested name of converted.html means that exact file will exist. Inspect the directory, then open the generated frames HTML in a browser or pass it to your next HTML-processing step.
4. Choose a single HTML document when needed
Use -s when you want one HTML document containing all selected pages instead of the normal frames-style set:
$ pdftohtml -s -f 1 -l 3 source.pdf single.html
Page-1
Page-2
Page-3
$ find . -maxdepth 1 -type f -name 'single*' -printf '%f %s bytes\n' | sort
single.html ... bytes
Output details such as image count, whitespace and generated CSS vary with the PDF. Check both that the file is non-empty and that its first lines look like HTML:
$ test -s single.html && echo 'single HTML exists'
single HTML exists
$ sed -n '1,5p' single.html
<!DOCTYPE html>
<html ...>
If your PDF has a large number of pages, a single document can be less convenient to process than the default per-page output. Select a smaller range with -f and -l when you are testing.
5. Send HTML to standard output
Use -stdout when another command should receive the HTML through a pipe. This avoids choosing generated filenames, but it does not automatically make the output safe to store as a single complete document in every mode.
$ pdftohtml -s -f 1 -l 1 -stdout source.pdf > page.html
$ test -s page.html && echo 'page.html exists'
page.html exists
$ sed -n '1,4p' page.html
<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml" ...>
Shell redirection truncates page.html before pdftohtml runs. For an output you cannot easily recreate, write to a new name and replace the old one only after checking it:
$ pdftohtml -s -stdout source.pdf > page.html.new
$ test -s page.html.new && mv page.html.new page.html
If conversion fails, leave the existing page.html untouched and inspect the error. Remove the incomplete page.html.new only after you have confirmed it is the failed temporary output.
6. Produce XML for position-aware processing
Use -xml when a later program needs text and coordinates rather than presentation HTML:
$ pdftohtml -xml -f 1 -l 1 source.pdf page.xml
Page-1
$ file page.xml
page.xml: XML 1.0 document, Unicode text, UTF-8 text
Do not treat XML output as HTML. The XML describes extracted page content and its positions, so it is usually a better input for a layout-aware parser. Add -noroundcoord only when retaining unrounded coordinates matters; the manual limits that option to XML output.
7. Handle images, passwords and warnings safely
Images are normally written as separate files. -i tells the converter to ignore them, while -dataurls asks for data URLs instead of external images where that build supports the option. If you need predictable downstream files, check the result rather than assuming that either choice was applied.
For an encrypted PDF, -upw supplies the user password and -opw supplies the owner password. Do not put a real password directly into a shell command that will be saved in history or process-monitoring logs. Use a protected execution environment appropriate to your workflow, and avoid sharing command output that includes credentials.
The -nodrm option overrides document DRM settings. That is a rights and policy boundary, not a routine troubleshooting switch. Use it only when you are authorised to do so. Likewise, -q hides messages and errors; avoid it until a conversion works, because silence can conceal a bad input or incomplete output.
8. Diagnose the likely failures
A non-zero exit status or an empty output needs investigation. Run the same command without -q, check the input and destination permissions, and verify that the selected page range exists. A successful exit status is not a visual-quality check, so inspect representative pages and images before deleting the source.
$ printf 'exit status: %s\n' "$?"
exit status: 0
$ find . -maxdepth 1 -type f -size 0 -print
The second command should print nothing. If it prints an output file, rerun into a fresh directory and check the PDF itself. If pages appear blank, try the documented -hidden option for PDFs whose useful text is marked hidden, or compare the visual PDF with the extracted result. Do not assume hidden text is trustworthy or that extraction preserves the original layout.
Checkpoint: you have identified the output mode, checked the files on disk, and kept the source PDF. Only then should you move generated files into a permanent destination. If you need to undo the conversion, remove or archive the generated HTML, XML and image files; the source PDF remains unchanged.
Done means
pdftohtmland its Poppler version were checked.- The PDF was copied or converted from a readable, preserved source.
- The page range and output mode were chosen deliberately.
- Generated filenames, file sizes and file types were inspected.
- Any password, DRM override or quiet-mode decision was treated as a security or troubleshooting boundary.
- The resulting HTML or XML was opened or parsed before the generated files were moved or the source was discarded.