Inspect HTML Structure with htmltree

htmltree turns a messy HTML file into an indented parse tree, including the nodes the parser quietly inserts on your behalf. The examples use the installed libhtml-tree-perl package, version 5.07-3, on this machine. Allow about ten minutes. You need a shell, a readable HTML file and the package installed.

Checkpoint: This is an inspection tool. It reads files and prints a tree; it does not rewrite the input. The package on this host installs the example program at /usr/share/doc/libhtml-tree-perl/examples/htmltree, but does not provide an htmltree command on PATH. Use that explicit path unless your distribution has installed a wrapper elsewhere.

1. Check the installed package and program

Confirm the package version and the example path with ordinary, unprivileged commands:

$ dpkg-query -W -f='${Package} ${Version}\n' libhtml-tree-perl
libhtml-tree-perl 5.07-3
$ test -x /usr/share/doc/libhtml-tree-perl/examples/htmltree && echo 'htmltree example is executable'
htmltree example is executable

If the second check fails, stop and locate the program supplied by your operating system package. Do not download a replacement into a system directory merely to make this guide's path work: the manual page describes the program's options, while the installed example is the runnable file verified here.

2. Create a small input you can recognise

Use a temporary file when learning the output shape. This example includes a normal element, text, a nested element and an image:

$ workdir=$(mktemp -d /tmp/htmltree.XXXXXX)
$ printf '%s\n' '<html><body><p>Hello <b>world</b>.</p><img src="x.png"></body></html>' > "$workdir/sample.html"
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree "$workdir/sample.html"

The output starts with a separator and the file name, then shows the tree. On this machine the useful part is:

<html> @0
  <head> @0.0 (IMPLICIT)
  <body> @0.1
    <p> @0.1.0
      "Hello "
      <b> @0.1.0.1
        "world"
      "."
    <img src="x.png" /> @0.1.1

The @ values are tree positions, not source line numbers. Notice the implicit head: the parser builds a full document tree even though the input never contained a head element. Text appears as quoted leaf values, and attributes sit on the element line.

3. Parse a real file

Pass one or more existing files after the options. Each regular file is reported separately:

$ /usr/share/doc/libhtml-tree-perl/examples/htmltree /path/to/page.html
==============================================================================
Parsing /path/to/page.html...
... parse tree ...

Replace the placeholder with a path you can read. The program skips arguments that are not regular files, so a misspelled path can produce no parse section at all. Check the path first when output is unexpectedly empty:

$ test -f /path/to/page.html && test -r /path/to/page.html && echo readable

For a comparison, pass several files in one invocation:

$ /usr/share/doc/libhtml-tree-perl/examples/htmltree old.html new.html

Keep the separator and Parsing ... line in any saved output. They make it obvious which tree belongs to which input.

4. Turn on parser warnings

Use -w before the file names to enable warnings on each newly created tree:

$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -w /path/to/page.html

Warnings are diagnostic output only. They do not repair the file, and a tree being printed does not prove the source is valid HTML for a browser or application. Save a copy of the original before making any later edit: there is no undo in htmltree because it never modifies the input in the first place.

5. Increase parser debugging when the tree is surprising

Use -D followed immediately by a number to set HTML::TreeBuilder::Debug:

$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -D3 /path/to/page.html
Debug level 3
==============================================================================
Parsing /path/to/page.html...

The debug level prints before the parse. Higher levels can get noisy fast, so start small and redirect the output to a new file if you need to compare runs:

$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -D3 /path/to/page.html > /tmp/page-tree.txt
$ test -s /tmp/page-tree.txt && echo 'tree output saved'

Do not confuse -D3 with a file named 3. The option must be written as one argument, with the number attached, and the program's option parser only consumes its documented switches before the file list.

6. Ask for help or clean up temporary output

Use -h for the installed usage message:

$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -h
htmltree - Parse the given HTML file(s) and dump the parse tree

When you have finished checking the sample, remove only the temporary directory you created, after verifying its exact path:

$ printf 'temporary directory: %s\n' "$workdir"
$ find "$workdir" -maxdepth 1 -type f -print
$ rm -r -- "$workdir"

Warning: The final command deletes that directory and its contents. Do not substitute a home directory, a repository path or an unresolved variable. Keep real source files elsewhere; htmltree never needs elevated privileges unless the file itself is readable only by root.

Done means