htmltree turns a messy HTML file into an indented parse tree, including the nodes the parser quietly inserts on your behalf. The examples use the installed libhtml-tree-perl package, version 5.07-3, on this machine. Allow about ten minutes. You need a shell, a readable HTML file and the package installed.
Checkpoint: This is an inspection tool. It reads files and prints a tree; it does not rewrite the input. The package on this host installs the example program at /usr/share/doc/libhtml-tree-perl/examples/htmltree, but does not provide an htmltree command on PATH. Use that explicit path unless your distribution has installed a wrapper elsewhere.
Confirm the package version and the example path with ordinary, unprivileged commands:
$ dpkg-query -W -f='${Package} ${Version}\n' libhtml-tree-perl
libhtml-tree-perl 5.07-3
$ test -x /usr/share/doc/libhtml-tree-perl/examples/htmltree && echo 'htmltree example is executable'
htmltree example is executable
If the second check fails, stop and locate the program supplied by your operating system package. Do not download a replacement into a system directory merely to make this guide's path work: the manual page describes the program's options, while the installed example is the runnable file verified here.
Use a temporary file when learning the output shape. This example includes a normal element, text, a nested element and an image:
$ workdir=$(mktemp -d /tmp/htmltree.XXXXXX)
$ printf '%s\n' '<html><body><p>Hello <b>world</b>.</p><img src="x.png"></body></html>' > "$workdir/sample.html"
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree "$workdir/sample.html"
The output starts with a separator and the file name, then shows the tree. On this machine the useful part is:
<html> @0
<head> @0.0 (IMPLICIT)
<body> @0.1
<p> @0.1.0
"Hello "
<b> @0.1.0.1
"world"
"."
<img src="x.png" /> @0.1.1
The @ values are tree positions, not source line numbers. Notice the implicit head: the parser builds a full document tree even though the input never contained a head element. Text appears as quoted leaf values, and attributes sit on the element line.
Pass one or more existing files after the options. Each regular file is reported separately:
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree /path/to/page.html
==============================================================================
Parsing /path/to/page.html...
... parse tree ...
Replace the placeholder with a path you can read. The program skips arguments that are not regular files, so a misspelled path can produce no parse section at all. Check the path first when output is unexpectedly empty:
$ test -f /path/to/page.html && test -r /path/to/page.html && echo readable
For a comparison, pass several files in one invocation:
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree old.html new.html
Keep the separator and Parsing ... line in any saved output. They make it obvious which tree belongs to which input.
Use -w before the file names to enable warnings on each newly created tree:
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -w /path/to/page.html
Warnings are diagnostic output only. They do not repair the file, and a tree being printed does not prove the source is valid HTML for a browser or application. Save a copy of the original before making any later edit: there is no undo in htmltree because it never modifies the input in the first place.
Use -D followed immediately by a number to set HTML::TreeBuilder::Debug:
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -D3 /path/to/page.html
Debug level 3
==============================================================================
Parsing /path/to/page.html...
The debug level prints before the parse. Higher levels can get noisy fast, so start small and redirect the output to a new file if you need to compare runs:
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -D3 /path/to/page.html > /tmp/page-tree.txt
$ test -s /tmp/page-tree.txt && echo 'tree output saved'
Do not confuse -D3 with a file named 3. The option must be written as one argument, with the number attached, and the program's option parser only consumes its documented switches before the file list.
Use -h for the installed usage message:
$ /usr/share/doc/libhtml-tree-perl/examples/htmltree -h
htmltree - Parse the given HTML file(s) and dump the parse tree
When you have finished checking the sample, remove only the temporary directory you created, after verifying its exact path:
$ printf 'temporary directory: %s\n' "$workdir"
$ find "$workdir" -maxdepth 1 -type f -print
$ rm -r -- "$workdir"
Warning: The final command deletes that directory and its contents. Do not substitute a home directory, a repository path or an unresolved variable. Keep real source files elsewhere; htmltree never needs elevated privileges unless the file itself is readable only by root.
libhtml-tree-perl version and the actual executable path.-w, -Dnumber and -h.