xmllint answers the question "is this XML actually well formed" faster than opening it in an editor and squinting. You will finish with a workflow for checking a file, pulling a value out with XPath, and reformatting a copy without touching the original. Allow about fifteen minutes.
You need a shell and a readable XML file. Every check here is an ordinary unprivileged command: no sudo, no DTD installed, no service altered, no input file overwritten. Examples use xmllint from libxml2 2.9.14, packaged here as libxml2-utils 2.9.14+dfsg-1.3ubuntu3.9.
Check which binary your shell will actually run, and note the library details while you are there:
$ command -v xmllint
/home/linuxbrew/.linuxbrew/bin/xmllint
$ xmllint --version
xmllint: using libxml version 21504
compiled with: Threads Tree Output Push Reader Patterns Writer SAXv1 DTDValid HTML C14N Catalog XPath XPointer XInclude Iconv ISO8859X Regexps Automata RelaxNG Schemas Schematron Modules Debug Zlib
Your path will differ. What matters is that --version succeeds and lists XPath and whatever validation features you plan to use. Note that the package version and the library version are related numbers, not the same string.
Start with a read-only parse. --noout suppresses the normal serialised output, leaving just the diagnostics if something is wrong:
$ xmllint --noout /path/to/input.xml
$ printf 'exit status: %s\n' "$?"
exit status: 0
Status 0 means the parser accepted the document. It does not mean the document satisfies a DTD or XML Schema, only that the markup itself is sound. Deliberately break the input and the command reports a parser error with a non-zero status:
$ printf '<root>\n' | xmllint --noout -
-:2: parser error : Premature end of data in tag root line 1
^
$ printf 'exit status: %s\n' "$?"
exit status: 4
The filename - means standard input. Capture the status straight after xmllint, because the next command overwrites it. A parser diagnostic is useful in a script, so do not swallow it with 2>/dev/null until you have actually finished investigating.
Checkpoint: run the first command against your real file. Non-zero means fix or replace the XML before touching XPath queries. A partially printed document from a failed parse is not trustworthy output.
By default, xmllint prints its result tree to standard output. Add --format to reindent it, then redirect to a new file rather than the one you started with:
$ xmllint --format /path/to/input.xml > /path/to/input.formatted.xml
$ xmllint --noout /path/to/input.formatted.xml
$ file /path/to/input.formatted.xml
/path/to/input.formatted.xml: XML 1.0 document, ASCII text
These commands leave the input untouched. The exact file description varies with document encoding, so the parse check is the verification that actually matters.
Warning: shell redirection with > truncates its destination before xmllint even runs. Never use the input path as the output path. If you do need to replace it, write a temporary sibling, validate that, then rename it as a controlled change:
tmp='/path/to/input.xml.new'
xmllint --format /path/to/input.xml > "$tmp" && xmllint --noout "$tmp"
status=$?
if [ "$status" -eq 0 ]; then
mv -- "$tmp" /path/to/input.xml
else
rm -f -- "$tmp"
exit "$status"
fi
This one changes state: the final mv replaces the original. Keep a backup if the source matters, and never run it against a live configuration file without that file's own recovery plan.
Reach for --xpath when you need a value, not a whole document. This example pulls the second item's text from standard input:
$ printf '<root><item id="a">one</item><item id="b">two</item></root>\n' | \
xmllint --xpath 'string(/root/item[@id="b"])' -
two
Quote the XPath in single quotes so the shell leaves its brackets, dollar signs and other punctuation alone. A node-set result serialises in full; an empty one prints XPath set is empty and returns an error status. Make that distinction explicit in a script rather than assuming blank output means success.
Namespaces are a classic trap here. An XPath such as /root/item will not automatically match elements sitting in a default XML namespace. Inspect the document and adapt the expression before you conclude the data is simply missing.
XML can point at external DTDs or entities, and --nonet stops the parser fetching them over the internet:
$ xmllint --nonet --noout /path/to/input.xml
$ printf 'exit status: %s\n' "$?"
exit status: 0
That is a good default for inspection and batch checks where the document should not trigger network fetches. It can also make a previously accepted file fail if it relies on a remote resource, but that failure is a dependency or policy issue, not a reason to rip the safeguard out.
Catalogues and local search paths can move where DTDs and entities are found. The installed manpage documents /etc/xml/catalog, XML_CATALOG_FILES and --path: treat those as part of a reproducible validation job's input, record them, and do not trust a result produced with an unexpected catalogue.
Well-formedness and schema validity answer different questions. Got a W3C XML Schema file? Pass it with --schema:
$ xmllint --noout --schema /path/to/schema.xsd /path/to/input.xml
/path/to/input.xml validates
$ printf 'exit status: %s\n' "$?"
exit status: 0
Use --relaxng for a Relax NG schema instead, --dtdvalid to select a DTD, or --valid to check the document's own included DTD. Do not bolt on one of these options just to make a parser check feel more thorough: you need the correct schema and its referenced resources, not a placebo.
Validation errors carry their own non-zero statuses in the documented range, but a script should treat every non-zero result as failure unless it deliberately handles a specific code. A clean parse without a schema is not evidence that the application's required fields are actually present.
xmllint --version identified the installed libxml2 build.--noout ran without producing misleading output.--nonet was weighed for untrusted or automated input.