Home / Alt manpages / preconv(1)

  • preconv(1)
  • User command
  • linux

Make roff Encoding Predictable with preconv

You will finish with a repeatable way to feed non-ASCII roff source to GNU troff, using the encoding that the file actually has. The examples use GNU preconv from groff 1.23.0, installed here as package version 1.23.0-3build2, with iconv and uchardet support.

Allow about ten minutes. You need a shell, a roff input file or a pipe, and the groff-base package. These examples only read input and write converted output. They do not alter the source file or system configuration.

1. Check the installed version and interface

Start by checking the binary you will actually run:

$ preconv --version
GNU preconv (groff) version 1.23.0 with iconv support and with uchardet support

The version matters because encoding detection and the available conversion libraries are build and release details. The useful command form is preconv [-dr] [-D fallback-encoding] [-e encoding] [file ...]. With no file, or with - as a file name, it reads standard input and writes converted roff to standard output.

Checkpoint: run preconv --help if you are unsure whether an option belongs to this installed release. The help output also identifies UTF-8 as the default fallback reported by this build.

2. Convert a known UTF-8 file

For a normal source file, let preconv inspect it and send the result to another command or file. This example creates a temporary input with a UTF-8 accented character, then displays the converted roff escape:

$ printf 'caf\303\251\n' > /tmp/example.roff
$ preconv -r /tmp/example.roff
caf\[u00E9]

Characters in the ASCII range remain ordinary text. Other characters are represented as groff Unicode escapes such as \[u00E9], which troff can interpret. The -r option means raw output: it suppresses the default .lf request that identifies the input for later diagnostics.

If you leave out -r, the same input begins with a line marker:

$ preconv /tmp/example.roff
.lf 1 /tmp/example.roff
caf\[u00E9]

That marker is useful when later groff diagnostics need to identify the original file and line. It is output, not a change to the source. Do not remove it merely because it is unfamiliar.

3. Declare the encoding when detection is not trustworthy

Detection is convenient, but a short file or a pipe may not contain enough evidence. Use -e when you know the input encoding and need to skip detection:

$ printf 'caf\351\n' > /tmp/latin1.roff
$ preconv -r -e latin-1 /tmp/latin1.roff
caf\[u00E9]

The input above contains byte 0xE9, which is the accented character in Latin-1. Declaring it as UTF-8 would be wrong, and declaring it as US-ASCII would discard that byte with this iconv-enabled build:

$ preconv -r -e us-ascii /tmp/latin1.roff
caf

This is a data-loss boundary. Treat -e as a statement about the bytes, not as a repair switch. If you do not know the encoding, preserve the original and establish it from the producing application, file metadata or a controlled sample before converting.

4. Make fallback behaviour explicit

When no explicit encoding, byte-order mark or coding tag settles the question, preconv tries detection and then falls back. You can choose that final fallback with -D:

$ LC_ALL=C preconv -r -D latin-1 /tmp/latin1.roff
caf\[u00E9]

The documented order is explicit -e, Unicode byte-order mark, a recognised GNU Emacs coding tag in the first or second line of a seekable file, uchardet when available, -D, and finally the locale. For the C, POSIX or empty locale, the manual describes Latin-1 as the fallback. In this installation, --help reports UTF-8 as the default fallback when no more specific choice succeeds. The practical rule is simple: use -e for a known file encoding and -D when you deliberately want a final fallback.

A coding tag can document the source for tools that support the GNU Emacs convention. It must be in a roff comment on the first or second line, using the default control and escape characters, for example:

\" -*- coding: latin-1 -*-
caf\351

Keep the tag with the source file. It is an input hint, not a conversion result, and it is considered only for a seekable file.

5. Treat pipes as a separate case

A regular file can be sought back to its opening lines. A pipe cannot. When input arrives through a pipe, coding-tag and uchardet detection are skipped, so a short or ambiguous stream can be interpreted using the wrong locale fallback.

$ printf 'caf\351\n' | preconv -r -e latin-1
caf\[u00E9]
$ printf 'caf\351\n' | preconv -r -D latin-1
caf\[u00E9]

Prefer -e when the producer's encoding is known. Use -D only when treating every undecided input as one known fallback is genuinely safe. Do not assume that adding a coding tag to data already in a pipe will make preconv see it.

6. Connect it to groff without duplicating work

In ordinary use, you usually do not need to invoke preconv yourself. The preprocessor is intended to be selected through groff's -k or -K options. Use direct preconv when you want to inspect or record the converted stream, or when a pipeline needs this stage explicitly.

For a direct pipeline, keep standard output separate from diagnostics and choose an output destination deliberately:

$ preconv -r -e utf-8 document.roff > document.converted.roff
$ groff -Tutf8 document.converted.roff > document.txt

Redirection truncates an existing destination before the command runs. Before replacing a valuable generated file, write to a new name and compare it. To recover from an unwanted conversion, use the original source or remove only the derived output after confirming its exact path. preconv itself has no in-place editing mode.

7. Know what preconv cannot see

preconv converts the input it receives. It cannot transform text that is inserted later by soelim, by troff's so request, or through strings supplied with troff's -d option. Put the conversion stage where the bytes enter the pipeline, and do not infer that included files received the same treatment.

It also assumes the default backslash escape character and emits special-character escapes accordingly. If your roff source changes that escape convention, inspect the pipeline design rather than assuming generated Unicode escape sequences will still be interpreted as intended.

Done means

  • preconv --version identifies the expected groff release and conversion support.
  • You know whether each input is UTF-8, Latin-1 or another supported encoding.
  • You use -e for a known encoding, or -D for an intentional fallback.
  • You understand that pipes skip seekable-file detection.
  • You choose -r knowingly, preserving .lf markers when useful for diagnostics.
  • You keep the original source and write converted output to a separate destination.