Home / Alt manpages / tr(1)

  • tr(1)
  • User command
  • linux

Translate, Delete and Clean Text Safely with tr

You will finish with a small set of reliable tr patterns for changing characters in a shell pipeline: converting case, removing unwanted characters and collapsing repeated separators. This guide uses GNU tr from coreutils 9.4, the version installed on this machine.

Allow about ten minutes. You need a shell and ordinary input that can be passed through standard input. These examples do not need elevated privileges and do not alter an input file unless you explicitly redirect output over it.

1. Confirm the installed command

Check which executable the shell will run and record its version:

$ command -v tr
/usr/bin/tr
$ tr --version | head -n 1
tr (GNU coreutils) 9.4

The manual describes tr as a standard-input to standard-output filter. It translates, squeezes or deletes characters according to one or two character arrays. It does not edit a named file in place.

Checkpoint

If command -v tr points somewhere unexpected, stop and inspect your PATH before putting the command into a script. A different implementation can have different extensions or locale behaviour.

2. Translate one set of characters into another

Give tr two arrays to translate characters from the first array to corresponding characters in the second. The following turns lower-case ASCII letters into upper-case ASCII letters:

$ printf '%s\n' 'server status: ok' | tr 'a-z' 'A-Z'
SERVER STATUS: OK

The shell supplies the string to printf, and the pipe supplies it to tr. There is no output file yet. This makes the command safe to experiment with: the original text is still only the input to the pipeline.

When the second array is shorter, GNU tr extends it by repeating its final character. Excess characters in the second array are ignored. Do not rely on either rule accidentally in a script; make both arrays explicit when the mapping matters.

For case conversion, the installed manual also documents character classes such as [:lower:] and [:upper:]. Locale rules can make class expansion broader than the simple ASCII ranges above. Use the ASCII form when the data is a protocol field, identifier or other deliberately ASCII value.

3. Delete characters you do not want

Add -d and provide only the array of characters to remove. This example removes every digit:

$ printf '%s\n' 'job-1842 finished' | tr -d '0-9'
job- finished

Deletion is literal at the character-array level. The range 0-9 means the characters from 0 through 9 in ascending order. It does not mean "find a number" and it does not remove surrounding whitespace.

To remove a small literal set, put those characters in the array:

$ printf '%s\n' 'name: [email protected]' | tr -d ':'
name [email protected]

Take care with shell quoting. Single quotes keep the shell from expanding characters in the array. A backslash, bracket expression or range can have meaning to tr and to the shell, so test the exact command with representative input before embedding it in a longer pipeline.

4. Squeeze repeated characters

Use -s to replace each run of a repeated character in the last specified array with one occurrence. This is useful for making a separator predictable:

$ printf '%s\n' 'alpha:::beta::::gamma' | tr -s ':'
alpha:beta:gamma

With only -s, the array is the first and only string argument. The command does not remove a leading or trailing separator; it only shortens each run. If you need both operations, combine deletion and squeezing deliberately. For example, this removes carriage returns and collapses repeated spaces:

$ printf '%s\n' 'alpha   beta   gamma' | tr -d '\r' | tr -s ' '
alpha beta gamma

Squeezing happens after translation or deletion. That ordering matters when a translation creates repeated characters. Run the stages separately while debugging a pipeline so you can see which stage changed the data.

5. Handle whitespace and newlines explicitly

The manual documents escapes such as \n for newline, \t for tab and \r for carriage return. This converts tabs to spaces:

$ printf 'one\ttwo\n' | tr '\t' ' '
one two

Do not confuse a newline with the shell's display of a command. To remove newlines from a small, controlled input, use:

$ printf 'one\ntwo\n' | tr -d '\n'
onetwo

Removing record separators from arbitrary input can join records into one long value. That is often irreversible once the pipeline has discarded the boundaries. If line structure matters, preserve a copy of the original and inspect the transformed output before sending it to another command.

6. Avoid locale surprises

GNU tr has full support only for safe single-byte locales, where every possible input byte represents one character. The installed manual specifically identifies the C locale as safe on GNU systems. For byte-oriented work, set it for this command:

$ printf '%s\n' 'abc123' | LC_ALL=C tr -d '[:digit:]'
abc

Character classes such as [:alpha:], [:space:] and [:punct:] are locale-aware. That can be useful for human text, but it can also produce a different result when the same script runs under another locale. Choose an explicit locale when reproducibility matters, and mention that choice in the script.

Checkpoint

Test with non-ASCII input if your real data contains it. Do not assume that a byte-oriented cleanup command is a Unicode normaliser. tr is a character-array filter, not a general text-encoding conversion tool.

7. Write output without destroying the input

Redirect to a new path while testing:

$ tr -s ' ' < raw.txt > cleaned.txt
$ cmp -- raw.txt cleaned.txt
raw.txt cleaned.txt differ: byte 6, line 1

The cmp result above is expected when repeated spaces were present. If it prints nothing and returns status 0, the files are byte-for-byte identical. Inspect the result with sed, od or the application that consumes it before replacing the original.

Warning

Shell redirection with > truncates its destination before tr runs. Never use tr ... > raw.txt when raw.txt is also the input. Use a temporary file in the same directory, verify it, then replace the original only if that replacement is genuinely intended and recoverable from backup.

There is no undo operation in tr. If you have overwritten a file and have no backup or source copy, the command cannot reconstruct the removed characters.

Done means

  • You confirmed the executable and GNU coreutils version used by the script.
  • You can distinguish translation, deletion and squeezing, and have tested their ordering.
  • Ranges, escapes and character classes match the data you actually process.
  • You selected an explicit locale for byte-sensitive or reproducible work.
  • Your test pipeline writes to a new output path and does not truncate its own input.