Collapse Adjacent Duplicates with uniq
GNU uniq collapses repeated lines, counts each run, prints only duplicates, or keeps only lines that occur once. The one trap that catches almost everyone: it only compares lines sitting next to each other, not every line in the file.
The route
Jump straight to the step you need, or tick off Done means at the end.
- Time: about ten minutes.
- You need: a shell and a text file or command that produces line-oriented output.
- Version used here: GNU coreutils 9.4, package version 9.4-3ubuntu6.3. None of the examples needs
sudo.
1. Check the installed command
Confirm the binary and its version before relying on a newer option. This is read-only:
$ command -v uniq
/usr/bin/uniq
$ uniq --version | head -1
uniq (GNU coreutils) 9.4
$ dpkg-query -W -f='${Package} ${Version}\n' coreutils
coreutils 9.4-3ubuntu6.3
The command accepts an input file and an optional output file. With neither, it reads standard input and writes standard output:
$ printf '%s\n' apple apple pear apple | uniq
apple
pear
apple
The first and third apple lines stay separate because pear sits between them. Checkpoint: if you need duplicates found anywhere in the file, not just next to each other, plan to sort or otherwise group the input first.
2. Collapse a file safely
For a file whose repeated lines are already next to each other, pass its name directly:
$ uniq input.txt > collapsed.txt
$ test -s collapsed.txt && echo 'created non-empty output'
created non-empty output
The output file is a new result; the input is untouched. Watch the shell redirection: > input.txt truncates the input before uniq can even read it, so never use the same path on both sides.
If the destination already holds something useful, write to a fresh temporary name and swap it in only after checking the result:
$ uniq input.txt > collapsed.txt.new
$ diff -u collapsed.txt collapsed.txt.new
$ mv collapsed.txt.new collapsed.txt
diff prints nothing when the old and new results match. If something looks wrong, remove only collapsed.txt.new; the previous output stays intact. The mv itself changes state, but needs no elevated privileges when both files sit in a directory you can write.
3. Count each adjacent run
Use -c or --count when the number of consecutive occurrences matters:
$ printf '%s\n' apple apple pear apple | uniq --count
2 apple
1 pear
1 apple
The count is printed before the line, separated by whitespace, and it counts the current run only, not the total matches elsewhere in the input. That distinction matters in logs, where the same message can recur after other messages sit in between.
For machine processing, treat the first whitespace-separated field as the count and the remainder as the original line. Do not assume the displayed padding has a fixed width.
4. Select duplicate or unique runs
-d/--repeated: print one copy of each run that contains duplicates.-u/--unique: print only runs that occur once.-D/--all-repeated: print every line from duplicate groups, not just one copy.
$ printf '%s\n' apple apple pear apple apple | uniq --repeated
apple
apple
$ printf '%s\n' apple apple pear apple apple | uniq --unique
pear
There are two separate adjacent apple groups in that input, so --repeated prints one apple for each group rather than merging them into a single result. --unique discards every repeated group and keeps only pear.
$ printf '%s\n' apple apple pear apple apple | uniq --all-repeated
apple
apple
apple
apple
GNU uniq also supports --all-repeated=prepend and --all-repeated=separate to place an empty separator around groups. Keep the default when another program expects only the original duplicate lines.
5. Group unsorted input before using uniq
If the file is not arranged by the value you want to compare, sort it into a separate result first, then pass that to uniq:
$ printf '%s\n' apple pear apple banana pear | sort | uniq
apple
banana
pear
For just one copy of each sorted value, sort -u can replace the whole pipeline:
$ printf '%s\n' apple pear apple banana pear | sort -u
apple
banana
pear
Both approaches change the order, which may be unacceptable for a report or a log, and sorting can also change how locale rules order text. If input order is meaningful, do not sort merely to make uniq find more matches; decide first whether you actually need adjacent-run analysis or a set of globally distinct values.
6. Compare only the part of each line you actually mean
By default the entire line is compared. Three options narrow that:
-i/--ignore-case: ignore case when comparing.-f N: skip the firstNfields, where a field is a run of blanks followed by non-blank characters.-s N: skip the firstNcharacters instead.
$ printf '%s\n' 'A item' 'a item' 'B item' | uniq --ignore-case
A item
B item
$ printf '%s\n' '2026-09-27 one' '2026-09-27 two' '2026-09-28 three' | uniq --skip-fields=1
2026-09-27 one
2026-09-28 three
$ printf '%s\n' aa-1 aa-2 bb-1 | uniq --check-chars=2
aa-1
bb-1
These options change comparison only, not what gets printed. With --skip-fields=1, the complete first line of the matching group is what appears in the output. With --check-chars=2, only the first two characters decide whether adjacent lines match.
Check the input shape before choosing a field count. A timestamp, identifier or leading space may be part of the value you actually need to preserve, and an off-by-one field can silently collapse unrelated lines.
7. Handle records that use NUL delimiters
Use -z or --zero-terminated when records are separated by a NUL byte rather than a newline, which matters most when filenames or other values may themselves contain newlines:
$ printf 'one\0one\0two\0' | uniq --zero-terminated | od -An -t x1
6f 6e 65 00 74 77 6f 00
The output stays NUL-delimited. Keep -z on every later command that must understand the same record boundary. Do not inspect this output with an ordinary line-oriented tool and then assume the records are unchanged.
8. Diagnose misleading results
If a repeated value refuses to collapse, display the relevant input with markers or line numbers:
$ nl -ba input.txt
1 apple
2 pear
3 apple
Those apple lines are not adjacent, so the result is correct as it stands. If invisible whitespace might differ, inspect with a tool that makes it visible, or compare the exact bytes: a line containing apple and a line containing apple are different as far as ordinary uniq is concerned.
A non-zero status usually means a file could not be opened or the output could not be written. Check paths and permissions before reaching for sudo:
$ test -r input.txt && echo readable
readable
$ test -w . && echo current-directory-writable
current-directory-writable
Elevated privileges are not part of normal text filtering. If a source file belongs to another account, copy it into a controlled working directory through your normal access process instead of running a broad command as root.
Done means
- Know the core rule.
uniqcompares adjacent records, not the whole file. - Picked the right mode. You can choose between collapsing, counting, duplicate-only, unique-only and all-duplicate output.
- Checked sorting was safe. You verified whether reordering is acceptable before using
sort | uniqorsort -u. - Matched the input shape. You selected field, character, case or NUL comparison rules from the actual data, not a guess.
- Protected the original. You wrote output to a different path before replacing an existing file.
- Left the system alone. No service, permission, source file or persistent configuration was changed.