Join Sorted Files Safely with GNU join
You will finish with a repeatable way to combine two sorted text files on a shared field, select the fields that appear in the result, and find records that do not match. The examples use GNU coreutils 9.4, the version installed on this machine.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about fifteen minutes. You need a shell, join, sort, and permission to read the input files and write a temporary working directory. The commands below do not edit either input file. No elevated privileges are needed.
1. Check the installed command
Confirm that the command in your PATH is the GNU implementation and record its version:
$ command -v join
/usr/bin/join
$ join --version | head -n 1
join (GNU coreutils) 9.4
Your path may differ. The important check is that the command reports GNU coreutils 9.4 or another version whose local manual you have checked. This guide uses options documented by that installed command.
Checkpoint
You have identified the executable and know which version's behaviour you are testing.
2. Create two small, sorted inputs
Use a temporary directory for the demonstration. The first file contains customer names and the second contains account status. The first field is the shared key in both files:
$ workdir=$(mktemp -d)
$ printf '%s\n' \
'1001 Alice' \
'1002 Ben' \
'1004 Dana' > "$workdir/people.txt"
$ printf '%s\n' \
'1001 active' \
'1003 pending' \
'1004 paused' > "$workdir/status.txt"
$ cat "$workdir/people.txt"
1001 Alice
1002 Ben
1004 Dana
$ cat "$workdir/status.txt"
1001 active
1003 pending
1004 paused
These files are already sorted on their first field. Keep the key spelling and comparison rules consistent between the files. By default, join treats leading blanks as separators and ignores them, then uses the first field as the join field.
Safety note: the shell redirection operator > truncates an existing destination. The names above are inside a new temporary directory. Do not point them at a valuable file unless overwriting it is intentional.
3. Perform the default join
Pass the people file first and the status file second:
$ join "$workdir/people.txt" "$workdir/status.txt"
1001 Alice active
1004 Dana paused
Only keys present in both files are printed. The default output begins with the join field, followed by the non-key fields from file 1 and then file 2. The unmatched 1002 and 1003 records are omitted.
There is no output for an unmatched pair, but that is not an error. Check the status explicitly when a script needs to distinguish an empty result from a failed command:
$ join "$workdir/people.txt" "$workdir/status.txt" > "$workdir/joined.txt"
$ test -s "$workdir/joined.txt" && echo 'matched records written'
matched records written
4. Select fields from each file
Use -1 and -2 when the shared key is not in the same position. Use -o to specify the output order. Format items use FILENUM.FIELD; field numbers start at 1.
For the current files, output just the name, key, and status in that order:
$ join -o 1.2,0,2.2 "$workdir/people.txt" "$workdir/status.txt"
Alice 1001 active
Dana 1004 paused
Here, 1.2 is the name from file 1, 0 is the common join field, and 2.2 is the status from file 2. The equivalent -j FIELD option sets both join fields to the same field number, but it does not by itself change output order.
If file 1 has a name before its key, for example Alice 1001, and file 2 still has 1001 active, the command would use -1 2 -2 1. Both files must still be sorted according to their selected join fields.
5. Use a delimiter for CSV-like records
For files separated by a character rather than whitespace, pass that character with -t. This example uses a tab, represented by the shell's $'...' quoting:
$ printf '1001\tAlice\n1002\tBen\n1004\tDana\n' > "$workdir/people.tsv"
$ printf '1001\tactive\n1003\tpending\n1004\tpaused\n' > "$workdir/status.tsv"
$ join -t $'\t' -o 1.2,0,2.2 "$workdir/people.tsv" "$workdir/status.tsv"
Alice 1001 active
Dana 1004 paused
With -t, fields are separated by that exact character rather than by runs of blanks. Do not mix a whitespace join with a delimiter-aware sort: the files must be sorted using the same field definition that join uses.
6. Keep or inspect unmatched records
Use -a 1 or -a 2 to include unpairable lines from one file. Missing fields are empty by default. Add -e when a visible marker is more useful:
$ join -a 1 -a 2 -e '<missing>' -o 0,1.2,2.2 \
"$workdir/people.txt" "$workdir/status.txt"
1001 Alice active
1002 Ben <missing>
1003 <missing> pending
1004 Dana paused
To output only the rows that have no partner, use -v 1 for file 1 or -v 2 for file 2:
$ join -v 1 "$workdir/people.txt" "$workdir/status.txt"
1002 Ben
$ join -v 2 "$workdir/people.txt" "$workdir/status.txt"
1003 pending
This is useful for reconciliation. Treat the result as a report until you have checked why the keys differ; a spelling, case, padding, or sort-order mistake can look like missing data.
7. Diagnose sorting and locale problems
The inputs must be sorted on their join fields. GNU join compares according to LC_COLLATE, so a file sorted under one locale may not be correctly ordered under another. Make the locale explicit and use the matching sort command:
$ LC_ALL=C sort -k 1b,1 "$workdir/people.txt" > "$workdir/people.sorted"
$ LC_ALL=C sort -k 1b,1 "$workdir/status.txt" > "$workdir/status.sorted"
$ LC_ALL=C join "$workdir/people.sorted" "$workdir/status.sorted"
1001 Alice active
1004 Dana paused
The b in -k 1b,1 makes leading blanks part of the field definition used by the default whitespace join. For a delimiter-aware join, use the same delimiter with sort -t and the appropriate key specification.
If an input is not sorted and some lines cannot be joined, join can warn about the order. --check-order asks it to check even when every line happens to be pairable. Do not silence a warning with --nocheck-order until you have established that the ordering is deliberate and safe.
8. Clean up the demonstration
When you have finished checking the output, remove only the temporary directory you created:
$ rm -rf -- "$workdir"
$ test ! -e "$workdir" && echo 'temporary files removed'
temporary files removed
This cleanup is irreversible for those temporary files. Do not substitute a broad path or an unreviewed variable. The original source files used in the examples were never modified.
Done means
- You confirmed the installed GNU
joinversion. - You joined two files sorted on their shared field and checked the matched output.
- You can select fields with
-o, change key positions with-1and-2, and set a delimiter with-t. - You can report unmatched rows with
-a,-v, and an explicit missing-field marker. - You understand that locale and field-separator choices must match between
sortandjoin. - You removed only the temporary demonstration directory and left source data unchanged.