Split Files by Lines, Bytes or Records with GNU split
You will finish with a repeatable way to divide a large file into smaller pieces, choose readable output names, and check that the pieces can be joined or processed as intended. The examples use GNU split from coreutils 9.4, installed here as package version 9.4-3ubuntu6.3.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about fifteen minutes. You need a shell, a readable input file, and enough free space for the output pieces. The normal examples do not need sudo. They write new files in the current directory, so choose a working directory deliberately.
1. Check the installed command
Confirm the implementation and read the local option list before putting a command into a script:
$ split --version
split (GNU coreutils) 9.4
$ dpkg-query -W -f='${Package} ${Version}\n' coreutils
coreutils 9.4-3ubuntu6.3
$ split --help
GNU split reads a named file, or standard input when the file is omitted or written as -. By default it writes 1,000 lines per piece, using the prefix x and two-letter suffixes such as xaa and xab.
Checkpoint: identify the binary that will run in a script or terminal:
$ command -v split
/usr/bin/split
2. Split a text file by lines
Make a small test input so the result is easy to inspect. This changes only files in the current directory:
$ mkdir -p split-work
$ cd split-work
$ printf 'alpha\nbeta\ngamma\ndelta\nepsilon\nzeta\n' > records.txt
$ split --lines=2 records.txt part-
$ ls -1 part-*
part-aa
part-ab
part-ac
--lines=2 puts two newline-terminated records in each output file. The final piece may contain fewer records. The last argument, part-, is the output prefix. Without it, the files would be named xaa, xab and so on.
Inspect the pieces and count their lines:
$ for file in part-*; do printf '%s: ' "$file"; wc -l < "$file"; done
part-aa: 2
part-ab: 2
part-ac: 2
$ cat part-aa part-ab part-ac
alpha
beta
gamma
delta
epsilon
zeta
Shell glob order matters when reassembling alphabetic suffixes. The command above works because these suffixes sort in their creation order. Keep the original input until the pieces have been checked.
3. Prevent accidental overwrites
split opens its output names for writing. A matching name from an earlier run can therefore be replaced. Use a new prefix or check for existing files before starting:
$ if compgen -G 'batch-*' >/dev/null; then
> printf 'refusing to overwrite existing batch-* files\n' >&2
> exit 1
> fi
$ split --lines=2 records.txt batch-
$ ls -1 batch-*
batch-aa
batch-ab
batch-ac
This guard is especially useful in automated jobs. It does not protect against every naming scheme, so review the prefix and directory as part of the job configuration. Do not add sudo to solve a naming mistake. Elevated access can make an accidental overwrite harder to recover.
Recovery checkpoint: the generated files are ordinary files. If this test run is no longer needed, remove only the known test outputs:
$ rm -- batch-aa batch-ab batch-ac part-aa part-ab part-ac records.txt
$ cd ..
$ rmdir split-work
That deletion is irreversible unless another copy exists. For valuable data, move the pieces to a quarantine directory or keep them until the consumer has verified them.
4. Use numeric suffixes for scripts
Alphabetic suffixes continue from aa through zz and then need more suffix space. Numeric names are often easier for another program to sort and parse:
$ printf 'one\ntwo\nthree\nfour\nfive\n' > input.txt
$ split --lines=2 --numeric-suffixes=1 --suffix-length=3 input.txt chunk-
$ ls -1 chunk-*
chunk-001
chunk-002
chunk-003
--numeric-suffixes=1 selects numeric suffixes and starts at 1. --suffix-length=3 makes the width predictable. The short form -d selects numeric suffixes starting at 0, while the long form can also accept a starting value. Use the long forms in scheduled jobs when clarity matters.
5. Split by bytes without confusing units
Use --bytes when the pieces need an approximate or exact byte limit rather than a line limit:
$ split --bytes=1K --numeric-suffixes=1 --suffix-length=2 \
large.bin binary-
$ wc -c binary-*
1024 binary-01
1024 binary-02
In GNU split, K means 1,024 bytes. Decimal units such as KB mean 1,000 bytes; binary spellings such as KiB are also accepted. The final file is shorter when the input size is not an exact multiple. Do not use byte splitting for a format that must remain record-aligned.
For text or another record-based format, --line-bytes=SIZE puts at most that many bytes of records in each output file without cutting a record. A single record longer than the limit can still make a piece exceed it. Verify this behaviour with the format's actual record boundaries rather than relying only on file sizes.
6. Keep records together or distribute them
The --number option divides input into a chosen number of outputs. Its mode controls whether records are preserved and whether they are assigned sequentially or round-robin:
$ split --number=l/3 --numeric-suffixes=1 --suffix-length=2 \
records.txt group-
$ for file in group-*; do printf '%s: ' "$file"; wc -l < "$file"; done
group-01: 2
group-02: 2
group-03: 1
l/3 makes three pieces without splitting lines or records. The plain form 3 divides by input size, and r/3 distributes records round-robin. The --elide-empty-files option suppresses empty files when a numbered mode would otherwise create them. Treat the mode as part of the data contract: a consumer expecting sequential batches will not get the same data order from round-robin output.
7. Read standard input and verify a reassembly
Omit the input file or use - to read a pipe. Give the output prefix explicitly so the generated names are obvious:
$ printf 'red\ngreen\nblue\nyellow\n' | split --lines=2 - stdin-
$ cat stdin-aa stdin-ab
red
green
blue
yellow
For a byte-for-byte check, calculate a digest before splitting, concatenate the pieces in the intended order, then compare the digest. This example creates a temporary reassembly and leaves the original input untouched:
$ sha256sum input.txt
<hash> input.txt
$ cat chunk-001 chunk-002 chunk-003 > reassembled.txt
$ sha256sum reassembled.txt
<hash> reassembled.txt
$ cmp --silent input.txt reassembled.txt && echo 'reassembly matches'
reassembly matches
The displayed hash is a placeholder for the real value printed by your machine. If cmp reports a difference, stop before deleting anything. Check the suffix order, the selected split mode, and whether the input ended with a newline.
8. Diagnose common failures
An error saying that the input cannot be opened usually means the path, permissions or current directory are wrong. Check without changing the file:
$ pwd
$ ls -l /path/to/input.txt
$ test -r /path/to/input.txt && echo readable
An error about an invalid size, line count or suffix length means the argument was not accepted. Check the option spelling and value with split --help. If an output file already exists, change the prefix or move the old output out of the way after confirming it is safe to do so.
For a service or scheduled job, write to a dedicated directory and use a temporary prefix until verification is complete. split itself does not encrypt, compress, checksum or delete the source. Those are separate operations with separate failure and recovery paths.
Done means
- You confirmed the local GNU coreutils version and the input path.
- You chose line, byte, line-byte or numbered splitting to match the data format.
- You selected an explicit prefix and suffix style that the consumer can sort.
- You checked the generated names and contents before removing the source or pieces.
- You know that
Kis 1,024 bytes and that record-preserving modes can exceed a byte limit for one long record. - You have a digest or
cmpcheck when reassembly must be exact.