Split a Text File at Matching Lines with csplit
You will finish with separate files for each marked section of a text file, with predictable names and a quick way to check that no content disappeared. The examples use GNU csplit from coreutils 9.4, installed here as package version 9.4-3ubuntu6.3.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about ten minutes. You need a shell and a readable text file whose sections have a recognisable marker line. The commands normally run as an ordinary user. No elevated privileges are needed unless your input or output directory is deliberately restricted.
1. Inspect the input before splitting it
Start by viewing the file and checking the command version. These are read-only commands:
$ sed -n '1,30p' /path/to/report.txt
$ csplit --version
csplit (GNU coreutils) 9.4
This guide assumes markers such as # alpha appear at the start of a line. Replace the path and marker format with values from your file. Do not guess the pattern: a regular expression that matches ordinary content can create more pieces than you intended.
Checkpoint: identify the marker lines and decide whether each marker belongs in the section that follows. By default, csplit keeps a matching line at the start of the new piece.
2. Split at every marker
Use a slash-delimited regular expression followed by {*}. The repetition means to apply the preceding pattern as many times as possible:
$ csplit --prefix=part- --suffix-format='%02d.txt' \
/path/to/report.txt '/^# /' '{*}'
0
12
11
14
csplit prints the byte count written for each output file. With the default naming scheme, the files would be xx00, xx01 and so on. This command instead creates part-00.txt, part-01.txt and later files. The first zero-byte file is expected when the input begins with a marker: splitting happens before the first matching line.
Inspect the result before using it downstream:
$ for file in part-*.txt; do
printf '%s: ' "$file"
wc -c < "$file"
done
$ sed -n '1,5p' part-01.txt
# alpha
one
Output order follows the input order. A section containing a marker and no body is still a real piece unless you choose the empty-file option later.
3. Remove marker lines when they are only separators
Some consumers need the section contents but not their headers. Add --suppress-matched to omit the line that triggered each split:
$ csplit --prefix=section- --suffix-format='%02d' \
--suppress-matched /path/to/report.txt '/^# /' '{*}'
4
4
6
$ sed -n '1,5p' section-00
one
The pattern still determines the boundaries; only the matching lines are left out. This is easy to confuse with the percent form of a pattern. A slash pattern copies up to a match. A percent pattern skips up to a match and is useful when discarding a preamble, but it does not by itself mean "remove every marker".
4. Avoid an unwanted empty first file
If a marker occurs on the first line, the initial split creates an empty piece. Add --elide-empty-files when empty output files have no useful meaning:
$ csplit --prefix=section- --suffix-format='%02d' \
--suppress-matched --elide-empty-files \
/path/to/report.txt '/^# /' '{*}'
4
4
6
$ test ! -e section-00 && echo 'no empty section written'
no empty section written
Do not use this option merely to hide a surprising result. First confirm that the input really starts with a marker and that an empty section is not meaningful to the program consuming these files.
5. Split once at a known line number
Patterns do not have to be regular expressions. An integer copies up to, but not including, that line number. For example, this divides after line 20:
$ csplit --prefix=page- --suffix-format='%02d.txt' \
/path/to/report.txt 20
20
...
The second count depends on the input size, so the displayed ellipsis is not literal output. Line numbers count from one. Use wc -l first if the boundary is calculated from a known file layout:
$ wc -l /path/to/report.txt
48 /path/to/report.txt
For a moving boundary, a regular expression is usually safer than a hard-coded line number.
6. Use offsets for a nearby boundary
A regular-expression pattern can have an integer offset, optionally preceded by + or -. This example places the split after the line immediately following the beta marker:
$ csplit --prefix=off- --suffix-format='%02d' \
/path/to/report.txt '/^# beta/+1'
19
18
Offsets are line positions relative to the match. Check a small input by hand before applying one to a large archive, especially when markers can occur close together. A negative offset can move the boundary before the matching line, but it is more likely to surprise a later reader of the command.
7. Keep scripts quiet and failures recoverable
Use --quiet when byte counts would interfere with a script's output. Use --digits when the number of output files needs more than two digits, for example --digits=4. The suffix format is controlled separately by --suffix-format; it is an sprintf format containing the numeric index.
csplit can remove output files when an error occurs. Add --keep-files if partial files are useful for diagnosis:
$ csplit --keep-files --prefix=debug- \
/path/to/report.txt '/^# /' '{*}'
$ printf 'exit status: %s\n' "$?"
exit status: 0
Do not treat --keep-files as a substitute for a clean destination. Existing names can collide with earlier runs, and redirection or later processing may overwrite useful data. Split into a new directory, or choose a unique prefix. If a run fails, keep the partial files until you have inspected them, then remove only the known temporary outputs with an explicit path.
There is no undo operation for files csplit has already written. Recovery is straightforward if the original remains untouched: delete the generated pieces and rerun into a clean directory. Do not delete the source file as part of a batch command.
Done means
- The installed csplit version and marker format were checked first.
- Each output piece has the intended boundary, and marker lines were either retained or suppressed deliberately.
- Empty leading output was accepted or removed with
--elide-empty-filesfor a stated reason. - Names, suffixes and digit width are predictable for the next command in the workflow.
- Byte counts, file listings or sample contents confirm the split, and the original input remains available for recovery.