Half a CSV file vanishes because a redirect landed before cut had read a single byte, which is the first trap worth knowing before the second one. You will use cut to pull delimited fields or fixed byte and character positions out of text, while dodging the defaults that quietly produce the wrong answer. The examples match GNU coreutils 9.4, package version 9.4-3ubuntu6.3 here. Allow about ten minutes; none of the normal examples needs elevated privileges.
Safety checkpoint: cut only ever writes to standard output. That is great for previews and pipelines, but shell redirection can truncate an existing file before cut has read anything from it. Write to a new destination until you have checked the result.
Pick exactly one selection mode: -f for delimiter-separated fields, -c for characters, -b for bytes. Lists count from 1, not 0, and can mix single positions with inclusive ranges.
$ printf '%s\n' 'ada:admin:1001' | cut -d: -f1
ada
$ printf '%s\n' 'abcdef' | cut -c2-4
bcd
$ printf '%s\n' 'abcdef' | cut -b1,6
af
-3 means positions 1 through 3.Checkpoint: structured input with separators wants -f, never a guessed character position. Raw protocol or binary offsets want -b. Ordinary human text is usually clearer with -c.
Field mode expects a single delimiter through -d; leave it out and the delimiter defaults to TAB. It is a literal character, not a regular expression and not a multi-character string.
$ printf '%s\n' 'alice,engineering,active' 'bob,support,paused' | cut -d, -f1,3
alice,active
bob,paused
GNU cut reuses the input delimiter in field output by default. Reach for --output-delimiter when the consumer downstream needs something else:
$ printf '%s\n' 'alice,engineering,active' | cut -d, -f1,3 --output-delimiter='|'
alice|active
An empty field still counts as a field. In a,,c, field 2 is empty and field 3 is c. That matters the moment output feeds another command, since dropping an empty field would shift every position after it.
In field mode, a line with no delimiter passes through unchanged by default, even when the field you asked for does not exist on that line. This is the most common surprise in a file that mixes real records with comments or malformed lines.
$ printf '%s\n' 'alice,admin' 'unstructured line' | cut -d, -f2
admin
unstructured line
$ printf '%s\n' 'alice,admin' 'unstructured line' | cut -d, -f2 -s
admin
Add -s, also spelled --only-delimited, to drop any line without a delimiter. It will not drop a line merely because the selected field happens to be empty, which is the distinction that lets you reject malformed records without losing valid ones that carry an empty value.
Checkpoint: preview a representative sample without -s first. If a passed-through line could be mistaken for real output, add -s and count or log what gets rejected separately.
With no file argument, or with - as the argument, cut reads standard input, which is what makes it fit naturally into a pipeline:
$ printf '%s\n' 'ada:1001' 'grace:1002' | cut -d: -f1
ada
grace
Newline is the record separator by default. For filenames or other data that may contain newlines themselves, add -z so NUL becomes the record separator instead. The producer earlier in the pipeline has to emit NULs too; switching only the consumer does not convert ordinary newline input.
$ printf '%s\0' 'alpha:one' 'beta:two' | cut -z -d: -f1 | od -An -t x1
61 6c 70 68 61 00 62 65 74 61 00
Those trailing NUL bytes in the dump are expected. Keep this kind of data inside a NUL-safe pipeline rather than a shell variable, since shell variables cannot hold NUL characters at all.
-c selects characters, -b selects bytes, and the two can disagree on multibyte text because the active locale governs how characters are read. Check the locale before trusting positions in non-ASCII input:
$ printf '%s\n' 'café' | LC_ALL=C cut -c1
c
$ printf '%s\n' 'café' | LC_ALL=C.UTF-8 cut -c1
c
The first character happens to match in both cases; the difference only shows once the selected range reaches a multibyte character. Use -b only when byte offsets are genuinely the format's rule, and -c with a known locale when the requirement is really about characters. -n is not a useful switch here: GNU coreutils 9.4 accepts it but documents it as ignored.
Look at a pipeline's output first, then write it to a new file:
$ cut -d, -f1,3 records.csv > records-selected.csv.new
$ test -s records-selected.csv.new && mv records-selected.csv.new records-selected.csv
That test -s check only stops an empty result overwriting the old file; it cannot prove the content is right. Compare a sample, or run diff --unified=3, before replacing anything that matters. If the command fails, remove the named temporary with rm records-selected.csv.new once you have confirmed it really is the failed output; the original is still sitting there untouched.
Do not reach for sudo just because a pipeline happens to read a system-owned file. Use elevated privileges only when permissions genuinely demand it, and write the result somewhere you can check ownership and permissions afterwards.