Inspect Git Objects Safely with git cat-file

git cat-file reads a commit or file exactly as Git stored it, bypassing whatever your working tree currently shows. In about 15 minutes you will identify an object, read it in a useful form, check whether it exists, and inspect many objects at once in a script. The examples use Git 2.43.0, installed with the git-man package on the machine used for this guide.

You need a Git repository and ordinary read access to it. No elevated privileges are required.

What you are inspecting

Git stores commits, trees, blobs and annotated tags in an object database. A commit points to a tree, a tree names files and subdirectories, and a blob contains file content. git cat-file reads those objects, rather than reading the current working-tree copy.

The most common distraction is the word 'file' in the command name. A blob has no filename by itself. A name such as HEAD:README.md combines a revision with a path and resolves to the blob currently recorded by that commit. Use a normal shell command such as cat README.md when you mean the working tree.

1. Choose a repository and a reference

Change to the repository you want to inspect, then choose a stable reference. HEAD is convenient for the current commit; a full or abbreviated object ID also works.

cd /path/to/repository
git rev-parse --verify HEAD

Expected output is a 40-character object ID for a SHA-1 repository, or the repository's configured object-format length in another format. If this fails, you are not in a repository or HEAD is not resolvable. That is a location or reference problem, not a reason to run the command as root.

Checkpoint: You have a repository path and a revision that Git can resolve.

2. Check type, size and existence

Ask for metadata before asking for content. The -t option prints the object type and -s prints its uncompressed size in bytes.

git cat-file -t HEAD
git cat-file -s HEAD
git cat-file -e HEAD

A normal commit produces output similar to this; the exact size depends on the repository:

commit
238

-e deliberately prints nothing. Its exit status is the useful result, which makes it suitable for a guard in a script:

if git cat-file -e 'HEAD:README.md'; then
    echo 'README.md exists in HEAD'
else
    echo 'README.md is not recorded in HEAD'
fi

An invalid object name returns a non-zero status. Do not treat a successful check as proof that a pathname is safe to use elsewhere: it only says that Git resolved a valid object in this repository.

Checkpoint: You can identify an object and test it without dumping its contents.

3. Read a commit or a recorded file

Use -p for type-aware, human-readable output. A commit becomes a readable header and message; a tree becomes a directory-like listing; a blob is emitted as its content.

git cat-file -p HEAD
git cat-file -p 'HEAD:README.md'

For a commit, expect fields such as tree, parent, author and committer, followed by the commit message. For a text blob, the second command prints the version recorded in HEAD, even if you have edited README.md locally.

If you know the actual type, supplying it returns the raw, uncompressed object contents:

git cat-file commit HEAD
git cat-file blob 'HEAD:README.md'

Use the type form when another program needs the stored bytes. Use -p when a person needs a useful representation. Do not redirect an unknown blob to a filename you care about without checking its contents first; Git objects can contain arbitrary data.

4. Inspect many objects with batch mode

Batch mode reads object names from standard input. --batch-check returns object ID, type and size without emitting content, so it is the safer and cheaper choice for an inventory.

printf '%s\n' HEAD 'HEAD:README.md' missing-object |
    git cat-file --batch-check

Each successful line has this shape:

<object-id> commit <size>
<object-id> blob <size>
missing-object missing

The literal input is retained for a missing object. An ambiguous short object ID is reported as ambiguous. Parse the status field rather than assuming every input produced a three-column success line.

Use --batch when you also need content. Its header is followed by exactly the number of content bytes in the size field and then a newline. This matters for binary data: do not parse a batch stream by looking for the next newline.

printf '%s\n' 'HEAD:README.md' |
    git cat-file --batch

For a program that needs explicit commands, --batch-command accepts info, contents and flush. Add --buffer only when you want to batch requests before output; then send flush or the process may appear to hang while waiting for work to be released.

5. Handle paths and filters deliberately

--textconv and --filters transform content using the repository's configured drivers. They require a path-aware object expression such as HEAD:path/to/file, because Git needs to know which path's filter applies. These options can execute configured filter commands, so treat a repository's configuration as code and review it before using them on untrusted material.

--follow-symlinks is available with batch modes when a tree path names a symlink. It can report missing, dangling, loop or notdir, and a link pointing outside the tree is reported as a symlink target rather than silently treated as a regular file. When output may contain newlines, use -Z so input and output are NUL-delimited. The older -z form only changes input and is deprecated in the installed manual.

These options inspect or transform data; they do not amend commits, reset files or change refs. The main safety boundary is your choice of filters and what you do with the bytes afterwards.

Common failure modes

Done means