Home / Alt manpages / git-hash-object(1)

  • git-hash-object(1)
  • User command
  • linux

Hash Files and Store Git Objects Safely

You will use git hash-object to calculate the object ID for a file, optionally store that object in a repository, and verify what Git wrote. The examples describe Git 2.43.0, installed here from git-man and git package version 1:2.43.0-1ubuntu7.3.

Allow about ten minutes. You need Git and a readable input file. The ordinary hashing examples are unprivileged and do not change the input. The -w example changes the selected repository's object database, so check the repository path before running it. There is no need for sudo.

1. Check the installed command

Confirm which executable and version your shell will use:

$ command -v git
/usr/bin/git
$ git --version
git version 2.43.0

The command is written as git hash-object, with a space, even though its manual page is named git-hash-object(1). Unless you supply another type, Git treats file content as a blob.

2. Calculate an ID without storing anything

Create or select a file whose contents you want to identify. This example writes a new file in a temporary directory:

$ mkdir -p /tmp/git-hash-demo
$ printf '%s\n' 'hello Git' > /tmp/git-hash-demo/input.txt
$ git hash-object /tmp/git-hash-demo/input.txt
7b5bbd989152e5bab6b5476f50133e16137d6b30

The 40-character result is the object ID for a blob containing the file's bytes. Without -w, Git reports the ID but does not add the object to the repository database. The file can be outside the current work tree, and a repository is not required for this read-only calculation.

Checkpoint: hash the same bytes through standard input:

$ printf '%s\n' 'hello Git' | git hash-object --stdin
7b5bbd989152e5bab6b5476f50133e16137d6b30

The IDs match because Git hashes the content, not the filename. A changed byte, including a changed final newline, produces a different ID.

3. Store the object in a repository

Only do this after checking the target repository. The write is not destructive to tracked files, but it does add an object to that repository's object database:

$ git -C /tmp/git-hash-demo init
$ git -C /tmp/git-hash-demo hash-object -w /tmp/git-hash-demo/input.txt
7b5bbd989152e5bab6b5476f50133e16137d6b30

The -w option means "write". It stores the object, but it does not create a branch, index entry, commit or working-tree file. Git may already have the object, in which case writing the same content simply leaves the same object available.

To verify the stored object, capture its ID and ask Git about its type and size:

$ oid=$(git -C /tmp/git-hash-demo hash-object /tmp/git-hash-demo/input.txt)
$ git -C /tmp/git-hash-demo cat-file -t "$oid"
blob
$ git -C /tmp/git-hash-demo cat-file -s "$oid"
10

The size is ten bytes for hello Git followed by a newline. cat-file is the verification step here; an ID printed by a command is not, by itself, proof that the object was stored.

4. Hash several paths from standard input

Use --stdin-paths when a list of filenames is easier to produce than a long command line. It reads one path per line and prints one object ID per input path:

$ printf '%s\n' /tmp/git-hash-demo/input.txt | git hash-object --stdin-paths
7b5bbd989152e5bab6b5476f50133e16137d6b30

Keep the input list unambiguous. This mode is line-oriented, so a filename containing a newline cannot be represented as one ordinary line. For a small, known list, passing quoted paths directly is easier to review.

5. Understand filters and exact bytes

By default, Git hashes a file's contents as a blob. Attributes can affect the content when you ask Git to treat the file as if it were at a particular path. The --path=<file> option supplies that path so Git can select applicable filters, which matters especially for temporary files or standard input.

Use --no-filters when the exact bytes supplied must be hashed without attribute-based input filtering:

$ git hash-object --no-filters /tmp/git-hash-demo/input.txt
7b5bbd989152e5bab6b5476f50133e16137d6b30

For standard input, --no-filters is implied unless you also give --path. Do not assume that a line-ending conversion or other filter has been applied merely because the file is in a Git work tree. If the ID is being used to compare exact bytes, state the filtering choice explicitly in your script.

6. Keep object types and dangerous options separate

The default type is blob. Git also accepts commit, tree and tag with -t, but those types have structured formats. Do not label arbitrary file data as another type unless you are deliberately constructing and testing a valid Git object.

--literally permits --stdin to write data that normal object parsing or git fsck might reject. That is for Git stress tests and reproducing corrupt objects. It is security-sensitive repository manipulation, not a repair shortcut. Avoid it in normal scripts, and never combine unfamiliar input with -w in a valuable repository.

Done means

  • The installed Git version and target repository were checked.
  • The expected blob ID was calculated without changing the input.
  • Standard input or --stdin-paths was used only when its line-oriented input was suitable.
  • -w was used only for an intentional repository object write.
  • cat-file -t and cat-file -s verified the stored object's type and byte count.
  • Filtering was chosen deliberately, and --literally was left for controlled Git testing.