Read Git's Chunk Tables Without Guessing at Offsets
By the end of this guide you will be able to recognise the table of contents in Git's chunk-based files, calculate where each payload ends, and verify the two common users of the format: commit-graph files and multi-pack indexes. The examples match the installed Git 2.43.0 manpages.
The route
Jump straight to the step you need, or tick off Done means at the end.
Allow about 15 minutes. You need Git 2.43.0 or a nearby release, a repository whose metadata you can inspect, and ordinary read access to its .git directory. None of the commands below needs sudo. Work on a disposable clone if you plan to experiment with files rather than only verify them.
1. Identify what the format is for
A chunk-based file is a container for several contiguous data sections. Git uses the shared format for commit-graph files and for the multi-pack-index, commonly abbreviated MIDX. The file-specific header comes first. It identifies the format, its version and the number of chunks, then the chunk table follows.
This is an on-disk format, not a command that you normally invoke directly. The useful operational commands are the writers and verifiers that understand the individual formats:
git commit-graph write
git commit-graph verify
git multi-pack-index write
git multi-pack-index verify
Checkpoint: if you only need to check repository metadata, stop here and use the relevant verify command. You do not need to decode raw bytes to diagnose every problem.
2. Read the table of contents
The chunk region contains one 12-byte row for each chunk, followed by one final row. Every row has a four-byte chunk identifier and an eight-byte offset. Integers use network-byte order, which means big-endian byte order.
Chunk ID (4 bytes) | Chunk Offset (8 bytes)
ID[0] | OFFSET[0]
... | ...
ID[C] | OFFSET[C]
0x0000 | OFFSET[C+1]
The identifier names the data beginning at its row's offset. It ends immediately before the next row's offset. Therefore chunk i has a size of OFFSET[i+1] - OFFSET[i]. The table requires payloads to appear contiguously and in the same order as their rows.
The last identifier is four zero bytes. That terminator tells a reader that the table has ended and supplies the offset immediately after the final chunk. A valid file also has at least a trailing hash after that offset. The hash is outside the chunk payload region, so do not treat it as another chunk.
Common decoding mistakes are predictable. Reading offsets as little-endian produces absurd positions. Treating the terminator as a payload identifier loses the final boundary. Assuming the last chunk extends to end of file includes the trailing hash and gives the wrong size.
3. Verify a commit-graph
From the root of a repository, write a graph for commits found in its packfiles, then ask Git to check it against the object database:
git commit-graph write
git commit-graph verify
A successful command returns to the shell without an error. With a terminal, Git may show progress while writing. The verifier checks the graph contents rather than merely checking that a file exists.
The graph is normally stored under the repository's object directory. A repository can also use a split graph chain, where several graph files live in an info/commit-graphs directory. To verify only the tip of such a chain, use the documented shallow check:
git commit-graph verify --shallow
Use the full verification when you are investigating corruption or a chain problem. --shallow is a narrower check, not a faster substitute for validating every layer.
Checkpoint: run git commit-graph verify after a write. If the write reports success but no useful graph appears, check whether core.commitGraph is disabled. The Git 2.43.0 documentation says a write can warn and return success without writing when that configuration is disabled.
4. Verify a multi-pack-index
A MIDX indexes objects across several packfiles instead of requiring one index per lookup path. Create or refresh it, then verify it:
git multi-pack-index write
git multi-pack-index verify
The current repository's packfiles are used by default. The MIDX is kept with them, under the object directory's packs area. If your repository uses an alternate object directory, pass that directory only when it is a known alternate for the current repository:
git multi-pack-index --object-dir /path/to/alternate-objects write
git multi-pack-index --object-dir /path/to/alternate-objects verify
Replace the placeholder with a real alternate object directory. The command expects its packs subdirectory to contain the pack and pack-index pairs. A random directory is not a safe test and will fail the alternate-object check.
A MIDX can also contain a bitmap when written with --bitmap. That adds another indexed data section, but it does not change the chunk-table rule: each section still gets a row, offsets remain big-endian, and the zero identifier still terminates the table.
5. Inspect bytes only when verification is not enough
Raw inspection is useful when developing a parser or explaining a damaged file. First identify the exact file Git is using. For a normal repository, list the metadata without changing it:
git rev-parse --git-path objects/info/commit-graph
git rev-parse --git-path objects/info/commit-graphs
git rev-parse --git-path objects/pack/multi-pack-index
Some paths may not exist, especially before the corresponding writer has run. Do not infer validity from the pathname alone. Run the relevant verifier and inspect the file only if you need byte-level evidence.
For a parser, read the format-specific header first, obtain its chunk count and the start of the chunk region, then read exactly C + 1 rows of 12 bytes. Decode each identifier as four bytes and each offset as an unsigned eight-byte big-endian integer. Check that offsets are ordered, lie within the mapped file, and leave room for the trailing hash. Reject a missing zero identifier, overlapping ranges and an offset that points into the table.
Git's C API follows the same lifecycle. A writer initialises a struct chunkfile with init_chunkfile(), registers each payload with add_chunk(), and calls write_chunkfile(). That write checks that each callback produced the expected amount of data. The caller then frees the chunk state, writes the trailing hash and closes the hashfile.
A reader memory-maps the complete file, initialises the state with init_chunkfile(NULL), and calls read_table_of_contents(). pair_chunk() returns a pointer to a present chunk without changing the pointer when the identifier is absent. read_chunk() instead calls a callback with the chunk's pointer and size, and does not call it for a missing chunk. Keep the mapping alive while those pointers are in use, then call free_chunkfile(); freeing the parser state does not unmap the file.
6. Recover from a failed or unwanted write
These writers change derived files inside .git, not tracked working-tree content. If a write fails, leave the files in place and run the matching verifier first. It may show whether the existing metadata is still usable. Do not delete packfiles or graph layers as a first response: pack maintenance can discard objects, and deletion is not required to understand the table format.
If you deliberately generated metadata in a disposable clone and need to return that clone to its previous state, remove only the exact generated file or directory after confirming the path with git rev-parse --git-path. This is an irreversible filesystem action, so take a backup or reclone instead when the repository contains valuable local state. A later git commit-graph write or git multi-pack-index write can recreate derived indexes from the object database, provided the underlying objects remain intact.
Done means
- You can locate the chunk region after the format-specific header.
- You can calculate a chunk size as the next offset minus the current offset.
- You recognise the four zero bytes as the table terminator, not a payload.
- You have run the relevant commit-graph or MIDX verifier.
- You have kept the trailing hash outside the final chunk range.