Git Rev News: Edition 139 (September 30th, 2026)

Welcome to the 139th edition of Git Rev News, a digest of all things Git. For our goals, the archives, the way we work, and how to contribute or to subscribe, see the Git Rev News page on git.github.io.

This edition covers what happened during the months of August and September 2026.

Discussions

General

Reviews

Johannes Schindelin, alias Dscho, sent a patch to the mailing list to fix a performance regression that appeared in Git 2.53 and that affects repositories containing a very large number of packfiles.

In the commit message, Dscho explained that since 589127caa7 (packfile: move list of packs into the packfile store, 2025-10-30), the packfile_store_add_pack() function calls packfile_list_remove_internal() to check whether a packfile is already in the list of packs, and, if so, to move it to the end of that list. As this check scans the whole list linearly before every insertion, loading N new packs has an O(N²) complexity.

In a case reported by a Microsoft Git user in a GitHub issue, N was 37,815, and a simple git rev-parse --short HEAD, which is regularly run by GIT_PS1 to display the current commit in the shell prompt, went from 0.4 seconds to 4.5 seconds. Dscho also reported that, in a heavily exercised CI scenario, clone times went from under two minutes to over half an hour.

The fix consisted of adding a fast path for packfiles known to be new. Dscho anticipated that readers might wonder why the check was not simply removed, since packfile_list_append() had only one caller left, which always passes new packs. He explained that there used to be a second caller in prepare_midx() that needed the check, but that it was removed by 6aff1f25a0 (packfile: always add packfiles to MRU when adding a pack, 2025-10-30). As the function is declared in a header file, he preferred to extend its signature with an is_new parameter, to avoid problems with in-flight topics or downstream callers.

The patch also added an “abbreviate with 10,000 packs” test, running git rev-parse --short HEAD, to the t/perf/p5303-many-packs.sh performance test script.

Some background

Git stores objects either as individual “loose” files or in packfiles. Each fetch or push usually creates a new packfile, and maintenance tasks like git gc or git maintenance regularly consolidate them into fewer packs. When maintenance doesn’t run, or doesn’t complete, packfiles can accumulate.

To look up objects, Git keeps an in-memory list of the packfiles it knows about. It also reorders that list to implement a “most recently used” (MRU) optimization: the pack where an object was last found is moved to the front, as the next object being looked up is likely to be in the same pack.

Commit 589127caa7 was part of Patrick Steinhardt’s work on refactoring the object database, so that different storage backends can eventually be plugged in. It moved the list of packs into a new “packfile store” structure.

Review of the first version

Junio Hamano, the Git maintainer, replied to the patch with a rolling eyes emoji, noting that “As we grow older, more and more extreme use cases that we initially thought were simply crazy become reality.” He agreed that, as long as the caller knows that a pack is new, there is no reason to walk through all the packs trying to remove it, and he found the fix “Clever and clean.”

Jeff King, alias Peff, pointed out that this was a regression of a problem that had already been dealt with by ec48540fe8 (packfile.c: speed up loading lots of packfiles, 2019-11-27). He showed that the regression could even be seen in Git’s existing performance test suite, as the “load 10,000 packs” test went from 0.13 to 0.45 seconds at commit 589127caa7, a 246% increase. Unfortunately, he noted, nobody pays close attention to the perf suite, partly because “it’s clunky and expensive to run”, and partly because deciding whether a change is real or just noise often requires human judgment.

Peff found the fix reasonable, but wondered what value the new perf test added, as it showed the same slowdown as the existing “load 10,000 packs” test.

Patrick replied that GitLab had set up continuous benchmarking with Bencher. But recent changes to their CI setup made the results flaky, as jobs seemed to alternate between two kinds of runners with different specs. He also admitted that their benchmarks lacked a test with lots of packfiles, which is why they didn’t catch this regression.

Dscho replied to Peff that the new test directly reflects what GIT_PS1 runs, and that it exercises a subtly different code path, as --short has to look for a unique abbreviation, while --verify can stop as soon as it has found the object. Peff answered that the regression was about creating the initial pack list, so it happened whether each pack was opened or not. He noted, though, that the existing tests that look at each object only did so with 1, 50 and 1,000 packs, not with 10,000, and in the end he was OK with the redundancy since the new test isn’t expensive.

Dscho also told Junio that he had to take back his claim about the slower clones in CI, as the patch didn’t fix that issue, which was still being investigated.

Why so many packs?

D. Ben Knoble asked whether enabling maintenance on the user’s repository could be an intermediate solution.

Dscho replied that the issue was actually about a Scalar clone, and more specifically a Microsoft Git Scalar clone. He explained that “a substantial part of Microsoft Git failed to get upstreamed to core Git”, including the “shared cache repository” feature. With it, a bare repository is set up as an alternate of the actual clone, and scheduled fetches go into that shared cache (see the commit introducing it). Maintenance usually runs on the shared cache, but Dscho suspected that it often takes too long to finish before machines are shut down for the day. As a result, “it is still not exactly rare to find setups with five-digit packfile counts. And since we can handle this more gracefully, we should ;-)”.

Ben clarified that he had meant maintenance would likely help the local case, like the shell prompt, more than the clones.

Naming and design discussions

Patrick reviewed the patch. Besides pointing out a typo in the commit message, he suggested renaming the is_new parameter to accept_duplicates, since the function would then just append the entry without ensuring that the packfile is unique in the list. He also sketched an alternative: tracking added packs in a hashmap. This would also cover packfile_list_prepend() and wouldn’t require callers to know about the mechanism. With a doubly-linked list, moving existing entries to the back or the front, which happens often to re-sort the list during object lookups, would also become cheap. He sent a patch implementing this idea, while wondering whether the added complexity was worth it.

Dscho agreed to fix the typo and drop the claim about CI clones. He disagreed with the new name though, as the function is not accepting duplicates: the callers know the packfiles cannot be duplicates. Interestingly, he said that his first reaction had also been to write a hashmap-based fix, until “the AI assistant pointed out that no duplicates could possibly exist yet.” He agreed that the added complexity wasn’t needed, at least not yet.

Patrick replied that, seen outside the context of its current caller, the parameter just tells whether packs should be deduplicated. He considered pursuing his patch anyway, as he thought it would speed up reordering significantly with 38k packfiles, in which case it would supersede Dscho’s patch. Dscho proposed the skip_dup_check name instead, and pointed out that even a hashset lookup is slower than skipping the search altogether. Patrick agreed to move forward with Dscho’s patch.

Junio also replied to Patrick’s naming suggestion. He had “the same thought”, as the current callers might have been vetted thoroughly, but future callers or code paths might break the promise that only new packs are added. He also asked whether it was well understood what bad things duplicate entries in a pack list could lead to.

Peff replied to Patrick that such a hashmap already exists: since ec48540fe8, packfile_store_add_pack() and packfile_store_load_pack() use one, and that is precisely why the new parameter can be set to true for the remaining caller. Otherwise, “reprepare” operations would create duplicates.

Patrick suggested moving that map from the packfile store into the packfile list, to make it more generally useful. Peff answered that the map protects more than adding packs to the list, as it avoids calling add_packed_git(), which allocates memory and performs a number of stat() calls. So the existence check would have to happen much earlier than in packfile_list_append(). He added that it would be easier to see which generalized pattern would be useful if there were more than one caller of packfile_list_append().

Patrick pointed out that there were other callers of packfile_list_prepend(), which has the same problem. Peff agreed that prepend() calls appear in some hot code paths, including the MRU adjustment in find_pack_entry(), and that this could be a candidate for the clone slowdown Dscho was still investigating. But he wouldn’t want to pay the cost of hash-based deduplication there, as no new pack is added. Moving an entry should instead be an O(1) operation using a doubly-linked list.

Peff also explained that it is harder to build a synthetic test for prepending, because of pack locality. If two consecutive lookups move the same pack to the front, the second one finds it there almost immediately. He showed, though, how to spread a history across many packs using git fast-import with fastimport.unpackLimit=0 and a checkpoint after each commit. Timing git rev-list --count then showed quadratic growth taking over around 2,000 packs, from 18ms with 500 packs to 6.3 seconds with 16,000 packs. He noted that this didn’t prove much about list management, as looking up objects across packs is linear anyway, so this situation is inherently quadratic. Still, he found it “prudent for these MRU updates to use a constant-time movement within the list, rather than an explicit duplicate check and removal.”

Version 2

Meanwhile, Dscho sent a version 2 of the patch. It fixed the typo found by Patrick, dropped the claim that the patch fixed the CI clone regression, and renamed the is_new parameter to skip_dup_check.

Patrick said he was happy with this version, and that the other parts of the discussion could be iterated on after the patch landed. Junio agreed and marked it for ‘next’.

Conclusion

A small patch was enough to fix a quadratic slowdown that made shell prompts noticeably slower in repositories with tens of thousands of packfiles. The discussion around it showed that this was the regression of a problem already fixed in 2019, and that the perf test suite had detected it, but that nobody noticed. Contributors discussed how to better catch such regressions with continuous benchmarking, and why some real-world setups, like Microsoft Git’s Scalar shared cache, can accumulate so many packs. Ideas for further improvements, like constant-time MRU moves in the packfile list, were also put forward for later.

The patch was merged into the ‘master’ branch and is part of the Git 2.56.0 release.

Developer Spotlight: Harald Nordgren

  • Who are you and what do you do?

    I am Harald Nordgren, software developer for 20+ years, most recently worked as CTO of Diet Doctor and have been an enthusiastic Git user for many years. I am married to Linda and we have 3 young children.

  • What would you name your most important contribution to Git?

    Most useful is status.comparebranches, which I see every time I run git status and it blows my mind that I was able to put it there.

  • What are you doing on the Git project these days, and why?

    I’m working on history squash and doing GitHub CI improvements.

  • If you could get a team of expert developers to work full time on something in Git for a full year, what would it be?

    Not sure! Git is very good already.

  • If you could remove something from Git without worrying about backwards compatibility, what would it be?

    Many settings (like diff.algorithm and branch.sort) should have more user friendly defaults, but it requires us to break backwards compatibility. Users are not getting a smooth experience and many are very afraid to mess with their Git settings. I would love to have a setup wizard with this as one possible preset:

    branch.sort=-committerdate
    tag.sort=-version:refname
    diff.algorithm=histogram
    diff.colormoved=zebra
    diff.compactionheuristic=true
    
  • What is your favorite Git-related tool/library, outside of Git itself?

    GitHub’s official gh tool is awesome. And before it came and took over, I loved using mislav/hub.

  • Do you happen to have any memorable experience w.r.t. contributing to the Git project? If yes, could you share it with us?

    My first merged commit in 2018 (patch) was an unreal experience, I couldn’t believe I got to be part of this project.

  • What is your toolbox for interacting with the mailing list and for development of Git?

    GitGitGadget for submitting patches, Gmail for answering emails and VSCode as editor.

    But AI writes the code for me nowadays. I bring the idea and get a first draft (if it’s horrible I start over) and when I have something that feels sound, I “quick save” by committing/pushing and then feedback on the solution until it’s nice. I use one AI session per topic, and keep them open for the reviews so it maintains the context. It’s incredible to have a sparring partner that never gets tired!

  • What is your advice for people who want to start Git development? Where and how should they start?

    Have an idea for something that you yourself need, something you would want in Git that is missing. Don’t focus on getting the credit, focus on an actual need and the rest will follow.

    If you don’t understand Git well as a user, it will be hard to contribute meaningfully, so start by reading up on Git. Back when Stack Overflow was still popular I used to love to read about Git there. I still love to dig into the Git documentation and ask AI about some new options, there are many to discover!

  • If there’s one tip you would like to share with other Git developers, what would it be?

    Commit early, commit often.

Other News

Various

  • What’s new in Git 2.56.0? by Karthik Nayak on GitLab Blog. Mentions Git Merge 2026 and schedule for Git 3.0, git history drop <commit>, git branch --delete-merged and git branch --forked, the new create, delete, update and rename subcommands of git ref, git replay now working with commit ranges containing merge via --linearize, results of four Google Summer of Code 2026 projects, and more.
  • Highlights from Git 2.56 by Elijah Newren on GitHub Blog. Mentions git add --resolved for marking conflicts as resolved without accidentally staging too much, faster finding of common ancestor(s), improving repacking, and more.
  • Disclosure of Vulnerability in the Radicle’s Network Protocol.

Light reading

Easy watching

Git tools and sites

  • gat: simple, fast, versioned large-file storage for Git. gat is what git-lfs would be if it didn’t need a special server, and what dvc would be if it did one thing. You commit small metadata files with Git, and upload large-file content with Gat. Written in Rust, under Apache 2.0 license.
  • BlackGit allows for locking files on a Git server. The CLI client lets you download only files you care about through sparse-checkout, the server serves a file level control layer as a proxy to upstream. Server written in Java, uses JGit and Netty 4.1; Client written in Python 3, requires Git 2.54+ to be installed. Under MIT license.
  • git-hooks-ext — semantic events for Git reference transactions. Git’s reference-transaction hook reports raw old and new values together with ref names. It does not tell a hook that a branch was created, a tag was deleted or a ref was renamed. git-hooks-ext turns those low-level updates into semantic events such as branch-created, tag-deleted, remote-head-updated, and ref-created. It also adds the worktree lifecycle events that Git does not provide. Provided as a Git hook and helper CLI tool. Written in C and shell, under GPL-2.0 license.
  • Git Reattribute is a small, cross-platform CLI for replacing one Git identity with another across a repository’s history, safely - without requiring you to hand-write a git filter-repo invocation. Also includes guard, a prevention companion that blocks a denied identity in CI or a local hook before it ever lands. This tool rewrites Git history. Written in Python (and shell), under MIT license.
  • Aldine is a slim, self-hosted, open-source LaTeX collaboration platform, an Overleaf alternative built for speed and simplicity. Real-time multi-cursor editing (CRDT-based, using Yjs), every project being a real Git repository with branches, native Zotero support. Provided as two containers (set up using Docker Compose) and flat files. There is a live demo that resets nightly. Written in TypeScript, under AGPL-3.0 license. Pre 1.0.0 version.
  • Ju! Ju! Tsu! is a book about Jujutsu, written by Arialdo Martini, under CC BY-SA 4.0 license. It is available in HTML and PDF.
  • Vendetect is a command-line tool for automatically detecting vendored and copy/pasted code between repositories. It uses similarity detection algorithms to compare code files and highlight matching sections. Written in Python, under AGPL-3.0 license.
  • crux is a standalone tool that finds the exact commit behind a behavior change, like git bisect run. After the search is performed, the found commit is minimized. Crux uses partial hunks on the parent and executes the command until there are only those lines left in the diff that cause behavioral changes, resulting in the causal diff. It handles cases git bisect can’t: behaviors that aren’t pass/fail tests, failures that need two commits together, and regressions caused by dependency updates. On crates.io as crux-finder. Written in Rust, under MIT license.
  • Foremerge is the open-source coordination protocol for coding agents, built on top of Git; Agents keep isolated worktrees while sharing intent, semantic claims, dependencies, provisional ChangeSets, decisions, validation, and provenance. The idea is to catch intent conflicts before code conflicts. Written in Rust, under Apache 2.0 license.
    Looks like GitLens, but for agents.
  • Riftri - Lightweight Git workspaces for parallel development. Riftri creates real Git worktrees without eagerly storing another full physical copy of every unchanged project file. It is designed for developers and coding agents working on several tasks at once. With it, you can keep using normal files, normal Git, and the tools you already have. Uses ReFS block clone on MS Windows, Btrfs or reflink or OverlayFS on Linux, and APFS clone on macOS - that is native copy-on-write backends. Written in Rust, under MIT license. Note: Riftri is experimental, pre-release software.
  • Dear Machine is the local client that lets you email back and forth with your computer. Machtiani is the experimental harness underneath it, using iteration to manage long-running agentic AI sessions without compaction. The idea is to use email to work with multiple agents at the same time, with each email thread a session, as described in The Shell & Email post. Written in Python and Go. Under MIT license.
  • nimblegate is a server that sits between your AI agent and your real Git host. It provides Git push guardrails for AI agents: block unsafe pushes consistently, forward safe ones, record every decision. Written in Go, under PolyForm Noncommercial License 1.0.0.
  • The Game of Trees Hub (GotHub), a transparently funded Git repository hosting service, with infrastructure on OpenBSD and the Game of Trees (GoT) VCS, which was first mentioned in Git Rev News Edition #131, got support for Mailing Lists.
  • Pushin.eu (in invite-only beta) and CodeFloe (public Git service running on Forgejo with a free tier) are Git hosting sites advocating that they are hosted in Europe, under EU law.
    The previous edition mentioned another forge hosted in Europe: Gitoro.
  • WebTerm Learn: free Git courses that run in a simulated terminal in the browser. The exercises check the resulting repository state (index, branches, commits) rather than the exact command typed, so equivalent commands are accepted. Covers first commit through rebase, stash, cherry-pick, conflicts and pull requests against a simulated remote. The first lesson of each course needs no account; the rest need a free account. Its sibling WebTerm is a no-signup browser terminal with short tutorials, including a set on Git troubleshooting.

Releases

Credits

This edition of Git Rev News was curated by Christian Couder <christian.couder@gmail.com>, Jakub Narębski <jnareb@gmail.com>, Markus Jansen <mja@jansen-preisler.de> and Kaartic Sivaraam <kaartic.sivaraam@gmail.com> with help from Harald Nordgren, Maciej Ciemborowicz, Toon Claes, @Sal-ami, @DaiAoki and Štěpán Němec.