zgba 站群
uv: Deduplicate all files in the wheel cache

uv: Deduplicate all files in the wheel cache

There was an error while loading. Please reload this page.

There was an error while loading. Please reload this page.

On main, we support content-addressed caching, but only at the wheel-level. That is, if you download the same wheel twice from different sources, they share a cache entry. But files within or across wheels are not deduplicated at all.

This PR adds deduplication at the file level: every file is now stored under its BLAKE3 hash in a files-v0 bucket. We hardlink these objects into their original locations in archive-v0, so the installation step doesn’t change at all — we’re just deduping within the cache (and cache cleanup removes file objects when their hardlink count drops to one).

In the prior proposal (#19694), we included the following table:

So we’re saving 545.2 MiB on my local machine, or about 10% of the cache.

In return, the net effect seems to be something like a <4% slowdown for cold installs (and no effect on warm installs), which I think is probably worthwhile here.

Sorry, something went wrong.

There was an error while loading. Please reload this page.

There was an error while loading. Please reload this page.

N.B. Slop comment for benchmark data.

We benchmarked the optimization in 85c3f7485 against the binary-only cache at a3f977ece, with —preview-features content-addressed-cache enabled on both. These are medians for the complete uv pip install process; positive changes mean slower.

Before this optimization, the measured cold regressions were +4.02% for AnyIO 4.9.0, +19.41% for SymPy 1.14.0, +12.75% for NumPy 2.2.6, +15.30% for PyTorch 2.7.1+cpu.

All eight median regressions are below 5%. The PyTorch cold 95% interval still reaches 5.18%, so its upper bound is not below 5%. Every other upper bound is below 5%.

We retained all 1,860 measured installs, including outliers, and combined every confirmation run for this candidate. Cold measurements have 360 paired rounds for AnyIO, 120 each for SymPy and NumPy, and 90 for PyTorch; warm measurements have 60 paired rounds per wheel. Baseline/candidate order alternates, with three warmups per case. Intervals use a paired percentile bootstrap of the ratio of medians with 10,000 resamples.

These measurements ran on Linux/ext4, an AMD EPYC-Milan VM pinned to eight CPUs, with Python 3.12.13. Both binaries use Rust 1.98.0 and the same profiling profile (optimized, no LTO). Each sample installs one pinned local wheel offline, without dependencies or bytecode compilation, using hardlinks into a fresh virtual environment. Cold removes the entire uv cache; warm retains a primed cache. Setup and cleanup are untimed, and the wheel OS page cache is warm. No compilation ran during the benchmarks. These results do not cover other platforms or network-inclusive installs.

All-file deduplication, content/executable identities, the cache layout, complete archives, and copy fallbacks are preserved. Inode checks confirmed that every archived file shares its file-store object for all four wheels. The five targeted integration tests passed ten stress iterations (50 executions), including local and streamed wheels with one and four workers, RECORD handling, cache cleanup, and cross-filesystem installation. Formatting and Clippy with warnings denied also passed.

Sorry, something went wrong.

There was an error while loading. Please reload this page.

There was an error while loading. Please reload this page.

✅ 25 untouched benchmarks ⏩ 12 skipped benchmarks1

Comparing charlie/dirhash-all-files (d12d2ad) with main (7c1d80e)

12 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Sorry, something went wrong.

There was an error while loading. Please reload this page.

This PR changes the tests when compared with the main base revision.

Sorry, something went wrong.

There was an error while loading. Please reload this page.

Okay, I experimented with a bunch of alternatives to try and make this more performant:

I also considered something like “guess whether a file is executable” during extraction (and then, if we’re wrong, copy it and change the mode after unzipping). I guess this could end up being more performant, but I wasn’t very happy with the heuristics.

Ultimately, I think what we have here is good.

Sorry, something went wrong.

There was an error while loading. Please reload this page.

If you bundle in #21340, I believe this is also faster than before (i.e., gains from #21340 outweigh the extra cost in this PR).

Sorry, something went wrong.

There was an error while loading. Please reload this page.

There was an error while loading. Please reload this page.

Successfully merging this pull request may close these issues.

There was an error while loading. Please reload this page.

View original article