Unpack a ZIP or tar archive you did not create, with path containment, a symlink policy that never walks through a link, and hard resource quotas enforced while the archive streams — one API for both container formats.
one archive: 199 KB on disk, 200 MB in one entry (1028.2x)
zipfile.extractall 57ms wrote 200 MB no ratio/size limit exists to refuse this
safeextract.extract 8ms wrote 0 MB refused — entry 'data/payload.bin' compression ratio 100.3x exceeded max_ratio 100.0 (203842 stored bytes inflated to 20447232)
That's python examples/bomb_demo.py, on this machine, today. The "naive"
column is not a straw man — it is exactly what calling Python's own
zipfile.extractall() does with this file, with nothing added or removed.
Python's zipfile has refused zip-slip (../, absolute paths, drive
letters) since 3.6. tarfile gained extraction filters in 3.12, and the
'data' filter — which blocks links to absolute paths or outside the
destination, and device files — is the default as of 3.14, with no
deprecation warning. And there is real, maintained prior art beyond the
stdlib: safezip (ZIP;
streaming quotas with a guard/streamer dual phase, symlink policies
REJECT/IGNORE/RESOLVE_INTERNAL, commits as recent as June 2026) and
its sibling safetar (tar;
same author, same quota shape), plus
zipguard (ZIP; policy
enforcement against ZipSlip/bombs/RTLO/extension drops, CLI-first) and
DefuseZip (ZIP; scanning and
detection with a ratio threshold and nested-zip limits). None of this is
vaporware, and this README does not pretend otherwise.
Here is what was actually checked, on Python 3.14.7, before writing a line of this package:
| Check | Result |
|---|---|
zipfile.extractall, entry named ../evil.txt |
contained — zipfile's own extraction code strips ../. components and drive/root prefixes before joining, a fix long-shipped by 3.14 (commonly cited to 3.6) |
zipfile.extractall, entry named /etc/evil.txt or C:\Windows\... |
contained — same code path |
zipfile.extractall, entry with the Unix-symlink bit set in external_attr |
not honored at all — written as a plain file containing the link-target text, no real symlink ever created |
zipfile.extractall, 199 KB deflating to 200 MB |
no limit of any kind — see the demo above; this is CVE-2019-9674, closed wontfix, "add documentation instead" |
tarfile.extractall(filter="data"), symlink pointing outside destination |
refused — LinkOutsideDestinationError |
tarfile.extractall(filter="data"), symlink contained inside destination, next entry writes through it |
allowed — the filter resolves-and-checks the link target, it does not refuse walking through a link that resolves inside the tree |
tarfile.extractall(filter="data"), 199 KB gzip inflating to 200 MB |
no limit of any kind — the docs explicitly say to use "external (e.g. OS-level) limits" |
safezip / safetar, ratio limit on a 20 KB all-zeros entry (ratio ≈ 200:1, ordinary and harmless) |
flagged — no grace floor; an absolute ratio cutoff with no exemption for small entries |
safezip and safetar together |
two separate packages, two separate APIs, two separate option objects — "no mention of safezip or shared API design patterns" in safetar's own README |
So: the path-traversal gap is closed in the stdlib and has been for a
decade — nothing here does anything zipfile hasn't already done since 3.6.
The symlink gap is closed for tar by the 'data' filter as of 3.12/3.14,
with one caveat (below). The quota gap is real and confirmed: neither
zipfile nor tarfile enforce any decompressed-size, entry-count, or
ratio limit, ever, and the one CVE filed about it was closed as
wontfix. safezip/safetar already close the quota gap well —
per format, if you're fine importing two packages with two APIs for zip
and tar, and if an absolute ratio cutoff with no grace floor doesn't bite
you on ordinary small compressible files (logs, JSON, sparse binaries).
What's actually left, after all of that:
- No grace floor on the ratio check, anywhere. A 20 KB file of zeros
compresses about 200:1 and is completely unremarkable;
safezipandsafetar's defaultmax_ratio=200would flag it immediately with no floor to hold the check back until there's enough signal to mean something. - No single package covers zip and tar under one API.
safezipandsafetarare two separate installs with two separate option shapes, from the same author, explicitly not unified. - tarfile's
'data'filter permits walking through a contained symlink. It resolves the link target and checks where it points, which is a real and useful check — but it is a check performed once, not a guarantee that nothing walks through the link component-by-component at the filesystem level. safeextract's stricter rule (below) never walks through any symlink, contained or not, which also stops the version of the two-step escape where the link was planted by a previous extraction into the same directory rather than by the archive currently being read — something a per-extraction resolve-and-check can't see.
That is a real gap, but a narrower one than "nobody has solved this." If
you only extract ZIP files and are fine with a dependency, use safezip.
If you only extract tar files, use safetar. If your archives are from
a source you already trust, or trust enough that path containment is the
only property you need, use the standard library directly — it is
free, it is already in your dependency tree, and as of 3.14 tarfile's
default filter is good. Reach for this package specifically when you want
one call that handles both formats, with a ratio check that has a grace
floor, and a symlink rule that is stricter than "does the target resolve
inside the tree right now."
Every field in an archive is written by whoever made the archive. The entry name, the file mode, the declared uncompressed size, the flag that says "this is a symlink" — all of it is input. Three shapes of this are still worth defending against explicitly even with the stdlib improvements above:
- Zip slip, cross-platform.
../../etc/cron.d/x,/etc/cron.d/x,C:\Windows\System32\x,..\..\x— the last one is a perfectly legal single filename on Linux, and a traversal the moment the same archive is extracted on Windows. Handled by bothzipfile(since 3.6) andtarfile's'data'filter (default since 3.14) — safeextract re-checks it anyway, independent of platform, as defense in depth. - Symlink escape, including the planted variant. One entry is a symlink; a later entry — in this archive, or a future one extracted into the same directory — writes through it. safeextract never walks through a symlink to create anything, under any policy, whether the link was just created by this extraction or was already sitting there.
- Resource bombs. 199 KB inflating to 200 MB, or 10,000 entries of it. Declared sizes are exactly as attacker-controlled as the entry name; reading them and trusting them defends nothing, which is the literal text of CVE-2019-9674's resolution.
name.includes('..') (or in, in Python) misses absolute paths, drive
letters, and UNC paths, and flags legitimate names like my..notes.txt.
resolved.startswith(dest) is the check almost everyone writes, and
it's wrong: /srv/dest-evil/x starts with the string /srv/dest. The
separator has to be part of the comparison — safeextract checks against
dest + os.sep.
Checking the resolved path and then writing still loses to symlinks.
dest/link/payload is inside dest by every string test there is; whether
it's really inside depends on what link turned out to be, and
os.makedirs walks straight through an existing symlink without a word.
Directories here are created one component at a time, lstat-ing anything
already present, so nothing is ever walked through.
Reading the entry's declared size first is the ZIP/tar version of
trusting Content-Length. It's a field the archive's author chose. A bomb
declares 1 KB and inflates until the disk is full — unless something checks
the bytes actually arriving, not the number claimed up front.
pip install safeextractPython ≥ 3.11. Zero runtime dependencies — both readers are the standard
library's own zipfile and tarfile, used through one containment and
quota layer.
from safeextract import extract
extract("upload.zip", dest, max_entries=10_000, max_total_bytes=500 << 20,
max_entry_bytes=100 << 20, max_ratio=100, symlinks="reject")Same call for a .tar, .tar.gz, .tar.bz2, or .tar.xz — the format is
detected from the file's contents, not its name.
from safeextract import extract, CompressionBombError, PathEscapeError
try:
extract("plugin.tar.gz", dest)
except PathEscapeError as exc:
print(f"entry {exc.entry_name!r} tried to escape: {exc.reason}")
except CompressionBombError as exc:
print(f"{exc.compressed_bytes} bytes inflated to {exc.decompressed_bytes}")Quotas are enforced while each entry streams out of the decompressor, so a bomb stops the moment it crosses a limit rather than after it has filled the disk. Defaults are conservative; nothing is unbounded.
| Option | Default | What it bounds |
|---|---|---|
max_entries |
10,000 | Entries in the archive. Free to check up front for ZIP (central directory); for tar, checked incrementally as members are discovered, since tar has no index to read ahead of the stream. |
max_total_bytes |
1 GiB | Decompressed bytes across all entries. |
max_entry_bytes |
max_total_bytes |
Decompressed bytes in any one entry. |
max_ratio |
100 |
Decompressed ÷ compressed. For ZIP, per entry (its own declared compressed size). For tar, archive-wide (raw bytes consumed from the underlying file so far) — tar compresses the whole stream, not one member at a time, so there is no such thing as one member's compressed size. float("inf") disables it. |
ratio_grace_bytes |
1 MiB | Bytes decoded before the ratio check applies at all. |
max_path_length |
1024 | Characters in an entry name. |
max_segment_length |
255 | Characters in any one path segment. |
symlinks |
"reject" |
"reject", "skip", or "allow-contained". Governs both symlink and hardlink entries. |
preserve_mode |
False |
Apply the archive's Unix permission bits, masked to 0o777. |
overwrite |
False |
Whether an existing regular file may be replaced. A symlink or directory already at that path is never silently replaced. |
cleanup_on_error |
True |
Remove everything this call wrote if extraction fails partway through. |
filter |
— | (entry) -> bool, run before any other check; entries it rejects are recorded in result.skipped. |
extract() returns a Result with files, directories, symlinks,
skipped (absolute paths, except skipped which lists archive entry
names), and total_bytes.
Declared sizes are checked first, because refusing an entry before opening it is free. They are a filter, never a guarantee: the same limits run again against bytes as they come out of the decompressor, and an entry that inflates past its own declared size is cut off on the spot, mid-chunk, before that chunk is written to disk.
20 KB of zeros compresses about 200:1 and is completely ordinary. Small,
repetitive files — JSON, logs, sparse binaries — do this constantly, and a
ratio limit applied to them without a floor produces nothing but false
positives (confirmed against safezip/safetar's defaults above, which
have no such floor). ratio_grace_bytes holds the check back until enough
has decoded for the ratio to mean something; below that floor, only
max_entry_bytes applies.
"reject" (default) fails the extraction on the first symlink or hardlink
entry. "skip" leaves it out and continues, recording it in
result.skipped — since the link is never created, nothing can later be
written "through" it; a later entry that names a path under the skipped
link's name just gets an ordinary directory. "allow-contained" creates a
symlink only if its target — resolved the way the kernel resolves it,
relative to the directory holding the link — stays inside the destination;
absolute targets are refused even when they'd currently resolve inside,
because they stop being contained the moment the tree is moved. Hardlinks
are resolved against already-extracted archive members and refused if the
target hasn't been written yet or isn't a plain file.
Under every policy, no symlink is ever walked through — not one this
extraction just created, and not one already sitting in the destination
from a previous run or planted directly. Directories are created one path
component at a time; anything already present is lstat-ed (never
stat, which would follow a link) and has to be a real directory before
the walk continues into it. Files are opened with O_EXCL. That's what
stops the two-step escape in both directions — the symlink planted by an
earlier entry in the same archive, and the symlink already there before
this call started.
All extend SafeExtractError.
PathEscapeError—entry_name,reasonPathTooLongError—entry_name,limitSymlinkError—entry_name,target,policyEntryTooLargeError—entry_name,limit,bytes_writtenArchiveTooLargeError—limit,bytes_written,entry_nameTooManyEntriesError—limit,entriesCompressionBombError—entry_name,ratio,limit,compressed_bytes,decompressed_bytesEntryConflictError—entry_name,pathUnsupportedArchiveError—feature,entry_nameMalformedArchiveError—reason
It does not out-parse the stdlib. Both readers are zipfile and
tarfile — no hand-rolled ZIP or tar parser, no zip64 restriction the way
a hand-rolled reader would need one, no bespoke handling of local-header
vs. central-directory disagreement (that's zipfile's problem to get right
and it already does, since the central directory is what infolist()
reads and streaming decompression checks actual bytes, not either header,
regardless).
Compression methods are whatever the installed stdlib supports —
stored, deflate, and (via zipfile's own bz2/lzma support) bzip2 and
LZMA for ZIP; gzip, bzip2, and XZ for tar. Zstd is not supported by
zipfile/tarfile in the versions this targets and is refused by name.
Encrypted ZIP entries are refused; there is no decryption path.
It is not a virus scanner. It bounds where bytes land, how many there are, and how fast they grow. It has no opinion about what they contain.
It is not tested on Windows. Path rules refuse Windows-shaped escapes
(drive letters, UNC paths, backslash separators) unconditionally, on any
host OS, because archives cross platforms — but CI runs on Linux, and the
symlink policies assume POSIX semantics (os.symlink, os.link, lstat).
Entries are extracted one at a time, synchronously. No parallelism, no resume, no async API.
max_entries is not equally cheap for both formats. ZIP has a central
directory, so the count is known — and checked — before a single byte is
touched. tar has no such index; reaching the Nth header in a compressed
tar stream requires decompressing everything before it, so the count is
enforced incrementally as members are discovered, not as one free
up-front check. This is a structural property of the tar format, not
something a smarter implementation here would fix.
python -m venv .venv && .venv/bin/pip install -e '.[dev]'
.venv/bin/python -m pytest -q
.venv/bin/python -m mypy src --strict
python examples/bomb_demo.pyEvery malicious archive in the test suite (tests/archives.py) is built
byte-by-byte in-process — traversal entries, absolute paths, planted
symlinks, the two-step escape (both within one archive and pre-planted by
a "previous" extraction), zip and tar.gz bombs, entry-count explosions.
There are no fixtures to commit and nothing taken on faith.
MIT