Skip to content

About

Unpack untrusted ZIP and tar archives with path containment, a real symlink policy, and streaming quotas.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

safeextract

Unpack a ZIP or tar archive you did not create, with path containment, a symlink policy that never walks through a link, and hard resource quotas enforced while the archive streams — one API for both container formats.

one archive: 199 KB on disk, 200 MB in one entry (1028.2x)

zipfile.extractall       57ms   wrote    200 MB   no ratio/size limit exists to refuse this
safeextract.extract       8ms   wrote      0 MB   refused — entry 'data/payload.bin' compression ratio 100.3x exceeded max_ratio 100.0 (203842 stored bytes inflated to 20447232)

That's python examples/bomb_demo.py, on this machine, today. The "naive" column is not a straw man — it is exactly what calling Python's own zipfile.extractall() does with this file, with nothing added or removed.

Honest positioning — read this before installing anything

Python's zipfile has refused zip-slip (../, absolute paths, drive letters) since 3.6. tarfile gained extraction filters in 3.12, and the 'data' filter — which blocks links to absolute paths or outside the destination, and device files — is the default as of 3.14, with no deprecation warning. And there is real, maintained prior art beyond the stdlib: safezip (ZIP; streaming quotas with a guard/streamer dual phase, symlink policies REJECT/IGNORE/RESOLVE_INTERNAL, commits as recent as June 2026) and its sibling safetar (tar; same author, same quota shape), plus zipguard (ZIP; policy enforcement against ZipSlip/bombs/RTLO/extension drops, CLI-first) and DefuseZip (ZIP; scanning and detection with a ratio threshold and nested-zip limits). None of this is vaporware, and this README does not pretend otherwise.

Here is what was actually checked, on Python 3.14.7, before writing a line of this package:

Check Result
zipfile.extractall, entry named ../evil.txt contained — zipfile's own extraction code strips ../. components and drive/root prefixes before joining, a fix long-shipped by 3.14 (commonly cited to 3.6)
zipfile.extractall, entry named /etc/evil.txt or C:\Windows\... contained — same code path
zipfile.extractall, entry with the Unix-symlink bit set in external_attr not honored at all — written as a plain file containing the link-target text, no real symlink ever created
zipfile.extractall, 199 KB deflating to 200 MB no limit of any kind — see the demo above; this is CVE-2019-9674, closed wontfix, "add documentation instead"
tarfile.extractall(filter="data"), symlink pointing outside destination refused — LinkOutsideDestinationError
tarfile.extractall(filter="data"), symlink contained inside destination, next entry writes through it allowed — the filter resolves-and-checks the link target, it does not refuse walking through a link that resolves inside the tree
tarfile.extractall(filter="data"), 199 KB gzip inflating to 200 MB no limit of any kind — the docs explicitly say to use "external (e.g. OS-level) limits"
safezip / safetar, ratio limit on a 20 KB all-zeros entry (ratio ≈ 200:1, ordinary and harmless) flagged — no grace floor; an absolute ratio cutoff with no exemption for small entries
safezip and safetar together two separate packages, two separate APIs, two separate option objects — "no mention of safezip or shared API design patterns" in safetar's own README

So: the path-traversal gap is closed in the stdlib and has been for a decade — nothing here does anything zipfile hasn't already done since 3.6. The symlink gap is closed for tar by the 'data' filter as of 3.12/3.14, with one caveat (below). The quota gap is real and confirmed: neither zipfile nor tarfile enforce any decompressed-size, entry-count, or ratio limit, ever, and the one CVE filed about it was closed as wontfix. safezip/safetar already close the quota gap well — per format, if you're fine importing two packages with two APIs for zip and tar, and if an absolute ratio cutoff with no grace floor doesn't bite you on ordinary small compressible files (logs, JSON, sparse binaries).

What's actually left, after all of that:

  1. No grace floor on the ratio check, anywhere. A 20 KB file of zeros compresses about 200:1 and is completely unremarkable; safezip and safetar's default max_ratio=200 would flag it immediately with no floor to hold the check back until there's enough signal to mean something.
  2. No single package covers zip and tar under one API. safezip and safetar are two separate installs with two separate option shapes, from the same author, explicitly not unified.
  3. tarfile's 'data' filter permits walking through a contained symlink. It resolves the link target and checks where it points, which is a real and useful check — but it is a check performed once, not a guarantee that nothing walks through the link component-by-component at the filesystem level. safeextract's stricter rule (below) never walks through any symlink, contained or not, which also stops the version of the two-step escape where the link was planted by a previous extraction into the same directory rather than by the archive currently being read — something a per-extraction resolve-and-check can't see.

That is a real gap, but a narrower one than "nobody has solved this." If you only extract ZIP files and are fine with a dependency, use safezip. If you only extract tar files, use safetar. If your archives are from a source you already trust, or trust enough that path containment is the only property you need, use the standard library directly — it is free, it is already in your dependency tree, and as of 3.14 tarfile's default filter is good. Reach for this package specifically when you want one call that handles both formats, with a ratio check that has a grace floor, and a symlink rule that is stricter than "does the target resolve inside the tree right now."

The problem, concretely

Every field in an archive is written by whoever made the archive. The entry name, the file mode, the declared uncompressed size, the flag that says "this is a symlink" — all of it is input. Three shapes of this are still worth defending against explicitly even with the stdlib improvements above:

  • Zip slip, cross-platform. ../../etc/cron.d/x, /etc/cron.d/x, C:\Windows\System32\x, ..\..\x — the last one is a perfectly legal single filename on Linux, and a traversal the moment the same archive is extracted on Windows. Handled by both zipfile (since 3.6) and tarfile's 'data' filter (default since 3.14) — safeextract re-checks it anyway, independent of platform, as defense in depth.
  • Symlink escape, including the planted variant. One entry is a symlink; a later entry — in this archive, or a future one extracted into the same directory — writes through it. safeextract never walks through a symlink to create anything, under any policy, whether the link was just created by this extraction or was already sitting there.
  • Resource bombs. 199 KB inflating to 200 MB, or 10,000 entries of it. Declared sizes are exactly as attacker-controlled as the entry name; reading them and trusting them defends nothing, which is the literal text of CVE-2019-9674's resolution.

Why the obvious defenses don't work

name.includes('..') (or in, in Python) misses absolute paths, drive letters, and UNC paths, and flags legitimate names like my..notes.txt.

resolved.startswith(dest) is the check almost everyone writes, and it's wrong: /srv/dest-evil/x starts with the string /srv/dest. The separator has to be part of the comparison — safeextract checks against dest + os.sep.

Checking the resolved path and then writing still loses to symlinks. dest/link/payload is inside dest by every string test there is; whether it's really inside depends on what link turned out to be, and os.makedirs walks straight through an existing symlink without a word. Directories here are created one component at a time, lstat-ing anything already present, so nothing is ever walked through.

Reading the entry's declared size first is the ZIP/tar version of trusting Content-Length. It's a field the archive's author chose. A bomb declares 1 KB and inflates until the disk is full — unless something checks the bytes actually arriving, not the number claimed up front.

Install

pip install safeextract

Python ≥ 3.11. Zero runtime dependencies — both readers are the standard library's own zipfile and tarfile, used through one containment and quota layer.

Use

from safeextract import extract

extract("upload.zip", dest, max_entries=10_000, max_total_bytes=500 << 20,
        max_entry_bytes=100 << 20, max_ratio=100, symlinks="reject")

Same call for a .tar, .tar.gz, .tar.bz2, or .tar.xz — the format is detected from the file's contents, not its name.

from safeextract import extract, CompressionBombError, PathEscapeError

try:
    extract("plugin.tar.gz", dest)
except PathEscapeError as exc:
    print(f"entry {exc.entry_name!r} tried to escape: {exc.reason}")
except CompressionBombError as exc:
    print(f"{exc.compressed_bytes} bytes inflated to {exc.decompressed_bytes}")

Quotas are enforced while each entry streams out of the decompressor, so a bomb stops the moment it crosses a limit rather than after it has filled the disk. Defaults are conservative; nothing is unbounded.

Options

Option Default What it bounds
max_entries 10,000 Entries in the archive. Free to check up front for ZIP (central directory); for tar, checked incrementally as members are discovered, since tar has no index to read ahead of the stream.
max_total_bytes 1 GiB Decompressed bytes across all entries.
max_entry_bytes max_total_bytes Decompressed bytes in any one entry.
max_ratio 100 Decompressed ÷ compressed. For ZIP, per entry (its own declared compressed size). For tar, archive-wide (raw bytes consumed from the underlying file so far) — tar compresses the whole stream, not one member at a time, so there is no such thing as one member's compressed size. float("inf") disables it.
ratio_grace_bytes 1 MiB Bytes decoded before the ratio check applies at all.
max_path_length 1024 Characters in an entry name.
max_segment_length 255 Characters in any one path segment.
symlinks "reject" "reject", "skip", or "allow-contained". Governs both symlink and hardlink entries.
preserve_mode False Apply the archive's Unix permission bits, masked to 0o777.
overwrite False Whether an existing regular file may be replaced. A symlink or directory already at that path is never silently replaced.
cleanup_on_error True Remove everything this call wrote if extraction fails partway through.
filter — (entry) -> bool, run before any other check; entries it rejects are recorded in result.skipped.

extract() returns a Result with files, directories, symlinks, skipped (absolute paths, except skipped which lists archive entry names), and total_bytes.

How the limits are actually enforced

Declared sizes are checked first, because refusing an entry before opening it is free. They are a filter, never a guarantee: the same limits run again against bytes as they come out of the decompressor, and an entry that inflates past its own declared size is cut off on the spot, mid-chunk, before that chunk is written to disk.

Why the ratio has a grace floor

20 KB of zeros compresses about 200:1 and is completely ordinary. Small, repetitive files — JSON, logs, sparse binaries — do this constantly, and a ratio limit applied to them without a floor produces nothing but false positives (confirmed against safezip/safetar's defaults above, which have no such floor). ratio_grace_bytes holds the check back until enough has decoded for the ratio to mean something; below that floor, only max_entry_bytes applies.

Symlinks

"reject" (default) fails the extraction on the first symlink or hardlink entry. "skip" leaves it out and continues, recording it in result.skipped — since the link is never created, nothing can later be written "through" it; a later entry that names a path under the skipped link's name just gets an ordinary directory. "allow-contained" creates a symlink only if its target — resolved the way the kernel resolves it, relative to the directory holding the link — stays inside the destination; absolute targets are refused even when they'd currently resolve inside, because they stop being contained the moment the tree is moved. Hardlinks are resolved against already-extracted archive members and refused if the target hasn't been written yet or isn't a plain file.

Under every policy, no symlink is ever walked through — not one this extraction just created, and not one already sitting in the destination from a previous run or planted directly. Directories are created one path component at a time; anything already present is lstat-ed (never stat, which would follow a link) and has to be a real directory before the walk continues into it. Files are opened with O_EXCL. That's what stops the two-step escape in both directions — the symlink planted by an earlier entry in the same archive, and the symlink already there before this call started.

Errors

All extend SafeExtractError.

  • PathEscapeError — entry_name, reason
  • PathTooLongError — entry_name, limit
  • SymlinkError — entry_name, target, policy
  • EntryTooLargeError — entry_name, limit, bytes_written
  • ArchiveTooLargeError — limit, bytes_written, entry_name
  • TooManyEntriesError — limit, entries
  • CompressionBombError — entry_name, ratio, limit, compressed_bytes, decompressed_bytes
  • EntryConflictError — entry_name, path
  • UnsupportedArchiveError — feature, entry_name
  • MalformedArchiveError — reason

What it does not do

It does not out-parse the stdlib. Both readers are zipfile and tarfile — no hand-rolled ZIP or tar parser, no zip64 restriction the way a hand-rolled reader would need one, no bespoke handling of local-header vs. central-directory disagreement (that's zipfile's problem to get right and it already does, since the central directory is what infolist() reads and streaming decompression checks actual bytes, not either header, regardless).

Compression methods are whatever the installed stdlib supports — stored, deflate, and (via zipfile's own bz2/lzma support) bzip2 and LZMA for ZIP; gzip, bzip2, and XZ for tar. Zstd is not supported by zipfile/tarfile in the versions this targets and is refused by name. Encrypted ZIP entries are refused; there is no decryption path.

It is not a virus scanner. It bounds where bytes land, how many there are, and how fast they grow. It has no opinion about what they contain.

It is not tested on Windows. Path rules refuse Windows-shaped escapes (drive letters, UNC paths, backslash separators) unconditionally, on any host OS, because archives cross platforms — but CI runs on Linux, and the symlink policies assume POSIX semantics (os.symlink, os.link, lstat).

Entries are extracted one at a time, synchronously. No parallelism, no resume, no async API.

max_entries is not equally cheap for both formats. ZIP has a central directory, so the count is known — and checked — before a single byte is touched. tar has no such index; reaching the Nth header in a compressed tar stream requires decompressing everything before it, so the count is enforced incrementally as members are discovered, not as one free up-front check. This is a structural property of the tar format, not something a smarter implementation here would fix.

Develop

python -m venv .venv && .venv/bin/pip install -e '.[dev]'
.venv/bin/python -m pytest -q
.venv/bin/python -m mypy src --strict
python examples/bomb_demo.py

Every malicious archive in the test suite (tests/archives.py) is built byte-by-byte in-process — traversal entries, absolute paths, planted symlinks, the two-step escape (both within one archive and pre-planted by a "previous" extraction), zip and tar.gz bombs, entry-count explosions. There are no fixtures to commit and nothing taken on faith.

License

MIT

About

Unpack untrusted ZIP and tar archives with path containment, a real symlink policy, and streaming quotas.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages