Skip to content

Clamp threads on extraction, and stop four ways to lose data quietly - #2

Merged
ValisSowilo merged 1 commit into
mainfrom
fix/streaming-and-self-overwrite
Sep 23, 2026
Merged

ValisSowilo merged 1 commit into
mainfrom
fix/streaming-and-self-overwrite

Conversation

@ValisSowilo

Copy link
Copy Markdown
Owner

Compression and extraction already stream one segment at a time. This PR adds the guard rails that were missing around that.

Extraction has no thread clamp. threads_for() only ran on the compress side. x -t0 on a -9 archive asked for a full model per core. Extraction now uses the same clamp, with the model sized for the worst-case (strided) geometry. On a 16 GB machine, -t64 at -9 now caps at -t13.

r bad.gl bad.gl destroyed the archive. The output was opened "wb" before any input was read. This case is now refused, and the check compares file identity, not spelling.

Non-regular files were archived as empty. A FIFO, a device, /dev/stdin or a /proc entry was stored as a 0-byte member with a valid SHA-256 and exit 0. These are now skipped with a message, and the run exits 1.

The output was archived into itself. A second run of c backup.gl . picked up the previous backup.gl and stored a partial copy of the new archive. The output file is now dropped from the inputs.

Segment rawlen is now bounded by the header's segmax. It used to be allowed up to 1 TB. A crafted index now gets exit 2 "index is corrupt" instead of a terabyte malloc followed by exit 1 "out of memory".

sfuzz.py gains rawlen_over_segmax. rawlen_over_wn_packed now declares a 512 MB segmax so that it still reaches the decoder check it tests.

Verified (Windows, MinGW gcc 16.1, static zlib 1.3.1)

  • Archives are byte-identical to the previous build at -1 -5 -9 -f1. Old and new builds read each other's archives.
  • sfuzz (93 runs), gfuzz (40 trials) and tfuzz (60 round trips) pass. The symlink cases were skipped because this host can't create links.
  • The old build reproduced each bug: repair onto itself wiped all 3 segments, a second c bk.gl . stored a 32 KB copy of itself, and c d.gl NUL exited 0 with an empty member.

Not run: Linux, macOS and the UBSan build. CI will cover those.

Found by a read-through of the streaming paths.  Compression and extraction
already stream a segment at a time; what was missing was the guard rails
around that.

Extraction had no thread clamp.  threads_for() sized -t to fit in RAM on the
compress side only; run_members took NTHREAD as given.  Decoding holds more
per worker than encoding, so `x -t0` on a -9 archive asked for a full model
per core and died in the allocator with no hint the thread count was why.
It now applies the same clamp, sizing the model for a strided geometry (the
largest, 30 contexts at -9) since the real one is not known until segments
are parsed.  On a 16 GB box -t64 at -9 now caps at -t13.

`r bad.gl bad.gl` destroyed the archive it was asked to repair.  do_repair
opens its output "wb" before reading a segment, so the input was truncated and
every segment then reported unrecoverable.  Refused now, by file identity
rather than spelling, so ./bad.gl and hard links count.

Non-regular files were archived as empty.  fsize64 seeks to the end, which on
a FIFO, device or /proc entry gives 0, so `c a.gl /dev/stdin` or a process
substitution stored a zero-byte member with a valid SHA-256 and exit 0.  The
walk now skips anything that is not a regular file, says so, and exits 1.

The output archive was archived into itself.  `c backup.gl .` run a second
time finds last run's backup.gl in the walk, truncates it, and stores whatever
partial copy of the new archive exists when it gets there.  Inputs that are
the output file are dropped with a note.

Segment rawlen was bounded at 1 TB instead of by the header.  The writer never
emits a segment over its -s (floored at MINCHUNK), so anything above that is a
corrupt index.  Before, a crafted one got a terabyte malloc per thread and
exited 1 "out of memory"; now it is "index is corrupt", exit 2.

sfuzz.py gains rawlen_over_segmax.  rawlen_over_wn_packed claims 256 MB, which
the new bound would now stop at the index, so it declares a 512 MB segmax to
keep reaching the decoder check it was written for.

Archives byte-identical to the previous build at -1 -5 -9 -f1, and old and new
binaries read each other's output.  sfuzz (93 runs), gfuzz (40) and tfuzz (60)
pass on Windows; the symlink cases were skipped on this host.
@ValisSowilo
ValisSowilo merged commit 739600b into main Sep 23, 2026
5 checks passed
@ValisSowilo
ValisSowilo deleted the fix/streaming-and-self-overwrite branch September 23, 2026 12:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant