Skip to content

Bound gzip header FNAME and FCOMMENT string sizes to prevent unbounded memory growth (issue #91) - #93

Merged
sile merged 1 commit into
masterfrom
fix-issue-91-bound-gzip-header-strings
Sep 6, 2026
Merged

sile merged 1 commit into
masterfrom
fix-issue-91-bound-gzip-header-strings

Conversation

@sile

@sile sile commented Sep 6, 2026 •

Copy link
Copy Markdown
Owner

Summary

Bound the length of the gzip header FNAME and FCOMMENT strings so that parsing untrusted input cannot allocate without limit. The shared read_cstring helper now rejects a string once it would exceed 64 KiB, returning io::ErrorKind::InvalidData.

Closes #91

Problem

RFC 1952 does not set an upper bound on the length of the gzip header FNAME / FCOMMENT strings. Header::read_from reads these via read_cstring, which appends bytes to a Vec until a NUL terminator or EOF. An attacker-controlled stream could therefore force an arbitrarily large allocation while the header is parsed, before any payload output is produced. This affects synchronous parsing in gzip::Decoder::new, the gzip::MultiDecoder member transition, and the non-blocking gzip decoder. Output-side decompression bounds do not help, since the allocation happens before any output exists.

Solution

  • Added MAX_HEADER_STRING_LEN = 64 * 1024 and a guard in the shared read_cstring helper. Once buf.len() >= MAX_HEADER_STRING_LEN and the just-read byte is non-NUL, read_cstring returns InvalidData. A NUL terminator immediately after exactly 64 KiB is still accepted, so a string of exactly 64 KiB is valid; a 65th non-NUL byte is rejected.
  • Because both FNAME and FCOMMENT go through read_cstring, the bound applies to the synchronous gzip::Decoder, gzip::MultiDecoder, and non_blocking::gzip::Decoder (which reads the header via Header::read_from) without any per-path duplication.
  • The maximum allocation retained while parsing a header is bounded by MAX_HEADER_STRING_LEN (plus one for the rejection-probe byte).

Why 64 KiB?

This is a policy choice, not a format requirement — RFC 1952 is silent on these fields' limits. The cap:

  • matches the scale of the format's largest fixed-width header field, the 16-bit FEXTRA length (XLEN), whose maximum is 65535 bytes;
  • is orders of magnitude above real-world filename and comment sizes (Linux PATH_MAX is 4096, Windows MAX_PATH is 260), so no legitimate input is rejected.

The value is an independent cap; it is not derived from the DEFLATE decoder's 64 KiB internal buffer (#90). That bound governs how much decoded output is buffered before read returns, which is a separate concern from how large a header field is accepted. The two numbers happen to coincide, but they are not logically coupled.

Public API

No public API changes.

Validation

  • cargo test --workspace
  • cargo test --workspace --no-default-features
  • cargo clippy --lib --all-features -- -D warnings
  • cargo fmt --all -- --check

New regression test gzip_header_string_is_bounded covers F_NAME and F_COMMENT independently, asserting that a string of exactly 64 KiB + NUL is accepted losslessly and that 64 KiB + 1 byte is rejected with InvalidData.

Credits / Acknowledgment

Thanks to @optiklab for reporting the issue and for their proposed patch (optiklab@4ef7fb7). The fix here reimplements the same approach and threshold — functionally equivalent — rather than copying the diff verbatim.

@sile
sile merged commit 1c3495f into master Sep 6, 2026
38 checks passed
@sile
sile deleted the fix-issue-91-bound-gzip-header-strings branch September 6, 2026 09:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bound memory used by gzip FNAME and FCOMMENT parsing

1 participant