Skip to content

fix(chunking): measure combine thresholds in tokens when max_tokens is set - #4528

Open
joaquinhuigomez wants to merge 1 commit into
Unstructured-IO:mainfrom
joaquinhuigomez:joaquinhuigomez/chunking-token-mode-can-combine
Open

joaquinhuigomez wants to merge 1 commit into
Unstructured-IO:mainfrom
joaquinhuigomez:joaquinhuigomez/chunking-token-mode-can-combine

Conversation

@joaquinhuigomez

@joaquinhuigomez joaquinhuigomez commented Oct 4, 2026 •

Copy link
Copy Markdown

PreChunk.can_combine() sizes its text with len(), so the combine_text_under_n_chars threshold and the hard-max check count characters even when max_tokens has selected token counting. Every sibling computation in the class — will_fit, _text_length — goes through ChunkingOptions.measure(), which branches on the counting mode; can_combine is the one that doesn't. In token mode a section of a few tokens is many more characters than the token budget, so every section looks too large to combine and becomes its own chunk. With max_tokens=60, chunk_elements fills the window while chunk_by_title produces twelve ten-token chunks — about 17% utilization — and the obvious workaround, combine_text_under_n_chars=240, is rejected because it exceeds the hard max.

The fix is the two measure() substitutions. Character-mode behavior is unchanged (measure() is len() there): a 3,000-chunk-list corpus across both chunkers and five option sets hashes identically before and after, and the existing character-mode can_combine cases pass untouched.

Tests: a deterministic offline token counter fixture (no tiktoken download), a token-mode can_combine parametrization mirroring the character-mode sibling, a chunk_by_title case asserting the token budget is filled and the chunk texts equal chunk_elements() on the same input, and a discriminating character-mode case that would flip if the measure were swapped. The first two fail on main. test_unstructured/chunking/test_base.py + test_title.py: 391 passed, 16 skipped; ruff clean. CHANGELOG entry under a new 0.27.18 section with the matching __version__ bump, following the one-patch-per-PR pattern of the recent entries.

Related but different: #4487 changes what the combine_text_under_n_chars threshold is (caps it at new_after_n_chars); this changes which unit it's compared in. No textual overlap.

Review in cubic

`PreChunk.can_combine()` sized its text with `len()` and compared the
result to `combine_text_under_n_chars` and to `hard_max`. Both of those
are token counts when `max_tokens` is given, so the comparison was
characters against a token budget. Every sibling computation in the
class -- `.will_fit()`, `._text_length`, `_TextSplitter` -- already
routes through `ChunkingOptions.measure()`.

A section of a few tokens is many more characters than the token budget,
so `can_combine()` answered `False` for every short pre-chunk and
`chunk_by_title()` emitted one chunk per section. With `max_tokens=60`,
twelve ten-token sections came out as twelve chunks at 17% of the
requested size, where `chunk_elements()` on the same elements produced
two full ones. Raising `combine_text_under_n_chars` is not a way around
it: values above `max_tokens` are rejected by option validation.

Send both measurements through `measure()`. Character mode is
unaffected, because `measure()` is `len()` there; the existing
character-mode `can_combine()` cases still pass unchanged, and chunk
output over a 3,000-case character-mode corpus is byte-identical.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant