Feature idea: token budgets on the content filters, and why pruning drops constraints first #2301
behrnt-slatgng
started this conversation in
Feature requests
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure: I build Grunz, a hosted chat and coding agent on open weights, so I am downstream of exactly what Crawl4AI produces. Posting because the library's core premise — markdown sized for an LLM — runs into a number almost nobody has measured, and the pruning/BM25 filters are closer to solving it than their framing suggests.
The target you are pruning toward is smaller than the model card says
PruningContentFilterandBM25ContentFilterexist to get a page down to what is worth feeding a model. The implicit target is the model's context window. But the window advertised on a model card is a property of the weights, and the window you are served is whatever your provider configured for that endpoint.In practice most hosted endpoints serve around 32K tokens no matter what the card claims. A few reach 256K. I have not found one genuinely serving the 1M numbers that get quoted.
It fails silently — no error, no truncation notice, the front of the context is simply evicted.
So someone tuning
thresholdagainst a claimed 200K window is tuning against a ceiling that does not exist, and the failure surfaces as the model ignoring the system prompt (which sits at the front and gets evicted first) rather than as anything that looks like a size problem.Two things that would make the filters much more useful
Report token counts, not just character counts, on
fit_markdown. The library already computes the reduction; expressing it in tokens makes the filter's output directly comparable to the constraint it exists to satisfy. A rough tiktoken/o200k estimate is enough to be actionable.Allow a token budget as a filter parameter.
PruningContentFilter(target_tokens=30000)is a much more natural knob than a relevance threshold, because the budget is the thing the user actually knows and the threshold is a proxy they have to discover by trial. It also composes properly across a multi-page crawl, where the per-page threshold does not.Neither requires knowing anything about the user's model, which is what makes them worth doing.
The pruning step is itself a compaction, and that has a known failure mode
Worth naming, because it is the same operation that bites agent loops: any lossy reduction of content tends to drop constraints preferentially, because constraints are short, low-frequency, and look unimportant to a relevance scorer. In an extraction pipeline that means the qualifier — "deprecated since v3", "applies only to self-hosted", "do not use in production" — is exactly the kind of sentence a density- or BM25-based filter discards, while keeping the surrounding prose that now reads as unconditional.
The downstream model then answers confidently and wrongly, and the pruning step that caused it is not logged anywhere.
For extraction specifically, that argues for keeping negations, version qualifiers and scope-limiting clauses even when they score low, or at minimum making it visible what was dropped.
fit_markdowncurrently gives you the result without the diff.One for anyone running open weights behind this
Refusal-ablated variants lose instruction-following and output-format adherence before they lose knowledge — so with
LLMExtractionStrategyand a JSON schema, expect the schema violations before you expect wrong content. Prose quality holds while structure drifts, and perplexity will not warn you. Test format compliance separately from accuracy.Happy to go into detail on any of it.
All reactions