Skip to content

Release v0.9.4 - #2281

Merged
unclecode merged 44 commits into
mainfrom
release/v0.9.4
Sep 23, 2026
Merged

unclecode merged 44 commits into
mainfrom
release/v0.9.4

Conversation

@ntohidi

@ntohidi ntohidi commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Summary

Release v0.9.4. Already tagged and published to PyPI from 133e1d9. This PR brings main up to the release.

  • Security: closes three coordinated-disclosure advisories. Two SSRF paths (robots.txt fetch and link preview through the URL seeder) now go through the Docker server's pinning egress proxy, and the untrusted-config gate re-checks {"type": "dict"} wrappers.
  • Performance: adds PruningContentFilterLXML, about 10x faster pruning with identical output, now the default. PruningContentFilter is deprecated.
  • Bug fixes: deep-crawl speed, tables with rowspan/colspan, robots.txt rules, pooled browser recycling, and the Docker Playground.

No breaking changes.

Release notes: docs/blog/release-v0.9.4.md

origin/main is an ancestor of this branch, so it merges as a fast-forward.

unclecode and others added 30 commits June 25, 2026 09:25
…er pruning, made default

The bs4 PruningContentFilter recomputes get_text() and encode_contents()
on every node while pruning, each walking the whole subtree -> O(N*subtree),
which is super-linear on wide/deep pages (a 6000-card page spent ~2.2s
inside filter_content, dominated by repeated subtree serialization).

PruningContentFilterLXML computes every per-node metric (text length,
inner-HTML length, link text, word count) once in a single bottom-up pass
cached on a __slots__ object, then scores and prunes top-down -> O(N).

Output is identical to the original: scoring math, thresholds, tag weights
and the bs4 quirks (per-fragment stripping, separator-less word count,
ASCII whitespace-only collapse, void-element self-closing, 256-depth cap
lifted via huge_tree) are reproduced 1:1. Validated node-by-node against
bs4 on real and pretty-printed pages plus a battery of adversarial shapes,
and end-to-end (fit_markdown byte-identical).

Made it the default where the code defaulted to PruningContentFilter: the
Docker server FIT filter and the CLI pruning filter. Exported from the
package and registered in the (de)serialization allowlists so configs
round-trip.

Speed (pruning only): medium 134->13ms, cards200 67->7ms, cards6000 2200->260ms.
…umented default

- PruningContentFilter now emits a DeprecationWarning on direct instantiation
  (suppressed when the PruningContentFilterLXML subclass calls super().__init__),
  pointing users to the lxml class and explaining the alias roadmap.
- Lazily re-export PruningContentFilterLXML from content_filter_strategy via a
  module __getattr__ so existing import paths keep working with no circular import.
- Update all active docs + example scripts to PruningContentFilterLXML and add
  deprecation/migration callouts to fit-markdown, markdown-generation and the SDK
  reference. Historical release notes (0.4.0) left unchanged.
page_timeout, wait_for_timeout and body_visibility_timeout are clamped to
60s for any config arriving over HTTP, and the value was a module literal
with no env var, no config.yml key, and no way for an operator to change
it. A deployment that is not public could not crawl a page that
legitimately takes longer.

CRAWL4AI_MAX_TIMEOUT_MS now sets the ceiling, defaulting to the same
60000ms. A value that is not a positive integer warns and keeps the
default, since a typo would otherwise silently widen a DoS bound.

Read per call rather than captured at import, so the setting applies
wherever the process picked its environment up.

Closes #2211
…yground

md/llm runs died before their request was sent: the pre-flight sends the legacy 'code' field, which 0.9.3 rejects on untrusted requests (fixes #2222).
…ress proxy

Two coordinated-disclosure reports, one root cause: a library HTTP client that
fetches a caller-influenced URL without going through the egress broker. The
browser, webhook and PDF paths were covered; these two channels were missed,
and both are reachable from an untrusted Docker API body (check_robots_txt and
link_preview_config are on the untrusted-config allowlist).

- AsyncUrlSeeder built a plain httpx client. Nine request sites, three of them
  following redirects blind, and the fetched page's parsed <head> came back to
  the caller in result.links[*].head_data - so this one exfiltrates, it is not
  blind SSRF.
- RobotsParser.can_fetch fetched robots.txt on a bare aiohttp session with
  ssl=False and redirects auto-followed, and cached the body for 7 days.

Both now take proxy=proxy_url() from the new crawl4ai/egress_policy module,
which deploy/docker/server.py points at the PinningProxy it already starts for
Chromium. A proxy rather than a set_url_validator hook because a validator
resolves the name, checks it, then throws the address away and lets the client
re-resolve at connect time - the rebinding window egress_broker.py exists to
close. The proxy dials the address it validated and re-checks every redirect
hop, since each hop is a fresh proxied request.

Unset (plain-library use) means proxy=None, the default on both httpx and
aiohttp: no behaviour change, and HTTP/2 survives because the kwarg is proxy=
and not transport=.

Also here:
- robots.txt no longer disables TLS verification. The except below fails open,
  so a bad certificate now means robots is ignored rather than respected.
- LinkPreviewConfig gets untrusted clamps (max_links, concurrency, timeout).
  Each link is a separate outbound fetch and all three were unbounded.
- egress_proxy sends Connection: close on the direct branch too, so a
  keep-alive client cannot put an absolute-form request line on the spliced
  origin socket. Same host either way, so this is correctness, not SSRF.
Pooled browsers get slower with sustained use and the janitor never
recycles them, because it only closes browsers that have been *idle*
past a TTL — and a server under continuous load never has an idle one.

Measured where the slowdown actually lives: after 1000 real page loads
on one pooled browser, a fresh context inside the SAME chromium process
navigated as fast as a brand-new process (7.1ms vs 7.0ms), while the
worn context took 32.0ms. Wiping that context's cookies/localStorage/
service workers in place recovered half of it (16.3ms). A control run
with pages that leave no state behind stayed flat over 2000 pages.

So the rot is context state, not the browser process, and the mechanism
to bound it already exists in browser_manager (version-based recycling,
which replaces the context under load without waiting for a quiet
moment). It was simply never enabled for the Docker server.

Enabling it is one config line — no second lifetime policy in the pool's
janitor, which would put the same rule in two places. Measured effect at
max_pages_before_recycle=200: penalty drops from ~30ms to ~10-14ms. It
bounds the damage to one recycle window rather than removing it, so the
threshold is the knob.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019EQGr8ey8riKFr4inp9B8f
…kwargs

The previous commit put max_pages_before_recycle in config.yml's
crawler.browser.kwargs, which does not reach the endpoints in the report:
/crawl and /crawl/stream build their BrowserConfig from the request body
(api.py:687, api.py:903) and nothing merges the server's browser kwargs
into it. Only the handful of endpoints that call
get_default_browser_config() would have picked it up.

Move the setting to crawler.pool (it is a pool policy, next to max_pages
and idle_ttl_sec) and apply it in get_crawler()/init_permanent(), which
every endpoint goes through. Pooling is what makes a browser long-lived,
so the pool is the right owner of its recycling policy.

Applied before _sig() so all requests still share one signature and
pooling is unchanged; an explicit per-request value wins.

Tests cover the call site too, not only the helper — dropping
_apply_pool_defaults() from get_crawler() now fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019EQGr8ey8riKFr4inp9B8f
Chrome-for-Testing SEGV_ACCERRs under --headless=new on macOS arm64 when
OptimizationHints is disabled. Keep MediaRouter and DialMediaRouteProvider.

Fixes #2239

Co-authored-by: Zsanz3 <Zsanz3@users.noreply.github.com>
…s past the untrusted gate

An untrusted request body could wrap a forbidden typed object in the plain-dict
envelope emitted by to_serializable_dict(). The unwrap recursed over
data["value"].items(), so the wrapped object was never seen as a whole typed
object and the UNTRUSTED_ALLOWED_TYPES gate never fired for it. from_kwargs()
then re-deserialized the surviving plain dict with the default TRUSTED
provenance, which constructs LLMConfig and resolves api_token="env:NAME"
through os.getenv - leaking OPENAI_API_KEY, and SECRET_KEY where it is set.

Two changes, defense in depth:

- from_serializable_dict(): after unwrapping {"type":"dict","value":X}, raise
  UntrustedConfigError under UNTRUSTED if the result is itself typed-object
  shaped. The shape test is factored into _is_typed_shape() and shared with the
  existing dispatch, so both places agree on what a typed object looks like.
- BrowserConfig.from_kwargs() and CrawlerRunConfig.from_kwargs() take the
  provenance and pass it down instead of silently defaulting to TRUSTED.
  load() threads it through.

TRUSTED (SDK / in-process) behavior is unchanged.

Reported by Adam Jordan (adamyordan). CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N (8.1).
Test: tests/unit/test_config_provenance.py

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1ZJkE8DVxFirNUrbUQKrQ
…gv-fba7

fix(browser): drop OptimizationHints from --disable-features
…signature

Build permanent browser with egress-hardened default config so its pool signature matches
Skip /config/dump pre-flight for md/llm endpoints in playground
…ders

DefaultTableExtraction.extract_table_data collected body cells with `.//td`,
so a `<th scope="row">` key column was dropped: every following cell shifted
left and the row was padded with a phantom trailing "". `rowspan` was ignored
entirely, so a spanning value never reached the rows it covers.

Body rows now go through _build_grid, which lays each cell into the rectangle
it spans -- including slots in rows not read yet -- and lets a later row fill
whatever slots are still free. No span state is carried between iterations, so
a short row cannot leave a stale span behind to overwrite a real value one row
further down.

The first row is only skipped as an implicit header when it holds no <td>. A
first row mixing `<th scope="row">` with `<td>` is a key/value data row (the
shape Wikipedia infoboxes use) and stays in `rows`, so the set of rows is
unchanged from before -- only the dropped key column comes back.

Scoring in is_data_table is untouched.

Fixes #2258.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
…editor

The Advanced Config panel still posted {type, code} to /config/dump. `code` is
forbidden under the untrusted trust boundary, so the pre-flight always 400'd:

    UntrustedConfigError: field 'code' is not permitted on CrawlerRunConfig
    from an untrusted request

A regex fallback in runCrawl() then scraped the editor text for `stream=True`
and sent {params: {stream: true}} instead, so crawl runs succeeded while
silently discarding every other field the user typed. BrowserConfig has no
`stream`, so no fallback fired and the run aborted outright.

/config/dump stopped eval-ing snippets on purpose -- it was a
gadget-construction oracle -- so the request shape could not simply be patched.
The editor now holds the `params` object itself, which is what the endpoint
accepts today. Bad JSON and rejected fields both stop the run and report the
server's message, naming the offending field.

- CodeMirror switched to JSON mode; templates reseeded from
  UNTRUSTED_FIELD_ALLOWLIST. The old BrowserConfig template used `extra_args`,
  which is not allowlisted and would have 400'd under any request shape.
- pyConfigToJson -> validateConfig, parsing locally before the round trip.
- Regex fallback removed; a rejected config aborts instead of shrinking.
- #cfg-status is cleared when the endpoint changes, so a stale
  '✖ config error' no longer lingers over md/llm, which skip validation.

Fixes #2260.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
…ig is empty

Two defects found while driving the panel in a real browser against the Docker
server:

- /config/dump's echo is not re-submittable. dump() fills in server-derived
  fields, and BrowserConfig gains a generated `headers` dict that is not on
  UNTRUSTED_FIELD_ALLOWLIST. Posting the echo to /crawl therefore 400'd with
  "field 'headers' is not permitted on BrowserConfig". CrawlerRunConfig only
  escaped by luck: every field dump() adds there happens to be allowlisted.
  /config/dump is a validator, so the page now forwards what the user typed
  and keeps the echo as a pass/fail signal.

- An empty editor returned early without touching #cfg-status, so a previous
  run's '✖ config error' stayed on screen while the run went ahead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
…itor

fix(playground): replace the Python config editor with a JSON params editor
_build_grid walked rowspan * colspan slots per cell. put() ignored writes past
the last row, but the loop still spun, so the cost was the product of the two
attributes rather than of the table. develop never read rowspan at all, so this
arrived with the grid.

Measured, one cell, two-row table:

    <td rowspan="1000000" colspan="1000000">   develop 0.02s   grid  >10s, no end
    <td colspan="10000000">                    develop 0.10s   grid  2.79s

Scraped HTML is untrusted input, so a crafted cell hangs the crawl -- and on the
Docker server that holds a pooled browser.

Both spans are now clamped: colspan to the HTML Standard's cap of 1000, which is
what browsers apply, and rowspan to the rows that actually exist, a tighter bound
than the standard's 65534. All three cases above drop to 0.00s and the cost is
linear in the grid again.

Junk values are read the way a browser reads them, as 1, instead of raising:
colspan="100%" used to escape extract_table_data as a ValueError.

Reported in review of #2261.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
fix(tables): expand rowspan/colspan into a grid and keep <th> row headers
…tions

The "Send to Google Apps Script (Stars only)" step uses `curl -fSs`, so an
HTTP error from GOOGLE_SCRIPT_ENDPOINT exits 22 and kills the job before the
Discord notification runs. The endpoint is currently returning 403, so every
`watch` event fails.

Spreadsheet tracking is best-effort, so mark the step continue-on-error.

Fixes #2255

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Utx8BTDqXTscjquLV7vgLq
…rkflow

fix(ci): don't let the stargazer tracking step block Discord notifications
…irst enqueue

BFS matched each fetched result back to its parent with
`next((parent for (u, parent) in current_level if u == url), None)`, re-scanning
the whole level once per result. On a page with high fan-out that pure-Python
bookkeeping dominates the level. current_level is already a list of
(url, parent) pairs, so build the lookup once per level instead. reversed()
keeps the old first-parent-wins behaviour if a resumed level repeats a URL.

BestFirst only added a URL to `visited` when it was dequeued, so two pages
linking to the same third page pushed it onto the priority queue twice. The
duplicate was dropped at dequeue time, so output stayed correct, but the URL was
scored and queued for nothing. BFS and DFS both already mark URLs at discovery
time; make BestFirst match and the dequeue guard becomes unreachable. Restoring
a checkpoint now de-dupes the saved queue and treats everything queued as seen,
so states written by older builds still resume without a repeat crawl.

Fixes #2242

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Utx8BTDqXTscjquLV7vgLq
Make the untrusted timeout ceiling configurable
Follow-up to #2212. Three loose ends from review:

- _cap_timeout returned the configured ceiling for any non-numeric or
  non-positive value, so raising CRAWL4AI_MAX_TIMEOUT_MS also raised what
  junk input became. An untrusted caller sending page_timeout 0 or "abc"
  got the full raised ceiling of held browser page. It now falls back to
  the 60s default, still bounded by a tightened ceiling.

- MIGRATION.md's 300000ms example collides with limits.wall_clock_s (300)
  and crawler.timeouts.batch_process (300.0), so an operator following it
  exactly still got a 504 at 300s. Say to raise those too.

- "60_000" sat in the "not a positive integer" parametrize list, but
  int("60_000") is 60000, so it was accepted, not refused. The test passed
  only because that value equals the default. Moved to its own accepted
  case and added coverage for the clamp fallback.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVaRrydnAhwnnRYpe5an1U
Keep malformed timeouts at the default ceiling
…irst links

The first cut of the best-first de-duplication marked a URL seen the moment it
was discovered, which froze it at the depth it was *first* found rather than the
shallowest depth reachable. Best-first does not visit levels in order, so a URL
can be discovered at depth 3 through a high-scoring branch and only later at
depth 2 through a slower one. Freezing depth 3 means link_discovery bails at
4 > max_depth and that whole subtree is silently lost.

The duplicate enqueue was doing double duty: it was also the depth relaxation.
So de-duplicate on depth instead of on identity. A URL already queued at an
equal or shallower depth is skipped, which is the waste the issue reported; a
strictly shallower re-discovery still re-queues, and the existing dequeue guard
drops the stale deeper copy.

This reverts the `visited`-at-discovery change, the removed dequeue guard and
the resume-state compatibility shim, none of which are needed now: `visited`
keeps its original crawled-only meaning and the whole fix is one guard in
link_discovery.

Found in review of #2265.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Utx8BTDqXTscjquLV7vgLq
fix(deep-crawl): drop the O(n²) parent scan and the duplicate best-first enqueue
Nalhin and others added 14 commits September 18, 2026 16:25
…versions-v2

docs: list 0.9.x as supported
fix: preventing disallow: /*? from blocking the whole website.
load_config() deep-merges config.yml over DEFAULT_CONFIG, so a mounted
custom config.yml without the new key left RECYCLE_PAGES at 0 and kept
the #2231 bug after upgrade. An explicit 0 in config.yml still disables
recycling.
…-recycle

fix(docker): recycle pooled browser contexts by pages served (#2231)
* fix(robots): don't patch robotparser on Python 3.14+

The wildcard monkey patch in utils.py overrides RuleLine.applies_to
unconditionally. Python 3.14 rewrote urllib.robotparser with native
wildcard, '$' and RFC 9309 longest-match support, where applies_to
returns the match *length* used to rank competing rules. The patch
returns a bool, so every wildcard rule collapses to the lowest
priority and Allow: overrides stop working:

    User-agent: *
    Disallow: /
    Allow: /public/*.html

denied /public/a.html on 3.14. Gate the patch to Python < 3.14, where
robotparser has no wildcard support and still needs it.

Follow-up to #2229, which fixed the 'Disallow: /*?' half of #2225.

Refs #2225

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(robots): don't rewrite bare-'?' rules on Python 3.14+ either

Review follow-up on the previous commit, which gated the RuleLine
monkey patch but left _preserve_bare_query running on every Python.

On 3.14 that rewrite is not just unnecessary, it is wrong. The stdlib
ranks rules by match length, and '/*?*' matches to end of string, so
it outranks a narrower competing Allow:

    User-agent: *
    Allow: /*?q=
    Disallow: /*?

/search?q=1 is allowed by the stdlib and denied after the rewrite, so
Crawl4AI skipped pages robots.txt permits. Gate the call with the same
sys.version_info < (3, 14) as the patch, and say so in the docstring,
which claimed the two forms were always equivalent.

Also move the RuleLine import inside the branch that uses it and drop
a duplicate 'import re'.

Refs #2225

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Apps Script endpoint that logged stargazers to a Google Sheet has
been returning 403 on every star since at least #2255, and the sheet is
no longer used. #2263 made the step non-blocking, but that just leaves a
curl that fails on every single star forever.

Remove the step entirely. The Discord stargazer notification is
unaffected. The GOOGLE_SCRIPT_ENDPOINT repo secret is now unused and can
be deleted.

Refs #2255
…-script

chore(ci): drop the dead Google Apps Script stargazer step
….9.3

SECURITY.md stopped at v0.8.1; add per-release tables for v0.8.5 to v0.9.3, tag features by version, and note that fixes are not backported.
docs(security): list fixed issues and features per release through v0.9.3
Bump the version to 0.9.4 and add the v0.9.4 fixes to SECURITY.md.
@unclecode
unclecode merged commit 86e6464 into main Sep 23, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants