Repository navigation
Release v0.9.4 - #2281
Merged
Merged
Release v0.9.4#2281
Conversation
…er pruning, made default The bs4 PruningContentFilter recomputes get_text() and encode_contents() on every node while pruning, each walking the whole subtree -> O(N*subtree), which is super-linear on wide/deep pages (a 6000-card page spent ~2.2s inside filter_content, dominated by repeated subtree serialization). PruningContentFilterLXML computes every per-node metric (text length, inner-HTML length, link text, word count) once in a single bottom-up pass cached on a __slots__ object, then scores and prunes top-down -> O(N). Output is identical to the original: scoring math, thresholds, tag weights and the bs4 quirks (per-fragment stripping, separator-less word count, ASCII whitespace-only collapse, void-element self-closing, 256-depth cap lifted via huge_tree) are reproduced 1:1. Validated node-by-node against bs4 on real and pretty-printed pages plus a battery of adversarial shapes, and end-to-end (fit_markdown byte-identical). Made it the default where the code defaulted to PruningContentFilter: the Docker server FIT filter and the CLI pruning filter. Exported from the package and registered in the (de)serialization allowlists so configs round-trip. Speed (pruning only): medium 134->13ms, cards200 67->7ms, cards6000 2200->260ms.
…ult pruning engine
…umented default - PruningContentFilter now emits a DeprecationWarning on direct instantiation (suppressed when the PruningContentFilterLXML subclass calls super().__init__), pointing users to the lxml class and explaining the alias roadmap. - Lazily re-export PruningContentFilterLXML from content_filter_strategy via a module __getattr__ so existing import paths keep working with no circular import. - Update all active docs + example scripts to PruningContentFilterLXML and add deprecation/migration callouts to fit-markdown, markdown-generation and the SDK reference. Historical release notes (0.4.0) left unchanged.
page_timeout, wait_for_timeout and body_visibility_timeout are clamped to 60s for any config arriving over HTTP, and the value was a module literal with no env var, no config.yml key, and no way for an operator to change it. A deployment that is not public could not crawl a page that legitimately takes longer. CRAWL4AI_MAX_TIMEOUT_MS now sets the ceiling, defaulting to the same 60000ms. A value that is not a positive integer warns and keeps the default, since a typo would otherwise silently widen a DoS bound. Read per call rather than captured at import, so the setting applies wherever the process picked its environment up. Closes #2211
…yground md/llm runs died before their request was sent: the pre-flight sends the legacy 'code' field, which 0.9.3 rejects on untrusted requests (fixes #2222).
…ress proxy Two coordinated-disclosure reports, one root cause: a library HTTP client that fetches a caller-influenced URL without going through the egress broker. The browser, webhook and PDF paths were covered; these two channels were missed, and both are reachable from an untrusted Docker API body (check_robots_txt and link_preview_config are on the untrusted-config allowlist). - AsyncUrlSeeder built a plain httpx client. Nine request sites, three of them following redirects blind, and the fetched page's parsed <head> came back to the caller in result.links[*].head_data - so this one exfiltrates, it is not blind SSRF. - RobotsParser.can_fetch fetched robots.txt on a bare aiohttp session with ssl=False and redirects auto-followed, and cached the body for 7 days. Both now take proxy=proxy_url() from the new crawl4ai/egress_policy module, which deploy/docker/server.py points at the PinningProxy it already starts for Chromium. A proxy rather than a set_url_validator hook because a validator resolves the name, checks it, then throws the address away and lets the client re-resolve at connect time - the rebinding window egress_broker.py exists to close. The proxy dials the address it validated and re-checks every redirect hop, since each hop is a fresh proxied request. Unset (plain-library use) means proxy=None, the default on both httpx and aiohttp: no behaviour change, and HTTP/2 survives because the kwarg is proxy= and not transport=. Also here: - robots.txt no longer disables TLS verification. The except below fails open, so a bad certificate now means robots is ignored rather than respected. - LinkPreviewConfig gets untrusted clamps (max_links, concurrency, timeout). Each link is a separate outbound fetch and all three were unbounded. - egress_proxy sends Connection: close on the direct branch too, so a keep-alive client cannot put an absolute-form request line on the spliced origin socket. Same host either way, so this is correctness, not SSRF.
Pooled browsers get slower with sustained use and the janitor never recycles them, because it only closes browsers that have been *idle* past a TTL — and a server under continuous load never has an idle one. Measured where the slowdown actually lives: after 1000 real page loads on one pooled browser, a fresh context inside the SAME chromium process navigated as fast as a brand-new process (7.1ms vs 7.0ms), while the worn context took 32.0ms. Wiping that context's cookies/localStorage/ service workers in place recovered half of it (16.3ms). A control run with pages that leave no state behind stayed flat over 2000 pages. So the rot is context state, not the browser process, and the mechanism to bound it already exists in browser_manager (version-based recycling, which replaces the context under load without waiting for a quiet moment). It was simply never enabled for the Docker server. Enabling it is one config line — no second lifetime policy in the pool's janitor, which would put the same rule in two places. Measured effect at max_pages_before_recycle=200: penalty drops from ~30ms to ~10-14ms. It bounds the damage to one recycle window rather than removing it, so the threshold is the knob. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019EQGr8ey8riKFr4inp9B8f
…kwargs The previous commit put max_pages_before_recycle in config.yml's crawler.browser.kwargs, which does not reach the endpoints in the report: /crawl and /crawl/stream build their BrowserConfig from the request body (api.py:687, api.py:903) and nothing merges the server's browser kwargs into it. Only the handful of endpoints that call get_default_browser_config() would have picked it up. Move the setting to crawler.pool (it is a pool policy, next to max_pages and idle_ttl_sec) and apply it in get_crawler()/init_permanent(), which every endpoint goes through. Pooling is what makes a browser long-lived, so the pool is the right owner of its recycling policy. Applied before _sig() so all requests still share one signature and pooling is unchanged; an explicit per-request value wins. Tests cover the call site too, not only the helper — dropping _apply_pool_defaults() from get_crawler() now fails. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019EQGr8ey8riKFr4inp9B8f
…fig so its pool signature matches
Chrome-for-Testing SEGV_ACCERRs under --headless=new on macOS arm64 when OptimizationHints is disabled. Keep MediaRouter and DialMediaRouteProvider. Fixes #2239 Co-authored-by: Zsanz3 <Zsanz3@users.noreply.github.com>
…s past the untrusted gate
An untrusted request body could wrap a forbidden typed object in the plain-dict
envelope emitted by to_serializable_dict(). The unwrap recursed over
data["value"].items(), so the wrapped object was never seen as a whole typed
object and the UNTRUSTED_ALLOWED_TYPES gate never fired for it. from_kwargs()
then re-deserialized the surviving plain dict with the default TRUSTED
provenance, which constructs LLMConfig and resolves api_token="env:NAME"
through os.getenv - leaking OPENAI_API_KEY, and SECRET_KEY where it is set.
Two changes, defense in depth:
- from_serializable_dict(): after unwrapping {"type":"dict","value":X}, raise
UntrustedConfigError under UNTRUSTED if the result is itself typed-object
shaped. The shape test is factored into _is_typed_shape() and shared with the
existing dispatch, so both places agree on what a typed object looks like.
- BrowserConfig.from_kwargs() and CrawlerRunConfig.from_kwargs() take the
provenance and pass it down instead of silently defaulting to TRUSTED.
load() threads it through.
TRUSTED (SDK / in-process) behavior is unchanged.
Reported by Adam Jordan (adamyordan). CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N (8.1).
Test: tests/unit/test_config_provenance.py
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1ZJkE8DVxFirNUrbUQKrQ
…gv-fba7 fix(browser): drop OptimizationHints from --disable-features
…signature Build permanent browser with egress-hardened default config so its pool signature matches
Skip /config/dump pre-flight for md/llm endpoints in playground
…ders DefaultTableExtraction.extract_table_data collected body cells with `.//td`, so a `<th scope="row">` key column was dropped: every following cell shifted left and the row was padded with a phantom trailing "". `rowspan` was ignored entirely, so a spanning value never reached the rows it covers. Body rows now go through _build_grid, which lays each cell into the rectangle it spans -- including slots in rows not read yet -- and lets a later row fill whatever slots are still free. No span state is carried between iterations, so a short row cannot leave a stale span behind to overwrite a real value one row further down. The first row is only skipped as an implicit header when it holds no <td>. A first row mixing `<th scope="row">` with `<td>` is a key/value data row (the shape Wikipedia infoboxes use) and stays in `rows`, so the set of rows is unchanged from before -- only the dropped key column comes back. Scoring in is_data_table is untouched. Fixes #2258. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
…editor
The Advanced Config panel still posted {type, code} to /config/dump. `code` is
forbidden under the untrusted trust boundary, so the pre-flight always 400'd:
UntrustedConfigError: field 'code' is not permitted on CrawlerRunConfig
from an untrusted request
A regex fallback in runCrawl() then scraped the editor text for `stream=True`
and sent {params: {stream: true}} instead, so crawl runs succeeded while
silently discarding every other field the user typed. BrowserConfig has no
`stream`, so no fallback fired and the run aborted outright.
/config/dump stopped eval-ing snippets on purpose -- it was a
gadget-construction oracle -- so the request shape could not simply be patched.
The editor now holds the `params` object itself, which is what the endpoint
accepts today. Bad JSON and rejected fields both stop the run and report the
server's message, naming the offending field.
- CodeMirror switched to JSON mode; templates reseeded from
UNTRUSTED_FIELD_ALLOWLIST. The old BrowserConfig template used `extra_args`,
which is not allowlisted and would have 400'd under any request shape.
- pyConfigToJson -> validateConfig, parsing locally before the round trip.
- Regex fallback removed; a rejected config aborts instead of shrinking.
- #cfg-status is cleared when the endpoint changes, so a stale
'✖ config error' no longer lingers over md/llm, which skip validation.
Fixes #2260.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
…ig is empty Two defects found while driving the panel in a real browser against the Docker server: - /config/dump's echo is not re-submittable. dump() fills in server-derived fields, and BrowserConfig gains a generated `headers` dict that is not on UNTRUSTED_FIELD_ALLOWLIST. Posting the echo to /crawl therefore 400'd with "field 'headers' is not permitted on BrowserConfig". CrawlerRunConfig only escaped by luck: every field dump() adds there happens to be allowlisted. /config/dump is a validator, so the page now forwards what the user typed and keeps the echo as a pass/fail signal. - An empty editor returned early without touching #cfg-status, so a previous run's '✖ config error' stayed on screen while the run went ahead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
…itor fix(playground): replace the Python config editor with a JSON params editor
_build_grid walked rowspan * colspan slots per cell. put() ignored writes past
the last row, but the loop still spun, so the cost was the product of the two
attributes rather than of the table. develop never read rowspan at all, so this
arrived with the grid.
Measured, one cell, two-row table:
<td rowspan="1000000" colspan="1000000"> develop 0.02s grid >10s, no end
<td colspan="10000000"> develop 0.10s grid 2.79s
Scraped HTML is untrusted input, so a crafted cell hangs the crawl -- and on the
Docker server that holds a pooled browser.
Both spans are now clamped: colspan to the HTML Standard's cap of 1000, which is
what browsers apply, and rowspan to the rows that actually exist, a tighter bound
than the standard's 65534. All three cases above drop to 0.00s and the cost is
linear in the grid again.
Junk values are read the way a browser reads them, as 1, instead of raising:
colspan="100%" used to escape extract_table_data as a ValueError.
Reported in review of #2261.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016bfBtYmF5ExJctGcqpKvWv
fix(tables): expand rowspan/colspan into a grid and keep <th> row headers
…tions The "Send to Google Apps Script (Stars only)" step uses `curl -fSs`, so an HTTP error from GOOGLE_SCRIPT_ENDPOINT exits 22 and kills the job before the Discord notification runs. The endpoint is currently returning 403, so every `watch` event fails. Spreadsheet tracking is best-effort, so mark the step continue-on-error. Fixes #2255 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Utx8BTDqXTscjquLV7vgLq
…rkflow fix(ci): don't let the stargazer tracking step block Discord notifications
…irst enqueue BFS matched each fetched result back to its parent with `next((parent for (u, parent) in current_level if u == url), None)`, re-scanning the whole level once per result. On a page with high fan-out that pure-Python bookkeeping dominates the level. current_level is already a list of (url, parent) pairs, so build the lookup once per level instead. reversed() keeps the old first-parent-wins behaviour if a resumed level repeats a URL. BestFirst only added a URL to `visited` when it was dequeued, so two pages linking to the same third page pushed it onto the priority queue twice. The duplicate was dropped at dequeue time, so output stayed correct, but the URL was scored and queued for nothing. BFS and DFS both already mark URLs at discovery time; make BestFirst match and the dequeue guard becomes unreachable. Restoring a checkpoint now de-dupes the saved queue and treats everything queued as seen, so states written by older builds still resume without a repeat crawl. Fixes #2242 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Utx8BTDqXTscjquLV7vgLq
Make the untrusted timeout ceiling configurable
Follow-up to #2212. Three loose ends from review: - _cap_timeout returned the configured ceiling for any non-numeric or non-positive value, so raising CRAWL4AI_MAX_TIMEOUT_MS also raised what junk input became. An untrusted caller sending page_timeout 0 or "abc" got the full raised ceiling of held browser page. It now falls back to the 60s default, still bounded by a tightened ceiling. - MIGRATION.md's 300000ms example collides with limits.wall_clock_s (300) and crawler.timeouts.batch_process (300.0), so an operator following it exactly still got a 504 at 300s. Say to raise those too. - "60_000" sat in the "not a positive integer" parametrize list, but int("60_000") is 60000, so it was accepted, not refused. The test passed only because that value equals the default. Moved to its own accepted case and added coverage for the clamp fallback. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PVaRrydnAhwnnRYpe5an1U
Keep malformed timeouts at the default ceiling
…irst links The first cut of the best-first de-duplication marked a URL seen the moment it was discovered, which froze it at the depth it was *first* found rather than the shallowest depth reachable. Best-first does not visit levels in order, so a URL can be discovered at depth 3 through a high-scoring branch and only later at depth 2 through a slower one. Freezing depth 3 means link_discovery bails at 4 > max_depth and that whole subtree is silently lost. The duplicate enqueue was doing double duty: it was also the depth relaxation. So de-duplicate on depth instead of on identity. A URL already queued at an equal or shallower depth is skipped, which is the waste the issue reported; a strictly shallower re-discovery still re-queues, and the existing dequeue guard drops the stale deeper copy. This reverts the `visited`-at-discovery change, the removed dequeue guard and the resume-state compatibility shim, none of which are needed now: `visited` keeps its original crawled-only meaning and the whole fix is one guard in link_discovery. Found in review of #2265. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Utx8BTDqXTscjquLV7vgLq
fix(deep-crawl): drop the O(n²) parent scan and the duplicate best-first enqueue
…versions-v2 docs: list 0.9.x as supported
fix: preventing disallow: /*? from blocking the whole website.
load_config() deep-merges config.yml over DEFAULT_CONFIG, so a mounted custom config.yml without the new key left RECYCLE_PAGES at 0 and kept the #2231 bug after upgrade. An explicit 0 in config.yml still disables recycling.
…-recycle fix(docker): recycle pooled browser contexts by pages served (#2231)
* fix(robots): don't patch robotparser on Python 3.14+
The wildcard monkey patch in utils.py overrides RuleLine.applies_to
unconditionally. Python 3.14 rewrote urllib.robotparser with native
wildcard, '$' and RFC 9309 longest-match support, where applies_to
returns the match *length* used to rank competing rules. The patch
returns a bool, so every wildcard rule collapses to the lowest
priority and Allow: overrides stop working:
User-agent: *
Disallow: /
Allow: /public/*.html
denied /public/a.html on 3.14. Gate the patch to Python < 3.14, where
robotparser has no wildcard support and still needs it.
Follow-up to #2229, which fixed the 'Disallow: /*?' half of #2225.
Refs #2225
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(robots): don't rewrite bare-'?' rules on Python 3.14+ either
Review follow-up on the previous commit, which gated the RuleLine
monkey patch but left _preserve_bare_query running on every Python.
On 3.14 that rewrite is not just unnecessary, it is wrong. The stdlib
ranks rules by match length, and '/*?*' matches to end of string, so
it outranks a narrower competing Allow:
User-agent: *
Allow: /*?q=
Disallow: /*?
/search?q=1 is allowed by the stdlib and denied after the rewrite, so
Crawl4AI skipped pages robots.txt permits. Gate the call with the same
sys.version_info < (3, 14) as the patch, and say so in the docstring,
which claimed the two forms were always equivalent.
Also move the RuleLine import inside the branch that uses it and drop
a duplicate 'import re'.
Refs #2225
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Apps Script endpoint that logged stargazers to a Google Sheet has been returning 403 on every star since at least #2255, and the sheet is no longer used. #2263 made the step non-blocking, but that just leaves a curl that fails on every single star forever. Remove the step entirely. The Discord stargazer notification is unaffected. The GOOGLE_SCRIPT_ENDPOINT repo secret is now unused and can be deleted. Refs #2255
…-script chore(ci): drop the dead Google Apps Script stargazer step
….9.3 SECURITY.md stopped at v0.8.1; add per-release tables for v0.8.5 to v0.9.3, tag features by version, and note that fixes are not backported.
docs(security): list fixed issues and features per release through v0.9.3
Bump the version to 0.9.4 and add the v0.9.4 fixes to SECURITY.md.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Release v0.9.4. Already tagged and published to PyPI from
133e1d9. This PR bringsmainup to the release.{"type": "dict"}wrappers.PruningContentFilterLXML, about 10x faster pruning with identical output, now the default.PruningContentFilteris deprecated.rowspan/colspan, robots.txt rules, pooled browser recycling, and the Docker Playground.No breaking changes.
Release notes: docs/blog/release-v0.9.4.md
origin/mainis an ancestor of this branch, so it merges as a fast-forward.