Skip to content

fix(docker): apply crawler_configs to single-URL and streaming requests - #2290

Open
talelboussetta wants to merge 3 commits into
unclecode:developfrom
talelboussetta:bugfix/crawler-configs-single-url-and-stream
Open

talelboussetta wants to merge 3 commits into
unclecode:developfrom
talelboussetta:bugfix/crawler-configs-single-url-and-stream

Conversation

@talelboussetta

@talelboussetta talelboussetta commented Sep 25, 2026 •

Copy link
Copy Markdown

Summary

Fixes #2287

crawler_configs (#1852) was only honoured when a non-streaming /crawl carried two or more URLs. With one URL the handler called arun(), which takes a single config, and /crawl/stream (and /crawl with stream: true) never passed the list down. Both answered 200 having run with crawler_config alone.

  • A request that carries a list goes to arun_many whatever its URL count; arun_many already matches each URL against the list.
  • handle_stream_crawl_request takes crawler_configs and applies it; stream_process passes it. Each entry gets stream=True, since arun_many reads the stream flag from the first config of a list.
  • Loading the list moves into _load_crawler_configs(), shared by both handlers the way _normalize_and_validate_seeds() is, so both apply the same untrusted boundary and wire the PDF URL validator into every entry.
  • With a list applied, the list also decides the crawler (_needs_pdf_crawler()): a list of PDFContentScrapingStrategy entries runs on PDFCrawlerStrategy, and since one crawler serves every URL of a request, a list mixing PDF and browser entries is refused with 400 instead of running its PDF entries on the browser crawler. The non-streaming handler now loads the list before it picks a crawler.

@hafezparast: this reverses what test_single_url_ignores_crawler_configs pinned. I read the rationale there ("arun only takes one config") as the reason for the branch rather than a wish to drop the list, so the test is rewritten to assert the new routing. If ignoring it for one URL was deliberate for another reason, I'd like to know.

Not changed: on the streaming path base_config is still not applied (it isn't applied to crawler_config there either), and a streaming deep crawl still runs from crawler_config, as arun_many itself does. A per-URL deep_crawl_strategy can't reach either path: it is on the untrusted forbidden list, so such a body is rejected when it is loaded.

List of files changed and why

  • deploy/docker/api.py: _load_crawler_configs() and _needs_pdf_crawler(); single-URL routing through arun_many when a list is present; the streaming handler's new parameter.
  • deploy/docker/server.py: stream_process forwards crawler_configs.
  • deploy/docker/tests/test_crawler_configs_routing.py (new): behavioural tests against the real handlers with a recording crawler: one URL with a list, one URL without, streaming with a list, the endpoint forwarding it, a list of PDF entries running on the PDF crawler, and a mixed list refused on both paths.
  • tests/test_issue_1837_config_list.py: the two source assertions that pinned the old branch are updated to the new one.

How Has This Been Tested?

pytest deploy/docker/tests/test_crawler_configs_routing.py tests/test_issue_1837_config_list.py
21 passed       (6 of the 7 new tests fail on develop; the 7th is the no-list control)

The offline suites (deploy/docker/tests without the live-server test_1–test_7 scripts, tests/unit, the pool tests) show no new failures against develop.

End to end, on an image built from develop @ 1f68e5b (docker build -t c4ai-dev .), against a local test site, so CRAWL4AI_ALLOW_INTERNAL_URLS=true, before and with this branch's api.py and server.py mounted over /app. The page has two sections, .only-a and .only-b; the list sets css_selector: ".only-a" with url_matcher: "*":

# develop
[configs] /crawl, one URL, per-URL css_selector=.only-a -> HTTP 200: A=True B=True
[configs] /crawl/stream, two URLs, same list -> HTTP 200: A=True B=True; A=True B=True

# this branch
[configs] /crawl, one URL, per-URL css_selector=.only-a -> HTTP 200: A=True B=False
[configs] /crawl/stream, two URLs, same list -> HTTP 200: A=True B=False; A=True B=False

cc @ntohidi

Checklist:

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • I have added/updated unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

The per-URL config list (unclecode#1837) was only honoured when a non-streaming
request carried two or more URLs. With one URL the handler called arun(),
which takes a single config, and /crawl/stream (and /crawl with stream=true)
never passed the list to handle_stream_crawl_request. Both requests ran with
crawler_config alone and still answered 200.

A request that carries a list now goes to arun_many whatever its URL count,
and the streaming handler takes and applies the list. Loading moves into
_load_crawler_configs(), shared by both handlers the way
_normalize_and_validate_seeds() is, so both apply the same trust boundary and
wire the PDF URL validator into every entry. On the streaming path each entry
gets stream=True, since arun_many reads the stream flag from the first config
of a list.

The legacy source-grep test that pinned the single-URL path to arun() is
updated to the new routing; behavioural tests are in
deploy/docker/tests/test_crawler_configs_routing.py.
Copilot AI lite review requested due to automatic review settings September 25, 2026 10:18

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The streaming path mishandles per-URL PDF crawler selection and silently ignores per-URL deep-crawl strategies.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

This PR updates Docker crawl routing so crawler_configs applies to single-URL and streaming requests.

Changes:

  • Centralizes config-list loading and validation.
  • Routes single-URL config lists through arun_many.
  • Forwards config lists through streaming handlers.
  • Adds and updates routing regression tests.
File Description
tests/​test_issue_1837_config_list.py Updates config-list routing expectations.
deploy/​docker/​tests/​test_crawler_configs_routing.py Adds routing behavior tests.
deploy/​docker/​server.py Forwards config lists to streaming requests.
deploy/​docker/​api.py Loads config lists and routes crawl requests.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread deploy/docker/api.py
With crawler_configs applied, the crawler was still chosen from the top-level
crawler_config alone, so a list entry using PDFContentScrapingStrategy ran on
the pooled browser crawler and failed before the PDF scraper could run.

_needs_pdf_crawler() decides from the list when there is one: a list of PDF
entries runs on PDFCrawlerStrategy, and since one crawler serves every URL of
a request, a list mixing PDF and browser entries is refused with 400 instead
of failing mid-crawl. The non-streaming handler now loads the list before it
picks a crawler, as the streaming one already did.
Wraps the new lines over black's limit. The source assertion in
test_issue_1837_config_list.py now matches the call rather than the whole
one-line conditional, so it no longer depends on how the line is wrapped.
No behaviour change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants