Skip to content

fix(models): default CrawlResult links and media to the scraped shape - #2289

Open
talelboussetta wants to merge 3 commits into
unclecode:developfrom
talelboussetta:bugfix/crawlresult-failed-shape
Open

talelboussetta wants to merge 3 commits into
unclecode:developfrom
talelboussetta:bugfix/crawlresult-failed-shape

Conversation

@talelboussetta

@talelboussetta talelboussetta commented Sep 25, 2026 •

Copy link
Copy Markdown

Summary

Fixes #2286

CrawlResult.links and .media defaulted to {}, so every result built without them (the robots.txt refusal, "All proxies failed", the exception path in arun, the anti-bot raw-HTML fallback) had no internal / external / images keys. They now default to the keys a scraped page already gets: {"internal": [], "external": []} and {"images": [], "videos": [], "audios": []} (arun moves media's tables to CrawlResult.tables on the success path, so it is not part of the shape). Changing the default covers every construction site, including ones added later, instead of patching four CrawlResult(...) calls.

Behaviour change to be aware of: a failed result's links is now {"internal": [], "external": []} rather than {}, so if result.links: is truthy for it. Iterating over it is unaffected.

List of files changed and why

  • crawl4ai/models.py: default_factory for links and media.
  • tests/unit/test_crawl_result_default_shape.py (new): a result built without links/media has the scraped shape; the default is not shared between instances; model_dump() keeps the keys; explicit values are kept.
  • docs/md_v2/api/crawl-result.md, docs/md_v2/core/crawler-result.md, and the CrawlResult sketch in /execute_js's docstring (deploy/docker/server.py): show the new defaults.

How Has This Been Tested?

pytest tests/unit/test_crawl_result_default_shape.py
4 passed        (3 of the 4 fail on develop)

The offline suites (deploy/docker/tests without the live-server test_1–test_7 scripts, tests/unit, the pool tests) show no new failures against develop.

End to end, on an image built from develop @ 1f68e5b: /crawl with check_robots_txt: true and two URLs of a local test site, so CRAWL4AI_ALLOW_INTERNAL_URLS=true, one disallowed by its robots.txt. Before, and with this branch's models.py mounted over the installed package:

# develop
[shape] HTTP 200
  /index.html            success=True status=200 links=['external', 'internal'] media=['audios', 'images', 'videos'] error=''
  /private/index.html    success=False status=403 links={} media={} error='Access denied by robots.txt'

# this branch
[shape] HTTP 200
  /index.html            success=True status=200 links=['external', 'internal'] media=['audios', 'images', 'videos'] error=''
  /private/index.html    success=False status=403 links={"internal": [], "external": []} media={"images": [], "videos": [], "audios": []} error='Access denied by robots.txt'

cc @nightcityblade @ntohidi

Checklist:

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • I have added/updated unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

CrawlResult.links and .media defaulted to {}. Successful crawls fill them from
the scraper's result, but the results AsyncWebCrawler builds itself -- the
robots.txt refusal, "All proxies failed", the exception path in arun, the
anti-bot raw-HTML fallback -- never set them, so those results had no
"internal"/"external" or "images" keys: result.links["internal"] raised
KeyError, and the Docker API returned "links": {} for exactly the entries of a
batch a client has to handle.

Default both to the keys a scraped page gets -- {"internal": [], "external": []}
and {"images": [], "videos": [], "audios": []} ("tables" is moved out of media
into CrawlResult.tables on the success path) -- so every construction site is
covered, including future ones.
Copilot AI lite review requested due to automatic review settings September 25, 2026 10:17

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The default-isolation test needs strengthening, and documentation must be updated.

Review effort: Lite
Findings: None

What changed in this PR

Updates CrawlResult defaults so failed results expose consistent links and media shapes.

Changes:

  • Adds per-instance defaults for links and media.
  • Adds regression tests for shape, serialization, isolation, and explicit values.
File Review finding
tests/​unit/​test_crawl_result_default_shape.py Moderate (1 vote): Test expectations should use fresh literals or explicitly verify nested containers are not shared.
crawl4ai/​models.py Nit (1 vote): Update user-facing documentation and schema snippets that still show {} defaults.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

…n test

The CrawlResult sketches in docs/md_v2/api/crawl-result.md,
docs/md_v2/core/crawler-result.md and the /execute_js docstring still showed
`= {}` for links and media. The isolation test now also asserts the nested
lists are distinct objects, and compares against fresh literals.
complete-sdk-reference.md repeats the CrawlResult sketch from
docs/md_v2/api/crawl-result.md and still showed `= {}` for links and media.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants