docs(docker): note /crawl result order, and that app.workers is not read - #2292
Open
talelboussetta wants to merge 2 commits into
Open
talelboussetta wants to merge 2 commits into
talelboussetta wants to merge 2 commits into
Conversation
/crawl returns results in the order the pages finished: MemoryAdaptiveDispatcher.run_urls appends each task as it completes. Say so next to the Simple Crawl example in the Docker README and the self-hosting guide, and to match results by url rather than by position. app.workers in config.yml is not read anywhere: supervisord.conf starts gunicorn with a fixed --workers 1, and server.py's direct-run path does not pass it. Mark it the way the README already marks the direct-run-only port.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Scope ordering guidance to regular multi-URL crawls and clarify that ties are unspecified.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 2
Open (2)
What changed in this PR
Updates Docker documentation for /crawl result ordering and the unused app.workers setting.
Changes:
- Adds URL-based result matching guidance.
- Documents that
app.workersis not read by Docker startup.
| File | Summary |
|---|---|
docs/md_v2/core/self-hosting.md |
Adds /crawl ordering guidance. |
deploy/docker/README.md |
Adds result-order guidance. |
deploy/docker/config.yml |
Documents the unused worker setting. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…t it is not request order
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Two small corrections to the Docker server docs.
/crawlresult order.resultsare not in the order ofurls:MemoryAdaptiveDispatcher.run_urlsappends each result asasyncio.wait(..., FIRST_COMPLETED)returns its task, so a slow page comes back after a fast one, and pages that finish together come back in no set order. Nothing in the Docker README or the self-hosting guide says so, and indexingresults[i]againsturls[i]files one page under another page's URL with no error. A note after the Simple Crawl example, in both, says to match byurl.app.workersinconfig.ymlis not read.supervisord.confstarts gunicorn with a fixed--workers 1, andserver.py's direct-runuvicorn.run(...)does not pass it. Marked the way the README already marks the direct-run-onlyport.List of files changed and why
deploy/docker/README.md,docs/md_v2/core/self-hosting.md: the result-order note.deploy/docker/config.yml: a comment onapp.workers.How Has This Been Tested?
Docs and a YAML comment only;
config.ymlstill loads, and the offline suites are unchanged againstdevelop. The order, on an image built fromdevelop@ 1f68e5b, against a local test site, soCRAWL4AI_ALLOW_INTERNAL_URLS=true: two URLs, the first held back 3 s by its own per-URL config:cc @ntohidi
Checklist: