fix(utils): treat a missing robots.txt (HTTP 404) as allow-all - #4171
Merged
Merged
Conversation
`RobotsTxtFile.load()` threw on every non-2xx status before reaching the `404` branch, so that branch was unreachable. A site without robots.txt made `RobotsTxtFile.find()` reject instead of returning an allow-all file. With `respectRobotsTxtFile`, the crawler catches that rejection but does not cache a result, so it re-fetched `/robots.txt` and logged a warning for every URL on such a site. Check for 404 before rejecting other non-2xx responses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follows RFC 9309, which treats an unavailable (4xx) robots.txt as allowing everything and an unreachable (5xx) one as disallowing everything.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RobotsTxtFile.load()throws on any non-2xx status before it reaches the404branch, so that branch can never run:So for a site with no robots.txt,
RobotsTxtFile.find()rejects instead of returning an allow-all file. This order came in with #3306, which replaced the got-scrapingHTTPErrorhandling.With
respectRobotsTxtFile: true,getRobotsTxtFileForUrl()catches the rejection and allows the URL, but it caches nothing. Every URL on that host fetches/robots.txtagain and logsFailed to fetch robots.txt for request .... I measured this with aBasicCrawlerand 5addRequests()URLs on a host whose robots.txt returns 404: 5 robots.txt fetches and 5 warnings onmaster, 1 fetch and no warnings with this change. Crawling each request runscheckRobotsTxtand repeats the same thing.The fix moves the 404 check above the generic non-2xx rejection. Other error statuses still reject as before. The new tests in
packages/utils/test/robots.test.tscover a 404 (allow-all, no sitemaps, no crawl delay) and a 500 (still rejects). The 404 test fails onmasterwithFailed to load robots.txt from http://no-robots.com/robots.txt: HTTP 404.This PR was prepared with help from an AI coding assistant; I reviewed the change and ran the tests.