Skip to content

fix(utils): treat a missing robots.txt (HTTP 404) as allow-all - #4171

Merged
barjin merged 3 commits into
apify:masterfrom
breken-ai:fix/robots-txt-404-allow-all
Oct 2, 2026
Merged

barjin merged 3 commits into
apify:masterfrom
breken-ai:fix/robots-txt-404-allow-all

Conversation

@breken-ai

Copy link
Copy Markdown
Contributor

RobotsTxtFile.load() throws on any non-2xx status before it reaches the 404 branch, so that branch can never run:

if (response.status < 200 || response.status >= 300) {
    throw new Error(`Failed to load robots.txt from ${url}: HTTP ${response.status}`);
}

if (response.status === 404) {
    return new RobotsTxtFile(url, { isAllowed() { return true; }, ... });
}

So for a site with no robots.txt, RobotsTxtFile.find() rejects instead of returning an allow-all file. This order came in with #3306, which replaced the got-scraping HTTPError handling.

With respectRobotsTxtFile: true, getRobotsTxtFileForUrl() catches the rejection and allows the URL, but it caches nothing. Every URL on that host fetches /robots.txt again and logs Failed to fetch robots.txt for request .... I measured this with a BasicCrawler and 5 addRequests() URLs on a host whose robots.txt returns 404: 5 robots.txt fetches and 5 warnings on master, 1 fetch and no warnings with this change. Crawling each request runs checkRobotsTxt and repeats the same thing.

The fix moves the 404 check above the generic non-2xx rejection. Other error statuses still reject as before. The new tests in packages/utils/test/robots.test.ts cover a 404 (allow-all, no sitemaps, no crawl delay) and a 500 (still rejects). The 404 test fails on master with Failed to load robots.txt from http://no-robots.com/robots.txt: HTTP 404.

This PR was prepared with help from an AI coding assistant; I reviewed the change and ran the tests.

breken-ai and others added 3 commits September 29, 2026 20:13
`RobotsTxtFile.load()` threw on every non-2xx status before reaching the
`404` branch, so that branch was unreachable. A site without robots.txt
made `RobotsTxtFile.find()` reject instead of returning an allow-all file.

With `respectRobotsTxtFile`, the crawler catches that rejection but does
not cache a result, so it re-fetched `/robots.txt` and logged a warning
for every URL on such a site.

Check for 404 before rejecting other non-2xx responses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follows RFC 9309, which treats an unavailable (4xx) robots.txt as allowing everything and an unreachable (5xx) one as disallowing everything.

@barjin barjin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm, thank you @breken-ai !

@barjin
barjin merged commit f3fd764 into apify:master Oct 2, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants