fix(api): record robots.txt denials on crawls so the status warning and robotsBlocked work - #4900
Draft
claude[bot] wants to merge 1 commit into
Draft
claude[bot] wants to merge 1 commit into
claude[bot] wants to merge 1 commit into
Conversation
The crawl status robots warning and the robotsBlocked list on /crawl/:id/errors matched one exact message text. Only the flag-gated scrape robots check used that text, so default crawls recorded nothing. Discovered links never reached the robots check either, because link extraction dropped them first. - Add isRobotsDenialReason() for the crawler robots denial. - Give CrawlDenialError a robots marker that survives serialization. The scrape robots check and the start-URL robots check set it. - Skip robots in link extraction. filterLinks applies robots after it and reports the denial. - Reword the warning to point at robotsBlocked on the errors endpoint. - Add a test-site robots fixture, snips tests and unit tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H5YqtAzmu8fHaWAUDQLZa4
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Micah Stairs · Slack thread
Summary
Before: For most crawls,
GET /v2/crawl/:idnever showed the robots.txt warning, androbotsBlockedonGET /v2/crawl/:id/errors(and v1) stayed empty. The worker recorded a robots block only when the denial text was exactly "URL blocked by robots.txt". Only the flag-gated scrape robots check uses that text. The crawler's own robots denial uses a different text, so default crawls recorded nothing. Discovered links had a second problem: link extraction dropped robots-blocked links beforefilterLinkscould report them. The warning also told users to use/scrape, which does not help.After: A start URL or a discovered link that robots.txt disallows goes into the
robotsBlockedlist, and the status shows the warning. The new warning is: "One or more pages could not be crawled because the site's robots.txt disallows them. See the robotsBlocked list on GET /v2/crawl/{id}/errors. Teams with the feature enabled can set ignoreRobotsTxt: true." The user-facing denial texts on each URL do not change.How
crawler.ts: addisRobotsDenialReason(), which compares toDenialReason.ROBOTS_TXT. Use the enum in place of the duplicate literal in the JS fallback.CrawlDenialErrorgets an optional{ robots: true }marker, kept through serialize and deserialize. The scrape robots check and the crawl start-URL robots check set it.scrape-worker.ts: both recording sites use the helper or the marker, not the message text.filterURLwithout the robots check.filterLinksruns right after with the same robots.txt and user agent, so it still drops these links, and now it reports them.crawl-status.ts(v2): new warning text. v1 has no robots warning, but v1/errorsreads the same Redis set, so it gets the fix too.robots.txtdisallows/robots-test/blocked. The new/robots-test/pages are not in the sitemap, so other tests do not see them.Tests
v2/crawl-robots.test.ts: disallowed start URL, disallowed discovered link, and a negative case with no warning. The two positive tests fail on main and pass with this change.WebScraper/__tests__/robots-denial.test.ts.scrape-robots.test.tsalso checks the marker.crawl.test.tsexpects the new warning text.crawl-robots,scrape-robotsandcrawlsnips pass.tscandknipare clean.🤖 Generated with Claude Code
https://claude.ai/code/session_01H5YqtAzmu8fHaWAUDQLZa4
Generated by Claude Code
Summary by cubic
Tracks robots.txt denials on crawls so the status warning and the
robotsBlockedlist on the crawl errors endpoints now report blocked pages. Previously, only the flag-gated scrape robots check recorded a block, because it matched the exact message text "URL blocked by robots.txt", which the crawler's own denial never used, and discovered links were dropped before the robots check could flag them.Bug Fixes
CrawlDenialErrornow carries arobotsmarker that survives serialization; recording sites use the marker orisRobotsDenialReason()instead of matching message text.filterLinksapply it, so denied links are reported instead of silently dropped.robotsBlockedon the errors endpoint instead of recommending/scrape.Written for commit 87e0872. Summary will update on new commits.