Skip to content

fix(api): record robots.txt denials on crawls so the status warning and robotsBlocked work - #4900

Draft
claude[bot] wants to merge 1 commit into
mainfrom
claude/api-robots-denial-flag
Draft

claude[bot] wants to merge 1 commit into
mainfrom
claude/api-robots-denial-flag

Conversation

@claude

@claude claude Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Requested by Micah Stairs · Slack thread

Summary

Before: For most crawls, GET /v2/crawl/:id never showed the robots.txt warning, and robotsBlocked on GET /v2/crawl/:id/errors (and v1) stayed empty. The worker recorded a robots block only when the denial text was exactly "URL blocked by robots.txt". Only the flag-gated scrape robots check uses that text. The crawler's own robots denial uses a different text, so default crawls recorded nothing. Discovered links had a second problem: link extraction dropped robots-blocked links before filterLinks could report them. The warning also told users to use /scrape, which does not help.

After: A start URL or a discovered link that robots.txt disallows goes into the robotsBlocked list, and the status shows the warning. The new warning is: "One or more pages could not be crawled because the site's robots.txt disallows them. See the robotsBlocked list on GET /v2/crawl/{id}/errors. Teams with the feature enabled can set ignoreRobotsTxt: true." The user-facing denial texts on each URL do not change.

How

  • crawler.ts: add isRobotsDenialReason(), which compares to DenialReason.ROBOTS_TXT. Use the enum in place of the duplicate literal in the JS fallback.
  • CrawlDenialError gets an optional { robots: true } marker, kept through serialize and deserialize. The scrape robots check and the crawl start-URL robots check set it.
  • scrape-worker.ts: both recording sites use the helper or the marker, not the message text.
  • Link extraction calls filterURL without the robots check. filterLinks runs right after with the same robots.txt and user agent, so it still drops these links, and now it reports them.
  • crawl-status.ts (v2): new warning text. v1 has no robots warning, but v1 /errors reads the same Redis set, so it gets the fix too.
  • Test site: robots.txt disallows /robots-test/blocked. The new /robots-test/ pages are not in the sitemap, so other tests do not see them.

Tests

  • New snips v2/crawl-robots.test.ts: disallowed start URL, disallowed discovered link, and a negative case with no warning. The two positive tests fail on main and pass with this change.
  • New unit test WebScraper/__tests__/robots-denial.test.ts. scrape-robots.test.ts also checks the marker. crawl.test.ts expects the new warning text.
  • Local self-hosted harness: crawl-robots, scrape-robots and crawl snips pass. tsc and knip are clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_01H5YqtAzmu8fHaWAUDQLZa4


Generated by Claude Code


Summary by cubic

Tracks robots.txt denials on crawls so the status warning and the robotsBlocked list on the crawl errors endpoints now report blocked pages. Previously, only the flag-gated scrape robots check recorded a block, because it matched the exact message text "URL blocked by robots.txt", which the crawler's own denial never used, and discovered links were dropped before the robots check could flag them.

Bug Fixes

  • CrawlDenialError now carries a robots marker that survives serialization; recording sites use the marker or isRobotsDenialReason() instead of matching message text.
  • Link extraction now skips the robots check and lets filterLinks apply it, so denied links are reported instead of silently dropped.
  • The robots warning text now points to robotsBlocked on the errors endpoint instead of recommending /scrape.

Written for commit 87e0872. Summary will update on new commits.

Review in cubic

The crawl status robots warning and the robotsBlocked list on
/crawl/:id/errors matched one exact message text. Only the flag-gated
scrape robots check used that text, so default crawls recorded nothing.
Discovered links never reached the robots check either, because link
extraction dropped them first.

- Add isRobotsDenialReason() for the crawler robots denial.
- Give CrawlDenialError a robots marker that survives serialization.
  The scrape robots check and the start-URL robots check set it.
- Skip robots in link extraction. filterLinks applies robots after it
  and reports the denial.
- Reword the warning to point at robotsBlocked on the errors endpoint.
- Add a test-site robots fixture, snips tests and unit tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5YqtAzmu8fHaWAUDQLZa4

This branch was successfully deployed

1 active deployment
Preview — 87e08723 Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant