Proxy Community
How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob - Printable Version

+- Proxy Community (https://proxycommunity.com/forum)
+-- Forum: Use Case (https://proxycommunity.com/forum/forum-use-case)
+--- Forum: Web Scraping (https://proxycommunity.com/forum/forum-web-scraping)
+--- Thread: How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob (/thread-how-to-scrape-a-whole-website-s-robots-txt-file-properly-or-best-way-to-scrape-a-whole-website-s-rob)

Pages: 1 2 3


“” - MaskedPath99 - 28-03-2025

OP reply:

Wow, thanks for all the suggestions!

Tried Scrapy + proxies and it’s way faster than my old script. Still hitting some blocks though—might give Screaming Frog a shot next.

Quick Q: Anyone know if Cloudflare always detects Scrapy, or can you sneak past with good headers?

Also, big thanks for the wget tip—saved me hours!