![]() |
|
How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob - Printable Version +- Proxy Community (https://proxycommunity.com/forum) +-- Forum: Use Case (https://proxycommunity.com/forum/forum-use-case) +--- Forum: Web Scraping (https://proxycommunity.com/forum/forum-web-scraping) +--- Thread: How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob (/thread-how-to-scrape-a-whole-website-s-robots-txt-file-properly-or-best-way-to-scrape-a-whole-website-s-rob) |
“” - MaskedPath99 - 28-03-2025 OP reply: Wow, thanks for all the suggestions! Tried Scrapy + proxies and it’s way faster than my old script. Still hitting some blocks though—might give Screaming Frog a shot next. Quick Q: Anyone know if Cloudflare always detects Scrapy, or can you sneak past with good headers? Also, big thanks for the wget tip—saved me hours! |