How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob

20 Replies, 1667 Views

OP reply:

Wow, thanks for all the suggestions!

Tried Scrapy + proxies and it’s way faster than my old script. Still hitting some blocks though—might give Screaming Frog a shot next.

Quick Q: Anyone know if Cloudflare always detects Scrapy, or can you sneak past with good headers?

Also, big thanks for the wget tip—saved me hours!



Users browsing this thread: 1 Guest(s)