[b]"How to scrape a whole website's robots.txt file properly?"[/b] or [b]"Best way to scrape a whole website's rob

20 Replies, 1448 Views

OP reply:

Wow, thanks for all the suggestions!

Tried Scrapy + proxies and it’s way faster than my old script. Still hitting some blocks though—might give Screaming Frog a shot next.

Quick Q: Anyone know if Cloudflare always detects Scrapy, or can you sneak past with good headers?

Also, big thanks for the wget tip—saved me hours!



Users browsing this thread: 1 Guest(s)