How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob

18 Replies, 2195 Views

OP reply:

Wow, thanks for all the tips! Didn’t expect so many options.

Tried the curl loop and it worked for a few sites, but got blocked on bigger ones.

Gonna test Scrapy + proxies next.

Quick Q: Anyone know if Cloudflare bans you faster for robots.txt scraping vs other pages?

Also, shoutout to the Node.js suggestion—didn’t even think of that!

Messages In This Thread



Users browsing this thread: 1 Guest(s)