Is it possible to scrape all pages of a website using robots.txt? What are the limits? or How effecti

14 Replies, 1727 Views

Thanks for all the tips, folks! Didn’t realize how sneaky some sites could be even with a permissive robots.txt.

I tried scraping all pages website robots.txt allows with Scrapy and added delays + proxies like some of you suggested. So far, no bans!

Quick follow-up: How do you guys handle CAPTCHAs when they pop up? Just switch IPs, or is there a smarter way?

Messages In This Thread



Users browsing this thread: 1 Guest(s)