Thanks for all the tips, folks! Didn’t realize how sneaky some sites could be even with a permissive robots.txt.
I tried scraping all pages website robots.txt allows with Scrapy and added delays + proxies like some of you suggested. So far, no bans!
Quick follow-up: How do you guys handle CAPTCHAs when they pop up? Just switch IPs, or is there a smarter way?
I tried scraping all pages website robots.txt allows with Scrapy and added delays + proxies like some of you suggested. So far, no bans!
Quick follow-up: How do you guys handle CAPTCHAs when they pop up? Just switch IPs, or is there a smarter way?
