Thanks for all the tips, guys! I tried Scrapy with rotating proxies, and it worked way better than my old requests script.
Still figuring out how to parse the disallowed paths cleanly, but the `robotsparser` library someone mentioned is a game-changer.
Quick Q: Anyone know how to handle sites that block all bots but still have a robots.txt? Like, what’s the point? Lol.
Still figuring out how to parse the disallowed paths cleanly, but the `robotsparser` library someone mentioned is a game-changer.
Quick Q: Anyone know how to handle sites that block all bots but still have a robots.txt? Like, what’s the point? Lol.
