[b]"Best way to scrape whole website robots.txt – how to do it properly?"[/b] or [b]"How to scrape whole website r

18 Replies, 1039 Views

Thanks for all the tips, guys! I tried Scrapy with rotating proxies, and it worked way better than my old requests script.

Still figuring out how to parse the disallowed paths cleanly, but the `robotsparser` library someone mentioned is a game-changer.

Quick Q: Anyone know how to handle sites that block all bots but still have a robots.txt? Like, what’s the point? Lol.

Messages In This Thread



Users browsing this thread: 1 Guest(s)