How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob

18 Replies, 2213 Views

Honestly, half the battle is avoiding IP bans.

When I scrape whole website robot.txt how to, I use Tor + requests in Python.

Slow? Yes. Effective? Also yes.

For big jobs, AWS Lambda with rotating IPs works wonders.

Messages In This Thread



Users browsing this thread: 1 Guest(s)