[b]"Best way to scrape whole website robots.txt – how to do it properly?"[/b] or [b]"How to scrape whole website r

18 Replies, 1103 Views

If you’re new to this, try `scraperapi.com`. It handles proxies, headers, and CAPTCHAs for you.

Just plug in the URL, and it’ll fetch the robots.txt for you. No coding needed unless you wanna parse it yourself.

For parsing, I’d write a small script to extract disallowed paths. Not hard once you get the hang of it.

Messages In This Thread



Users browsing this thread: 1 Guest(s)