Best Way to Scrape All Pages from a Website Using robot.txt – How to Do It?

16 Replies, 1386 Views

Yo, I’ve been in the same boat! Robot.txt is basically a “do not enter” sign for certain pages, so scraping all pages from a website robot.txt how to do it isn’t really the right approach.

I’ve used Octoparse before, and it’s pretty solid for scraping without getting banned. Just make sure to set delays between requests. Ethical? Depends on what you’re doing with the data. If it’s for research or personal use, you’re good.

Messages In This Thread



Users browsing this thread: 1 Guest(s)