Yo, I’ve been in the same boat! Robot.txt is basically a “do not enter” sign for certain pages, so scraping all pages from a website robot.txt how to do it isn’t really the right approach.
I’ve used Octoparse before, and it’s pretty solid for scraping without getting banned. Just make sure to set delays between requests. Ethical? Depends on what you’re doing with the data. If it’s for research or personal use, you’re good.
I’ve used Octoparse before, and it’s pretty solid for scraping without getting banned. Just make sure to set delays between requests. Ethical? Depends on what you’re doing with the data. If it’s for research or personal use, you’re good.
