Best Way to Scrape All Pages from a Website Using robot.txt – How to Do It?

16 Replies, 1387 Views

Hey, robot.txt is more of a guideline for what you *shouldn’t* scrape, not a roadmap for scraping all pages from a website.

I’ve had good luck with DataMiner for smaller projects. It’s super user-friendly and doesn’t require coding. For bigger stuff, Scrapy is the way to go, but you’ll need to tweak the settings to avoid getting banned.

Ethics? If you’re not causing harm or stealing, it’s probably okay. Just don’t be a jerk about it.

Messages In This Thread



Users browsing this thread: 1 Guest(s)