How to scrape all pages from a website using robots.txt? Need guidance! or Best way to scrape all pag

18 Replies, 1665 Views

Thanks for all the tips, guys! I checked the robots.txt and found a sitemap link—didn’t even notice it before lol.

Gonna try Screaming Frog first since a few of you recommended it. If that doesn’t work, I’ll dig into Scrapy.

Quick Q: How long should I set the delay between requests to avoid getting blocked? 2 seconds enough?

(And yeah, I’ll avoid the disallowed paths—don’t wanna be *that* guy.)

Messages In This Thread



Users browsing this thread: 1 Guest(s)