Is it possible to scrape all pages of a website using robots.txt? What are the limits? or How effecti

14 Replies, 1586 Views

Title: Can you scrape all pages website robots.txt allows it?

Hey folks,

Quick question—if a site’s robots.txt doesn’t block scraping, does that mean it’s *totally* fine to scrape all pages website robots.txt permits? Or are there hidden limits?

I’ve seen some sites with super permissive rules, but I’m not sure if crawling everything won’t get me banned anyway. Like, just because you *can* scrape all pages website robots.txt allows, should you?

Also, what’s the best way to do it without hammering their servers? Slow and steady, or are there better tricks?

Kinda new to this, so any tips or horror stories welcome!

Thanks!

Messages In This Thread
Is it possible to scrape all pages of a website using robots.txt? What are the limits? or How effecti - by - 13-01-2025, 08:35 PM



Users browsing this thread: 1 Guest(s)