Wow, thanks everyone for the awesome replies! I didn’t realize robot.txt was more about permissions than actual scraping. I’ll definitely check out Scrapy and BeautifulSoup since I’m comfortable with Python.
Also, thanks for the ethical advice. I’ll make sure to respect the site’s rules and not overload their servers.
One quick follow-up: How do you guys handle sites that block your IP? Is there a way to avoid that without being shady?
Cheers!
Also, thanks for the ethical advice. I’ll make sure to respect the site’s rules and not overload their servers.
One quick follow-up: How do you guys handle sites that block your IP? Is there a way to avoid that without being shady?
Cheers!
