How to scrape a whole website's robots.txt file properly? or Best way to scrape a whole website's rob

18 Replies, 2210 Views

If you’re into Node.js, "robots-parser" npm package is gold.

Fetch robots.txt, then parse rules programmatically.

For scrape whole website robot.txt how to, it’s way cleaner than regex hacking.

Bonus: It handles wildcards and disallows neatly.

Messages In This Thread



Users browsing this thread: 1 Guest(s)