If you’re trying to scrape all pages website robots.txt how to, here’s the deal: robots.txt doesn’t list pages—it tells crawlers where they *can’t* go.
But! If there’s a sitemap, that’s your shortcut. Tools like Screaming Frog can ingest it and spit out all the URLs.
Otherwise, you’ll need to crawl the site manually (painful, I know). Maybe start with the homepage and go from there?
But! If there’s a sitemap, that’s your shortcut. Tools like Screaming Frog can ingest it and spit out all the URLs.
Otherwise, you’ll need to crawl the site manually (painful, I know). Maybe start with the homepage and go from there?
