Man, I wasted weeks trying to scrape without crawling first. Big mistake.
For my job board project, I had to crawl Indeed to get URLs, THEN scrape the job details.
Used Scrapy for both, but it’s a steep learning curve.
If you’re lazy (like me), maybe try Diffbot—it does both crawling and scraping automagically.
For my job board project, I had to crawl Indeed to get URLs, THEN scrape the job details.
Used Scrapy for both, but it’s a steep learning curve.
If you’re lazy (like me), maybe try Diffbot—it does both crawling and scraping automagically.
