Hey! I’ve tried scraping Forbes Global 2000 before, and yeah, the JS rendering is a headache. Playwright worked way better for me than Scrapy because it handles dynamic content like a champ.
For anti-scraping, rotating proxies + adding random delays between requests saved me from getting blocked. Also, check if you can use their API—sometimes it’s hidden but exists.
Legal stuff? Grey area, but as long as you’re not hammering their servers, you’re probably fine.
For anti-scraping, rotating proxies + adding random delays between requests saved me from getting blocked. Also, check if you can use their API—sometimes it’s hidden but exists.
Legal stuff? Grey area, but as long as you’re not hammering their servers, you’re probably fine.
