What's the best way to handle dynamic content when doing R scraping? or How can I avoid getting block

18 Replies, 1834 Views

R scraping for 10k+ pages? Doable but messy.

`rvest` + `httr` works, but Python’s `scrapy` is way faster for bulk jobs. If you’re stuck with R, split the workload and use `future.apply` for parallel batches.

Also, watch out for IP bans—rotate proxies aggressively.

Messages In This Thread



Users browsing this thread: 1 Guest(s)