Best Practices for Web Scraping in R: How to Handle Dynamic Content and Avoid Getting Blocked?

16 Replies, 1660 Views

Wow, thanks for all the tips, everyone! RSelenium seems to be the consensus for dynamic content, so I’ll definitely give that a shot. I’ve been avoiding it because it seemed complicated, but I guess it’s time to dive in.

Also, the `V8` package sounds interesting—I’ll check that out too. And thanks for the proxy suggestions—I’ve been using free ones, but maybe it’s time to invest in something better.

One follow-up question: does anyone have a good tutorial or guide for setting up RSelenium? I’m a bit lost on where to start.

Thanks again, y’all are awesome!

Messages In This Thread



Users browsing this thread: 1 Guest(s)