Dude, I feel you. Web scraping in R was a headache at first. Here’s what worked for me:
- Use `rvest::html_nodes()` with Chrome’s "Copy selector" (right-click element in Inspect).
- For 403s, rotate user agents or try `httr:
et_config(httr::use_proxy())` if you’re blocked.
If rvest fails, `RSelenium` is your backup—but it’s slower. Also, `robotstxt` package tells you if scraping’s even allowed.
- Use `rvest::html_nodes()` with Chrome’s "Copy selector" (right-click element in Inspect).
- For 403s, rotate user agents or try `httr:
et_config(httr::use_proxy())` if you’re blocked. If rvest fails, `RSelenium` is your backup—but it’s slower. Also, `robotstxt` package tells you if scraping’s even allowed.
