What's the best way to parse HTML with Python lxml? or Python lxml vs. BeautifulSoup – which one do y

20 Replies, 1244 Views

Lxml’s XPath is unbeatable for complex scraping. Once you get the hang of it, you can extract data in one line that’d take 10 with BS4.

But yeah, the docs are... not great. StackOverflow is my real documentation lol.

If you’re dealing with *clean* HTML, python lxml is the winner. For broken pages, BS4’s parser is a lifesaver.

Messages In This Thread



Users browsing this thread: 1 Guest(s)