BeautifulSoup is great for *parsing html with python*, but if you’re looking for speed, lxml is the way to go.
Regex is a no-no for HTML. It’s just not worth the hassle.
For messy HTML, I’d recommend using `html5lib` with BeautifulSoup. It’s slower but handles broken tags really well.
Also, check out Scrapy if you’re doing more than just parsing. It’s a full framework and makes scraping way easier.
Regex is a no-no for HTML. It’s just not worth the hassle.
For messy HTML, I’d recommend using `html5lib` with BeautifulSoup. It’s slower but handles broken tags really well.
Also, check out Scrapy if you’re doing more than just parsing. It’s a full framework and makes scraping way easier.
