What's the best way to use Python to clean HTML into a text file? or How can I clean HTML into a text

18 Replies, 853 Views

Honestly, regex is a bad idea for HTML. It’s like using a spoon to dig a hole.

BeautifulSoup is the gold standard, but if you’re lazy (like me), just use `lxml`:
```python
from lxml import html
doc = html.parse("file.html")
print(doc.text_content())
```

Faster than BS4 and still gets the job done for python clean html into text file.

Messages In This Thread



Users browsing this thread: 1 Guest(s)