Which python xml parser is the most efficient for large files? or How do you handle XML parsing in Py

22 Replies, 1795 Views

For quick and dirty stuff, I use `xmltodict`. It’s not the fastest python xml parser, but it’s so easy.

For big files, though, lxml + `iterparse` is the way. Just make sure to `clear()` as you go!

Also, `pyarrow` can help if you’re moving data to pandas later.
Wow, thanks for all the suggestions! I gave lxml a shot with `iterparse`, and it’s *way* faster. Still getting the hang of clearing nodes properly, though—any good examples for that?

Also, `xmltodict` looks perfect for some smaller files I’m handling. Gonna try that next.

Really appreciate the tips!



Users browsing this thread: 1 Guest(s)