Hey, thanks for all the tips, everyone! This is super helpful.
I tried preprocessing a few books with Tesseract and regex like some of you suggested, and it’s already way cleaner. Still running into issues with footnotes, though—anyone have a magic script for that?
Also, love the idea of mixing in web data to avoid overfitting. Gonna give that a shot next.
Cheers!
I tried preprocessing a few books with Tesseract and regex like some of you suggested, and it’s already way cleaner. Still running into issues with footnotes, though—anyone have a magic script for that?
Also, love the idea of mixing in web data to avoid overfitting. Gonna give that a shot next.
Cheers!
