"Has anyone had success with training an LLM using books? Tips welcome!"
Hey folks,
So I’ve been digging into training an LLM using books—classic novels, textbooks, you name it. Seems like a goldmine for high-quality text, but I’m kinda stuck on the best way to prep the data.
Anyone here actually pulled this off? How’d you handle formatting issues (like footnotes or weird OCR errors)? And is it worth the effort compared to just scraping the web?
Also, did you chunk the text by chapters or just throw it all in raw? Low-key worried about overfitting if the model gets too cozy with one author’s style lol.
Would love to hear your war stories or any gotchas!
Cheers.
Hey folks,
So I’ve been digging into training an LLM using books—classic novels, textbooks, you name it. Seems like a goldmine for high-quality text, but I’m kinda stuck on the best way to prep the data.
Anyone here actually pulled this off? How’d you handle formatting issues (like footnotes or weird OCR errors)? And is it worth the effort compared to just scraping the web?
Also, did you chunk the text by chapters or just throw it all in raw? Low-key worried about overfitting if the model gets too cozy with one author’s style lol.
Would love to hear your war stories or any gotchas!
Cheers.
