Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Best approach for OCR → cleaning → Hindi–English translation pipeline? [R]

Via r/MachineLearning
Tuesday, Apr 14, 2026 · 6:49AM
Summary

I’m working with a large CSV dataset (~200k rows) containing OCR-extracted text from parliamentary documents. The data includes a mix of Hindi (Devanagari) and English, often within the same sentence, along with OCR noise (broken words, symbols, formatting artifacts). I’m looking for suggestions on:

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories