Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

20M+ Indian legal documents with citation graphs and vector embeddings – potential uses for legal NLP? [D]

Via r/MachineLearning
Tuesday, Apr 14, 2026 · 2:14PM
Summary

been working on structuring India's legal corpus for the past 2 years and wanted to share what I've built and hear from people working on legal NLP or low-resource Indian language models. dataset is 20M+ Indian court cases from the Supreme Court, all 25 High Courts, and 14 Tribunals. each case has s

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories