Applications
20M+ Indian legal documents with citation graphs and vector embeddings – potential uses for legal NLP? [D]
A r/MachineLearning discussion thread exploring the potential NLP applications of a large-scale dataset comprising over 20 million Indian legal documents, enriched with citation graphs and pre-compute
A r/MachineLearning discussion thread exploring the potential NLP applications of a large-scale dataset comprising over 20 million Indian legal documents, enriched with citation graphs and pre-computed vector embeddings. The post likely surfaces use cases such as legal judgment prediction, case similarity retrieval, citation link prediction, and document summarization — tasks well-aligned with the dataset's structure, as citation graphs enable graph-based reasoning over case-to-case and case-to-statute relationships, while vector embeddings support semantic search. Community discussion likely centers on opportunities for training or fine-tuning domain-specific legal language models and building retrieval-augmented generation (RAG) pipelines tailored to the Indian legal system.
Related
- Exploring Structural Complexity in Normative RAG with Graph-based approaches: A case study on the ETSI Standards
- NOMAD: Generating Embeddings for Massive Distributed Graphs
- ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
- Domain-Specific Data Generation Framework for RAG Adaptation
- [[d-large-scale-ocr-d|[D] Large scale OCR [D]]]
Source: r/MachineLearning | 2026-04-14