Hindi Relational Triple Extraction with Fine Tuned Indic Models and Human in the Loop Feedback
This project focuses on improving how relational triples are extracted from Hindi Wikipedia text for the DBpedia Hindi chapter. Currently, most structured data comes from infoboxes, while a large amount of useful information in free text remains unused. Existing approaches either rely on prompts, which are inconsistent, or rule-based systems that cannot handle the variety of Hindi sentence structures. There is also no proper system to use human feedback to improve results over time. To address this, I will fine-tune a small language model specifically for Hindi triple extraction, with a focus on correctly identifying predicates. The system will also include a layer that maps extracted relations to DBpedia ontology properties, ensuring compatibility with the knowledge graph. A simple interface will be built to allow users to review, correct, and improve extracted triples, creating better training data over time. The final outcome will be a working pipeline that extracts structured triples from Hindi text, a feedback system for continuous improvement, and a dataset of validated triples. This will help improve the coverage and quality of the Hindi DBpedia knowledge graph.
Project details
Technologies
Not listed in the archive