Back to DBpedia
GSoC 2026

Building the Amharic DBpedia Language Chapter with Large Language Models (LLMs)

The Amharic DBpedia chapter's past implementation was largely manual, relying on tedious template mapping that has proven impossible to scale given the complexity of the Amharic language and the messy Wikipedia markup. My project's goal is to transition the chapter away from this bottleneck and fully automate the entire Semantic Web pipeline. I'll build an Agentic Orchestration Pipeline (using LangGraph/Mastra) that fixes the problem in three core areas: AI-Powered Mapping: The pipeline will integrate the fine-tuned Afro-XLM-R model to automatically predict and align Amharic properties to the DBpedia Ontology, completely replacing the manual mapping effort. This process is preceded by a Python preprocessor that cleans the raw Wikipedia XML dumps to prevent parser crashes. Human-in-the-Loop Safety (HITL): I will wrap the entire system in a Human-in-the-Loop (HITL) interface, allowing community experts to verify low-confidence predictions to ensure data quality and continuously improve the AI model. Visualization and Accessibility: Finally, I will refactor the current am.dbpedia.org website into a dynamic dashboard that hooks into our deployed SPARQL endpoint, letting users actually query and visualize the live Knowledge Graph. Key deliverables will include the fully functional, automated extraction framework, the publication of the generated Amharic RDF triples to the DBpedia Databus, the deployment of a public-facing SPARQL endpoint (Tentris/Fuseki) for live querying, and the refactoring of the am.dbpedia.org website into a dynamic dashboard for visualizing the new Knowledge Graph.

Project details

Contributor

Natnael Yohanes

Mentors

Not available

Technologies

Not listed in the archive