Back to Genome Assembly and Annotation
GSoC 2026

Expand genome metadata in Ensembl with AI tools

The Ensembl Plants and Metazoa platforms face a significant metadata gap where critical biological context, such as ploidy, strain, and sex, is often documented in peer-reviewed literature but missing from formal INSDC sequence archives. This lack of structured metadata, particularly ploidy, creates technical bottlenecks in downstream comparative genomics and requires labor-intensive manual curation. To address this, I propose developing a standalone, Python-based "Retrieve-Extract-Verify" module that leverages Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) to automate metadata extraction from scientific papers. This tool will utilize literature APIs like Europe PMC, employing section-aware parsing and confidence scoring to ensure high data integrity with minimal human intervention. My previous experience in building Transformer-based bioprocess optimization models and genome-level embeddings directly informs my approach to handling complex biological datasets. The final outcome will be a ready-to-deploy Nextflow module designed for seamless integration into Ensembl’s production pipeline and provide the global research community with richer, more reliable genomic metadata.

Project details

Contributor

Soomin Lee

Mentors

Not available

Technologies

Not listed in the archive