Back to National Resource for Network Biology (NRNB)
GSoC 2026

ML Models that use SBOL for DNA annotation by Shreeya

Developing a scalable and reproducible research pipeline within SeqTrainer for bacterial DNA sequence modelling and annotation using SBOL-based datasets. The project addresses a key challenge in synthetic biology: the lack of accessible, machine learning–ready datasets and standardized workflows for genomic modelling. The primary focus is on E. coli and Gram-negative bacteria, with applications in promoter classification and sequence annotation. The work involves reproducing and benchmarking existing CNN-based baselines, followed by integrating and evaluating foundation models such as DNABERT2 and Evo 2 within structured experiment pipelines. Experiments explore challenges including overfitting and class imbalance through methods such as weighted loss functions, undersampling, learning-rate tuning, and early stopping. A major component of the project is the development of a modular and reproducible experiment framework within SeqTrainer, including configuration-driven workflows, reproducible dataset splitting, evaluation utilities, metric logging, and improved SBOL-to-dataset conversion pipelines with validation and provenance tracking. The final outcome aims to transform SeqTrainer into a more experiment-friendly framework for genomic ML research, enabling accessible, extensible, and reproducible machine learning workflows for synthetic biology researchers.

Project details

Contributor

Shreeya456

Mentors

Not available

Technologies

Not listed in the archive