ML Models that use SBOL for DNA annotation by Shreeya
Developing a scalable and reproducible research pipeline within SeqTrainer for bacterial DNA sequence modelling and annotation using SBOL-based datasets. The project addresses a key challenge in synthetic biology: the lack of accessible, machine learning–ready datasets and standardized workflows for genomic modelling. The primary focus is on E. coli and Gram-negative bacteria, with applications in promoter classification and sequence annotation. The work involves reproducing and benchmarking existing CNN-based baselines, followed by integrating and evaluating foundation models such as DNABERT2 and Evo 2 within structured experiment pipelines. Experiments explore challenges including overfitting and class imbalance through methods such as weighted loss functions, undersampling, learning-rate tuning, and early stopping. A major component of the project is the development of a modular and reproducible experiment framework within SeqTrainer, including configuration-driven workflows, reproducible dataset splitting, evaluation utilities, metric logging, and improved SBOL-to-dataset conversion pipelines with validation and provenance tracking. The final outcome aims to transform SeqTrainer into a more experiment-friendly framework for genomic ML research, enabling accessible, extensible, and reproducible machine learning workflows for synthetic biology researchers.
Project details
Technologies
Not listed in the archive