Back to Genome Assembly and Annotation
GSoC 2026

Building a Perturbation-Aware LLM for Multimodal In Silico Perturbation Modelling

Perturbation biology datasets across CRISPR screens, MAVE, and scPerturb-seq remain siloed in incompatible formats, making cross-modal reasoning about genetic perturbations nearly impossible at scale. This project builds a perturbation-aware LLM by fine-tuning BioMedLM on a curated multimodal training corpus derived from the EMBL-EBI Perturbation Catalogue, enabling natural language queries such as "what happens if gene X is knocked out in cell type Y?" Deliverables include a multimodal training corpus, a fine-tuned LLM prototype, a reproducible evaluation pipeline with gene-level splits distinguishing genuine generalisation from memorisation, and full open-source documentation for reuse by the Perturbation Catalogue team.

Project details

Contributor

Esra Ersan

Mentors

Not available

Technologies

Not listed in the archive