Building a Perturbation-Aware LLM for Multimodal In Silico Perturbation Modelling
Perturbation biology datasets across CRISPR screens, MAVE, and scPerturb-seq remain siloed in incompatible formats, making cross-modal reasoning about genetic perturbations nearly impossible at scale. This project builds a perturbation-aware LLM by fine-tuning BioMedLM on a curated multimodal training corpus derived from the EMBL-EBI Perturbation Catalogue, enabling natural language queries such as "what happens if gene X is knocked out in cell type Y?" Deliverables include a multimodal training corpus, a fine-tuned LLM prototype, a reproducible evaluation pipeline with gene-level splits distinguishing genuine generalisation from memorisation, and full open-source documentation for reuse by the Perturbation Catalogue team.
Project details
Technologies
Not listed in the archive