Foundation models for End-to-End event reconstruction
This project aims to develop a multi-modal foundation model training pipeline for end-to-end particle reconstruction in the CMS experiment. Building on last year's LorentzParT hybrid architecture (L-GATr + ParT) and its track-level masked autoencoder pretraining, I will design a unified tokenization scheme encoding particle kinematics and interaction features as distinct modality streams, and pretrain the shared encoder using a novel Cross-Modal JEPA objective that predicts masked token embeddings across modalities in latent space — combining, for the first time, JEPA-style self-supervision with Lorentz-equivariant representations. The pretrained encoder will be fine-tuned and benchmarked on multiple downstream tasks including event/jet classification and mass regression, with conditional generation as a stretch goal. Deliverables include the complete CM-JEPA pretraining pipeline, a multi-modal tokenizer, multi-task fine-tuning scripts, systematic ablation studies across pretraining methods and masking strategies, trained model checkpoints, and an open-source repository with documentation and a blog post.
Project details
Technologies
Not listed in the archive