Back to CERN-HSF
GSoC 2026

Automated Software Performance Monitoring for the ATLAS experiment

At present, software regression testing at ATLAS is predominantly performed manually via inspection of metrics on the ATLAS project management board. This process is both time-consuming and error-prone, and it does not provide any easy way to identify potential root causes of a given decrease in software performance. To ensure reliability and robustness, it is necessary to design mechanisms which can autonomously detect anomalies and failures at scale as they occur. I will resolve these shortfalls by modifying the SPOT orchestration scripts to incorporate an automated anomaly detection pipeline which can issue alerts upon detecting anomalies. Statistical and machine learning approaches will be employed to detect anomalies, and tested, validated and verified on existing data. Additionally, I will also upgrade the performance monitoring board to provide improved at-a-glance visibility of recent anomalies to ease diagnosis and resolution.

Project details

Contributor

Douglas Lindsay

Mentors

Not available

Technologies

Not listed in the archive