Extending AIDRIN: Zarr & ROOT Readers, HDF5 Hardening, and Pluggable Data Ingestion
This project extends AIDRIN’s data ingestion layer beyond CSV and basic tabular formats to better match real scientific workflows. I will (1) audit and harden the existing HDF5 path, (2) add new readers for Zarr and ROOT (via uproot), and (3) introduce a simple custom ingestion interface that normalizes user‑supplied data sources into a pandas DataFrame and metadata dictionary before AIDRIN’s metrics run. The work includes dataset/TTree selection UX, chunk‑aware reading for large N‑dimensional arrays, metadata extraction for FAIR-related metrics, and a full set of tests and documentation. By the end of the project, researchers will be able to use AIDRIN directly on common scientific formats and plug in their own loaders without modifying the core codebase.
Project details
Technologies
Not listed in the archive