Back to OpenVINO Toolkit
GSoC 2026

Auto-Labeling "Data Factory" & Edge Training Integration

This project introduces a ‘Cold Start’ Data Factory that turns zero/one-shot prompting into a full data-to-training workflow. The core contribution is a production-grade DatasetWriter pipeline that converts inference outputs into validated, reproducible annotations, supports format-aware serialization, and guarantees reliable flushing/finalization for large export jobs. The generated datasets can be exported in standard formats such as YOLO, COCO, and CVAT, with quality statistics and filtering controls for practical downstream use. Building on that, the project adds direct OTX training integration: once export is finalized, the system can automatically generate training configs, spawn OTX runs, stream logs/status, and track artifacts and run metadata. This shifts geti-instant-learn from a live-inference-only experience to a complete cold-start bootstrapping system where prompt-based “Teacher” labels are transformed into trainable “Student” datasets and model training jobs. The result is a faster and more reliable path from a few prompts to deployable edge models.

Project details

Contributor

Saptadip Saha

Mentors

Not available

Technologies

Not listed in the archive