Back to Kubeflow
GSoC 2026

Kubeflow SDK/SparkClient - Batch Jobs, Observability & Production Readiness

The current Kubeflow SparkClient (KEP-107) provides a solid foundation for running interactive Spark workloads on Kubernetes, but it is still missing key capabilities required for real-world usage, such as batch job submission, consistent job lifecycle management, and production-level observability. This project focuses on extending the SparkClient to support complete, production-ready Spark workflows. The primary goal is to implement a robust submit_job() API for batch job execution using the SparkApplication CRD, supporting both script-based and function-based workloads. Alongside this, I will design a consistent set of lifecycle APIs such as job listing, status tracking, log retrieval, and cleanup, ensuring they align with existing Kubeflow SDK patterns. To improve usability in production environments, the project will also introduce observability features, including metrics collection from the Spark REST API, structured event tracking, and basic health monitoring. These additions will help users better understand job execution, debug failures, and monitor resource usage. The implementation will build on the current architecture without breaking compatibility, focusing on reliability, validation, and consistency. The final outcome will be a more complete and production-ready SparkClient that enables users to run, monitor, and manage Spark workloads on Kubernetes in a simple and consistent way.

Project details

Contributor

Sameer_Yadav

Mentors

Not available

Technologies

Not listed in the archive