Kubeflow logo

Kubeflow

Our goal is to streamline the process of scaling and deploying machine learning (ML) models to production by leveraging Kubernetes' strengths: <br /><br /> - Seamless deployments: Effortless, repeatable, and portable deployments across diverse infrastructure, allowing you to experiment on local machines and then seamlessly transition to on-premises clusters or the cloud. <br /><br /> - Microservice management: Efficient deployment and management of loosely-coupled microservices for building modular and scalable ML pipelines. <br /><br /> - Dynamic scaling: Automated scaling based on demand, ensuring optimal resource utilization. <br /><br /> User-centric design: Recognizing the diverse tool preferences within the ML community, we prioritize user customization (within reasonable boundaries). Our system takes care of the repetitive tasks, freeing users to focus on the core ML challenges. <br /><br /> Starting focused, expanding rapidly: While we initially focused on specific technologies, we actively collaborate with various projects to integrate additional tools and broaden our reach. <br /><br /> Our vision: A future where simple manifests empower you to utilize a user-friendly ML stack anywhere Kubernetes runs, with self-configuration capabilities based on the deployment environment.

GSoC Participation History

Technologies

Topics

Past Projects

OptimizationJob CRD for Hyperparameter Optimization in Kubeflow Katib

This project addresses limitations in Kubeflow Katib current Experiment CRD, which uses a generic and loosely typed interface for hyperparameter...

Project 5: Helm Charts for Kubeflow Pipelines and Katib — Danish Siddiqui

KFP and Katib users deploying via Helm currently have no upstream-supported path — charts exist but drift from Kustomize baselines silently, with no...

Project 6: MCP Server for Kubeflow SDK

The Kubeflow SDK gives AI practitioners a clean Python interface to submit, monitor, and manage distributed training jobs on Kubernetes via...

Agentic RAG on Kubeflow — Multi-Index Retrieval with Kagent & MCP

Kubeflow's documentation, GitHub issues, and platform code are spread across dozens of repositories with no unified search. This project evolves...

Kubeflow SDK/SparkClient - Batch Jobs, Observability & Production Readiness

The current Kubeflow SparkClient (KEP-107) provides a solid foundation for running interactive Spark workloads on Kubernetes, but it is still missing...

Platform Scalability and Security

Kubeflow's adoption at enterprise scale exposes critical bottlenecks in controller efficiency, security posture, and operational overhead. This...

End-to-End ARM64 Support & Validation on Kubeflow

ARM64 is becoming increasingly common with the rise of Apple Silicon and cloud instances like AWS Graviton, but Kubeflow still doesn’t run...

Project 11 - Composable Kale Notebooks with Visual Pipeline Editor

Kale compiles a single Jupyter notebook into a Kubeflow Pipeline by parsing cell tags, detecting data dependencies with PyFlakes, and generating KFP...

Dynamic LLM Trainer Framework for Kubeflow — TRL Backend with Pluggable Multi-Framework Support

Kubeflow Trainer V2 currently supports only TorchTune as its LLM fine-tuning backend. TorchTune stopped adding new features in July 2025, leaving...

Frequently Asked Questions

Kubeflow | GSoC Org Profile & Stats - Learn about Kubeflow's involvement in Google Summer of Code (GSoC), their technologies, detailed reports.

|Currently Active|

Contributor Readiness

Participation

Projects

Top Programming Languages

python dominates with primary adoption

Project Difficulty Distribution

Beginner
0
Intermediate
0
Advanced
0

No difficulty data available