Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
-
Updated
Sep 8, 2026 - Go
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Hopsworks - Data-Intensive AI platform with a Feature Store
🪐 1-click Kubeflow using ArgoCD
Carbon Limiting Auto Tuning for Kubernetes
AWS SageMaker, SeldonCore, KServe, Kubeflow & MLflow, VectorDB
The Machine Learning Zoomcamp teaches foundational and advanced ML concepts using tools like NumPy, Pandas, Scikit-Learn, TensorFlow, XGBoost, Flask, Docker, AWS, Kubernetes, and KServe. It covers regression, classification, evaluation metrics, neural networks, deployment strategies, and end-to-end projects to bridge theory and practice.
My repo for the Machine Learning Engineering bootcamp 2022 by DataTalks.Club
End-to-end MLOps architecture built for polyp segmentation — featuring distributed Ray training, MLflow experiment tracking, and automated CI/CD with Kubeflow Pipelines and KServe (Triton) deployment on Google Kubernetes Engine.
Self-hosted Qwen3.8-27B (FP8) inference with vLLM, KServe and Envoy AI Gateway on RTX 6000 PRo or 2× RTX 4080 Super
Deploying machine learning model using 10+ different deployment tools
Seven hands-on labs for running AI workloads on Kubernetes: GPU scheduling, distributed training, model serving, and agents. Runs on a local kind cluster, no GPU required. Companion to a KubeCon EU 2026 talk.
A science inference cluster — every model, from protein folding to LLMs, behind one OpenAI/Anthropic-compatible endpoint and a single key. Built to run beside HPC so agentic AI researchers can drive simulations and model inference together. Self-deploying on RKE2 with KServe/Knative + HAMi GPU.
Client/Server system to perform distributed inference on high load systems.
Hands-on labs on deploying machine learning models with tf-serving and KServe
prokube.ai examples - Notebooks, Pipelines, Experiment Tracking, Model tuning, and more
A demo to accompany our blogpost "Scalable Machine Learning with Kafka Streams and KServe"
Observability for vLLM on KServe — dashboards, PromQL, SLOs, and a closed-loop load generator
Production-pattern Red Hat OpenShift AI 3.4.0 platform with bare-metal ESXi, GPU passthrough, KServe RawDeployment, DeepSeek R1 inference at 12–17 tok/s
To associate your repository with the kserve topic, visit your repo's landing page and select "manage topics."