Back to projects
Final Year Project
February 2025 - June 2025

KubeMind

Intelligent Kubernetes orchestration platform with predictive auto-scaling

Final year project (PFE) — a full-stack platform for monitoring and managing Kubernetes clusters in real time. Integrates ML-based predictive auto-scaling, multi-cluster visibility, automated alerting, and an interactive workload visualization dashboard.

KubernetesReactFastAPIPythonPrometheusGrafanaDockerHelmPostgreSQL

-5 min

Scaling reaction time

< 1s

Dashboard refresh

Multi

Clusters managed

< 3s

Alert latency

Context

Kubernetes has become the de-facto standard for container orchestration, but its operational complexity — especially for multi-cluster environments — remains a significant challenge for engineering teams. Manual scaling decisions lead to either over-provisioning (wasted cost) or under-provisioning (degraded performance).

Problem statement

"How do you build a system that gives operators a single pane of glass across multiple clusters, automates scaling decisions before incidents occur, and reduces mean time to detection (MTTD) for anomalies?"

Architecture

  • 1React + TypeScript frontend with real-time WebSocket updates
  • 2FastAPI backend exposing a REST + WebSocket API
  • 3Prometheus scraping cluster metrics every 15 seconds
  • 4Python ML service (time-series forecasting) predicting load spikes 5 min ahead
  • 5Helm chart for one-command cluster deployment
  • 6PostgreSQL for historical metrics and alert history

Key features

Real-time monitoring

Built a live dashboard showing CPU, memory, and network usage per pod/namespace using WebSocket streaming. Operators see cluster health change in real time without manual refresh.

Predictive auto-scaling

Trained a lightweight LSTM model on historical Prometheus metrics to forecast resource demand 5 minutes ahead. The system triggers Horizontal Pod Autoscaler adjustments proactively, before load spikes hit.

Multi-cluster management

Designed a namespace-isolated architecture allowing a single KubeMind instance to manage multiple clusters via kubeconfig context switching, with per-cluster RBAC.

Alerting engine

Configurable threshold-based and anomaly-based alerts with webhook notifications. Alert history stored in PostgreSQL with full audit trail.

Technical challenges

Designing a ML model light enough to run alongside cluster workloads without adding meaningful overhead — solved by using a compressed LSTM with a 15-min rolling window instead of a full time-series database scan.

WebSocket state synchronization across multiple cluster contexts — solved by a per-cluster event bus with client-side reconciliation.

Making Helm deployment truly one-command for any cluster size — required extensive templating and conditional resource generation.

What I learned

  • Production Kubernetes internals — controllers, informers, admission webhooks
  • Time-series forecasting with LSTM on operational metrics
  • Designing for operator experience, not just functionality

Interested in working together?

I am open to an end-of-year internship opportunity starting February 2027.