A Self-Healing Distributed Cloud Control Plane for Autonomous Fault Detection and Recovery
Keywords:
Self-healing systems, Distributed cloud, Kubernetes, FastAPI, Isolation Forest, AI-based anomaly detection, Automatic pod recovery, Microservices, Autonomous control planeAbstract
Cloud-native services increasingly run on distributed container platforms, where even short-lived failures can disrupt user-facing functionality and backend workflows. In Kubernetes-based microservice deployments, issues such as pod crashes, performance degradation, network instability, and sudden workload surges must be detected and mitigated automatically to maintain service continuity. This paper introduces a Self-Healing Distributed Cloud Platform (SHDCP) built around a layered control plane that supervises Fast API-powered microservices and orchestrates recovery actions at the pod level. The platform employs an AI-driven anomaly detection pipeline in which Isolation Forest and complementary time-series models analyse metrics and request traces to identify abnormal behavior before it becomes critical. Detected incidents are translated into recovery policies that trigger automatic pod restart, rescheduling, or horizontal scaling through Kubernetes, while a persistent knowledge base records decisions and outcomes for future optimization. In addition to describing the architecture and its core components, the work explains the AI integration for online anomaly detection, discusses practical design challenges, and presents an evaluation plan based on controlled failure injection scenarios including overload, node loss, and network partition. The proposed framework aims to deliver high availability, autonomous fault handling, and operational resilience for modern distributed cloud applications.