A Self-Healing Distributed Cloud Control Plane for Autonomous Fault Detection and Recovery

Authors

  • Aditya Sudhakar Naikwade Sandip Institute of Technology and Research Centre, Nashik, Maharashtra, India Author
  • Naresh C. Thoutam Sandip Institute of Technology and Research Centre, Nashik, Maharashtra, India Author
  • A. Ankita Karale Sandip Institute of Technology and Research Centre, Nashik, Maharashtra, India Author
  • Divyanshu Sachindra Padvi Sandip Institute of Technology and Research Centre, Nashik, Maharashtra, India Author
  • Priyanka Rajendra Mahire Sandip Institute of Technology and Research Centre, Nashik, Maharashtra, India Author
  • Vivek Ravindra Patil Sandip Institute of Technology and Research Centre, Nashik, Maharashtra, India Author

Keywords:

Self-healing systems, Distributed cloud, Kubernetes, FastAPI, Isolation Forest, AI-based anomaly detection, Automatic pod recovery, Microservices, Autonomous control plane

Abstract

Cloud-native services increasingly run on distributed container platforms, where even short-lived failures can disrupt user-facing functionality and backend workflows. In Kubernetes-based microservice deployments, issues such as pod crashes, performance degradation, network instability, and sudden workload surges must be detected and mitigated automatically to maintain service continuity. This paper introduces a Self-Healing Distributed Cloud Platform (SHDCP) built around a layered control plane that supervises Fast API-powered microservices and orchestrates recovery actions at the pod level. The platform employs an AI-driven anomaly detection pipeline in which Isolation Forest and complementary time-series models analyse metrics and request traces to identify abnormal behavior before it becomes critical. Detected incidents are translated into recovery policies that trigger automatic pod restart, rescheduling, or horizontal scaling through Kubernetes, while a persistent knowledge base records decisions and outcomes for future optimization. In addition to describing the architecture and its core components, the work explains the AI integration for online anomaly detection, discusses practical design challenges, and presents an evaluation plan based on controlled failure injection scenarios including overload, node loss, and network partition. The proposed framework aims to deliver high availability, autonomous fault handling, and operational resilience for modern distributed cloud applications.

Downloads

Download data is not yet available.

Downloads

Published

2026-10-11

How to Cite

A Self-Healing Distributed Cloud Control Plane for Autonomous Fault Detection and Recovery. (2026). Journal of Interdisciplinary Science & Technology, 1(3), 42-50. https://onlinejist.com/index.php/jist/article/view/45