Networking & Cloud · Flagship

Self-Healing Microservices Platform with Chaos Engineering

A microservices platform instrumented for automatic failure detection and recovery, validated with chaos-engineering experiments.

KubernetesIstioGoPrometheusChaos Mesh

In a system of many small services, something is always failing: a container crashes, a network link slows down or a dependency stops answering. What matters is whether users notice. It is not enough to hope that recovery works, it must be proven. This project builds a microservices platform that detects failures and recovers automatically, and validates it with chaos engineering.

The services, written in Go, run on Kubernetes with an Istio service mesh that handles traffic management, retries and timeouts. Circuit breakers stop calls to a failing service so that the failure does not spread, and automated failover shifts traffic to healthy instances while Kubernetes restarts the broken ones. Prometheus collects metrics that define the service levels the system must meet, and alerts fire when they are threatened. A chaos-testing suite built on Chaos Mesh then injects real failures such as killing pods, adding network delay and exhausting resources, and checks that the platform keeps its promises. Each experiment states a hypothesis, measures the effect on users and produces a report, and problems found are fixed and re-tested. The project also discusses running experiments safely.

You will learn service meshes, resilience patterns, observability and the practice of chaos engineering. The Project Reference Guide documents the design and the experiments, and the Reference Implementation includes the services, mesh configuration and chaos experiments.