You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
OpsAgent is a modern monitoring engine that uses AI to analyze cluster health, detect failures, and automatically heal pods. It bridges the gap between raw telemetry and actionable intelligence.
5
+
OpsAgent is an intelligent, autonomous backend engine that watches your Kubernetes clusters, analyzes pod failures using Large Language Models, and executes self-healing actions.
6
+
7
+
> **Note:** The backend engine (API, AI capabilities, Monitoring Worker, and Slack Alerts) is **fully implemented and operational**. The frontend dashboard is currently a work in progress.
6
8
7
9
---
8
10
9
-
## ✨ Features
11
+
## ✨ Fully Implemented Features
10
12
11
-
-**🤖 AI Diagnosis**: Analyzes pod failures using LLMs (Groq/Llama) to provide human-readable justifications for issues.
12
-
-**⚡ Self-Healing**: Automated pod restarts and healing actions based on AI-driven decisions.
13
-
-**📊 Real-time Dashboard**: A premium, atmospheric UI for monitoring cluster status at a glance.
14
-
-**📈 Prometheus Integration**: Native `/metrics` endpoint for long-term trend analysis and Grafana dashboards.
15
-
-**💬 Slack Notifications**: Instant, detailed alerts sent to your team for every critical event and healing action.
16
-
-**🎡 Helm Ready**: Packaged for production-grade deployment on any Kubernetes cluster.
17
-
-**🏗️ CI/CD Integrated**: GitHub Actions pipelines for automated testing and container image builds.
13
+
-**✅ Native Kubernetes Integration**: Connects seamlessly using your local `~/.kube/config` or in-cluster ServiceAccounts to monitor pod health, deployments, and restarts across all namespaces. (Includes a Mock Mode for testing without a cluster).
14
+
-**✅ AI-Driven RCA (Root Cause Analysis)**: Integrates with Groq (`llama-3.3-70b-versatile`) to provide plain-english, human-readable explanations of *why* pods are failing or restarting.
15
+
-**✅ Autonomous Self-Healing**: A background worker loop continuously monitors the cluster. When pods fail or enter crash loops, OpsAgent can automatically restart them based on AI justifications (configurable via API).
16
+
-**✅ Rich Slack Notifications**: Sends highly detailed, color-coded API-driven Slack alerts complete with pod metrics and AI analysis.
17
+
-**✅ Prometheus Metrics**: Exposes a standard `/metrics` endpoint for seamless integration into your existing Grafana & Prometheus monitoring stacks.
18
18
19
19
---
20
20
21
21
## 🛠️ Architecture
22
22
23
23
OpsAgent consists of three main components:
24
-
1.**The API**: A FastAPI backend that serves as the brain, managing settings and cluster interaction.
25
-
2.**The Worker**: A background monitoring loop that watches the cluster and triggers the AI analysis.
26
-
3.**The Frontend**: A sleek, React-powered dashboard for real-time visualization.
24
+
1.**The API (`main.py`)**: A lightning-fast FastAPI backend that serves as the control plane, exposing metrics, health checks, and configuration endpoints.
25
+
2.**The Worker (`worker.py`)**: An asynchronous background loop that actively polls Kubernetes, triggers LLM analysis on failures, issues Slack alerts, and executes healing actions.
26
+
3.**The AI Engine (`services/ai.py`)**: Bridges the raw K8s status data with prompt-engineered calls to Groq's low-latency inference endpoints.
27
+
4.**The Frontend**: *Currently undergoing a UI revamp.*
27
28
28
29
---
29
30
@@ -36,18 +37,24 @@ OpsAgent consists of three main components:
Contributions are welcome! Please feel free to submit a Pull Request.
75
+
Contributions are welcome! Whether it's adding a new AI provider, expanding the Kubernetes metrics we collect, or building out the UI, please feel free to submit a Pull Request.
0 commit comments