Skip to content

Commit 9530e99

Browse files
committed
Update README with backend features
1 parent e93f11d commit 9530e99

1 file changed

Lines changed: 32 additions & 20 deletions

File tree

‎README.md‎

Lines changed: 32 additions & 20 deletions
Original file line numberDiff line numberDiff line change
@@ -2,28 +2,29 @@
22

33
**AI-Powered Kubernetes Monitoring & Self-Healing**
44

5-
OpsAgent is a modern monitoring engine that uses AI to analyze cluster health, detect failures, and automatically heal pods. It bridges the gap between raw telemetry and actionable intelligence.
5+
OpsAgent is an intelligent, autonomous backend engine that watches your Kubernetes clusters, analyzes pod failures using Large Language Models, and executes self-healing actions.
6+
7+
> **Note:** The backend engine (API, AI capabilities, Monitoring Worker, and Slack Alerts) is **fully implemented and operational**. The frontend dashboard is currently a work in progress.
68
79
---
810

9-
## ✨ Features
11+
## ✨ Fully Implemented Features
1012

11-
- **🤖 AI Diagnosis**: Analyzes pod failures using LLMs (Groq/Llama) to provide human-readable justifications for issues.
12-
- **⚡ Self-Healing**: Automated pod restarts and healing actions based on AI-driven decisions.
13-
- **📊 Real-time Dashboard**: A premium, atmospheric UI for monitoring cluster status at a glance.
14-
- **📈 Prometheus Integration**: Native `/metrics` endpoint for long-term trend analysis and Grafana dashboards.
15-
- **💬 Slack Notifications**: Instant, detailed alerts sent to your team for every critical event and healing action.
16-
- **🎡 Helm Ready**: Packaged for production-grade deployment on any Kubernetes cluster.
17-
- **🏗️ CI/CD Integrated**: GitHub Actions pipelines for automated testing and container image builds.
13+
- **✅ Native Kubernetes Integration**: Connects seamlessly using your local `~/.kube/config` or in-cluster ServiceAccounts to monitor pod health, deployments, and restarts across all namespaces. (Includes a Mock Mode for testing without a cluster).
14+
- **✅ AI-Driven RCA (Root Cause Analysis)**: Integrates with Groq (`llama-3.3-70b-versatile`) to provide plain-english, human-readable explanations of *why* pods are failing or restarting.
15+
- **✅ Autonomous Self-Healing**: A background worker loop continuously monitors the cluster. When pods fail or enter crash loops, OpsAgent can automatically restart them based on AI justifications (configurable via API).
16+
- **✅ Rich Slack Notifications**: Sends highly detailed, color-coded API-driven Slack alerts complete with pod metrics and AI analysis.
17+
- **✅ Prometheus Metrics**: Exposes a standard `/metrics` endpoint for seamless integration into your existing Grafana & Prometheus monitoring stacks.
1818

1919
---
2020

2121
## 🛠️ Architecture
2222

2323
OpsAgent consists of three main components:
24-
1. **The API**: A FastAPI backend that serves as the brain, managing settings and cluster interaction.
25-
2. **The Worker**: A background monitoring loop that watches the cluster and triggers the AI analysis.
26-
3. **The Frontend**: A sleek, React-powered dashboard for real-time visualization.
24+
1. **The API (`main.py`)**: A lightning-fast FastAPI backend that serves as the control plane, exposing metrics, health checks, and configuration endpoints.
25+
2. **The Worker (`worker.py`)**: An asynchronous background loop that actively polls Kubernetes, triggers LLM analysis on failures, issues Slack alerts, and executes healing actions.
26+
3. **The AI Engine (`services/ai.py`)**: Bridges the raw K8s status data with prompt-engineered calls to Groq's low-latency inference endpoints.
27+
4. **The Frontend**: *Currently undergoing a UI revamp.*
2728

2829
---
2930

@@ -36,18 +37,24 @@ OpsAgent consists of three main components:
3637

3738
### 2. Local Setup
3839
```bash
39-
# Install dependencies
40+
# Clone the repository
41+
git clone https://github.com/yourusername/opsagent.git
42+
cd opsagent
43+
44+
# Install Python dependencies
4045
pip install -r requirements.txt
4146

42-
# Set environment variables
43-
export GROQ_API_KEY="your_key"
44-
export SLACK_WEBHOOK_URL="your_webhook"
47+
# Set required environment variables
48+
export GROQ_API_KEY="your_groq_api_key"
49+
export SLACK_WEBHOOK_URL="your_slack_webhook"
4550

46-
# Start the agent
51+
# Start the API & Worker
4752
python main.py
4853
```
54+
*Note: OpsAgent runs on port `8000` by default. You can access the API documentation at `http://localhost:8000/docs`.*
4955

5056
### 3. Deploy to Kubernetes (Helm)
57+
OpsAgent comes packaged and ready for production:
5158
```bash
5259
helm install opsagent ./charts/opsagent \
5360
--set env.groqApiKey="your_key" \
@@ -56,11 +63,16 @@ helm install opsagent ./charts/opsagent \
5663

5764
---
5865

59-
## 📊 Monitoring
60-
Metrics are available at `:8000/metrics`. Point your Prometheus instance here to start collecting data.
66+
## ⚙️ Configuration & API
67+
68+
The API exposes several endpoints to control the agent:
69+
- `GET /health`: Detailed system health (Python version, K8s connectivity, AI status).
70+
- `GET /settings`: View current worker settings.
71+
- `POST /settings/toggle-heal`: Enable or disable the autonomous Self-Healing loop.
72+
- `GET /metrics`: Prometheus-compatible telemetry endpoint.
6173

6274
## 🤝 Contributing
63-
Contributions are welcome! Please feel free to submit a Pull Request.
75+
Contributions are welcome! Whether it's adding a new AI provider, expanding the Kubernetes metrics we collect, or building out the UI, please feel free to submit a Pull Request.
6476

6577
---
6678

0 commit comments

Comments
 (0)