AIOps & ObservabilityStartupCNCF Sandbox
K8sGPT
CNCF Sandbox project that scans Kubernetes clusters and uses LLMs to explain failures in plain English and suggest fixes — CLI or continuous in-cluster operator
Mkt Cap / ValOpen Source
May 2026: v0.4.33 ships CNCF incubation prep
De facto open-source standard for AI-assisted Kubernetes diagnostics, under CNCF governance.
SWOT Analysis
Strengths
- CNCF Sandbox project with active, vendor-neutral governance
- 8k GitHub stars and 1k forks with steady release cadence into 2026
- Supports many AI backends including OpenAI, Bedrock, Azure, and local models
- Operator mode enables continuous in-cluster scanning, not just one-off CLI runs
- Free and open source, easy bottom-up adoption by platform teams
Opportunities
- CNCF incubation path underway per May 2026 governance prep
- Local model support for privacy-sensitive enterprises
- Integration into platform engineering golden paths
- Agent and MCP ecosystem integrations
Weaknesses
- No commercial support or SLA behind the project
- Diagnostic quality depends on the chosen LLM backend
- Scope limited to Kubernetes diagnostics, not full observability
- Still Sandbox maturity; incubation not yet achieved
Threats
- AI assistants built into kubectl and hyperscaler consoles
- Observability vendors bundling K8s AI troubleshooting
- General-purpose coding agents handling cluster debugging
- Maintainer and contributor sustainability risk
User Sentiment
Synthesized from G2, Gartner Peer Insights, and analyst review data.
What users love
- Plain-English explanations of cryptic cluster errors
- Fast triage that saves SRE investigation time
- Simple install as CLI or Kubernetes operator
- Built-in anonymization of sensitive cluster data
Common complaints
- LLM token costs add up on large clusters
- Remediation advice can be generic for complex failures
- Configuring AI backends and auth can be fiddly
Customer Profile
Who buys this
Typical segments
Platform engineering teamsKubernetes-heavy enterprises
Typical buyer
Platform engineering lead or SRE manager adopting bottom-up via OSS
Top use cases
- 1Triaging failing pods and misconfigurations
- 2Continuous in-cluster scanning via operator
- 3Helping junior engineers debug Kubernetes
Future Focus Areas
1
CNCF incubation and graduated maturity
2
Deeper automated remediation workflows
3
Expanded local and open-model backends
4
Tighter observability stack integrations