A Chaos Engineering Platform for Kubernetes.
-
Updated
Oct 3, 2026 - Go
Site reliability engineering (SRE) is a set of principles and practices that incorporates aspects of software engineering and applies them to infrastructure and operations problems. The main goals are to create scalable and highly reliable software systems. Site reliability engineering is closely related to DevOps, a set of practices that combine software development and IT operations, and SRE has also been described as a specific implementation of DevOps.
A Chaos Engineering Platform for Kubernetes.
Litmus helps SREs and developers practice chaos engineering in a Cloud-native way. Chaos experiments are published at the ChaosHub (https://hub.litmuschaos.io). Community notes is at https://hackmd.io/a4Zu_sH4TZGeih-xCimi3Q
Chaos testing, network emulation, and stress testing tool for containers
A chaos engineering platform for supporting the complete fault drill lifecycle.
The Skinny Distributed Lock Service
A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
The self-healing engine for modern infrastructure. Detects failures in milliseconds, heals autonomously with agentic AI, proves every action with a post-quantum signed audit trail. 79 Go packages, 12-view operator dashboard, single binary. Apache 2.0.
Terraform provider for Rootly - manage incident management, on-call schedules, workflows, and alerts as code
Stop burning your error budget. SRE config as code — SLOs, error budgets, runbooks, on-call, and dashboards in one sre.yaml file. Auto-remediates before alerts fire.
A typed, observable, policy-driven configuration system for Go services.
Endpoint monitoring and DNS failover agent written in Go
[INACTIVE] Terraform provider for Arachnys' Cabot. Create, manage, and manipulate status checks, and alerts for services.
Various static analysis and testing tools for managing PromQL compatible monitoring stacks
Control health checks and toggle upstream node status in load balancers with ease.
Automated and simplified SLO creation for product teams and developers. Full alerting support, OpenSLO YAML format.
Restores your Postgres backups from S3 into throwaway containers to prove they actually work — no data ever leaves your network.
srepulse CLI + TUI — terminal client for the autonomous Kubernetes SRE agent. kubectl plugin, distributed via Krew.
DevOps E / SRE 업무를 하면서 전문성을 갖추기 위하여 공부한 자료를 업로드하는 공간입니다. 개인적인 공부이지만 참고할 부분이 될 수 있었으면 좋겠습니다.
A tool for tracking troubleshooting notes for SREs