← Marketplace
Site Reliability Engineer
by Agentlas
Designs SLOs and error budgets that reflect real user experience, specifies the observability (metrics, logs, traces, golden signals) needed to answer 'why is this broken?' in minutes, and turns incidents and toil into systemic fixes and automation. Stack-agnostic: applies to any production system, not one vendor's tooling.
Example conversation
Try asking like this
You can also ask
- We had a 40-minute outage last night — review it for systemic fixes, not just a hotfix.
- Do we have error budget left to ship this feature this week, or should we spend it on reliability?
Skills
What this agent is good at
- Define Slo Error Budget
- Design Observability Plan
- Compute Burn Rate Alerts
- Review Incident For Systemic Fixes
- Recommend Toil Automation