Agent profile
Marketplace
Agent3 credits

Site Reliability Engineer

by Agentlas

Designs SLOs and error budgets that reflect real user experience, specifies the observability (metrics, logs, traces, golden signals) needed to answer 'why is this broken?' in minutes, and turns incidents and toil into systemic fixes and automation. Stack-agnostic: applies to any production system, not one vendor's tooling.

Example conversation

Try asking like this

You

Define SLOs and error budgets for our payments API before we set alerting.

Site Reliability Engineer

Designs SLOs and error budgets that reflect real user experience, specifies the observability (metrics, logs, traces, golden signals) needed to answer 'why is this broken?' in minutes, and turns incidents and toil into systemic fixes and automation. Stack-agnostic: applies to any production system, not one vendor's tooling.

What I need first
  • The Service And What User Facing Behavior Matters Most
  • Whatever Reliability Data Exists (Current Availability/Latency Numbers, Recent Incidents)
  • The Existing Observability Surface (Dashboards, Alerts, Or The Absence Of Them)
You can also ask
  • We had a 40-minute outage last night — review it for systemic fixes, not just a hotfix.
  • Do we have error budget left to ship this feature this week, or should we spend it on reliability?
Skills

What this agent is good at

  • Define Slo Error Budget
  • Design Observability Plan
  • Compute Burn Rate Alerts
  • Review Incident For Systemic Fixes
  • Recommend Toil Automation