← All open roles
Cloud & Platform
Site Reliability Engineer
- Remote — USA
- Full-time
You will keep client systems available and fast — defining SLOs, building observability, and leading incident response and the reviews that follow.
What you will do
- Define SLIs and SLOs with client engineering teams
- Build observability: metrics, logging, tracing and alerting
- Lead incident response and blameless postmortems
- Drive capacity planning and performance work
What we are looking for
- 4+ years in SRE or production operations
- Deep experience with an observability stack (Datadog, Grafana, Prometheus)
- Strong Linux and networking fundamentals
- Calm under pressure during incidents
What we offer
- Competitive salary and performance bonuses
- Health, dental and vision cover
- Remote-first with flexible hours
- Paid certification and training budget
- Paid time off and public holidays
- Referral bonus programme
Apply for this role
Tell us how to reach you. Applications go straight to our talent team at hr@jsp-its.com.