About the Role
As a Site Reliability Engineer, you will play a key role in owning the reliability, availability, and performance of production systems at scale. You will work closely with development, security, and compliance teams to design and enforce SLOs/error budgets, drive incident response, and build the automation that keeps the infrastructure resilient and audit-ready. This position is central to leading a modern SRE practice.
Key Responsibilities
- Own reliability for production systems at scale — defining and tracking SLIs/SLOs, error budgets, and reliability roadmaps
- Administer and harden Linux/Unix systems, and build automation and tooling in Python, Go, or similar languages to reduce manual toil
- Manage infrastructure as code using Terraform, Ansible, and operate containerized workloads with Kubernetes and Docker
- Build and maintain CI/CD pipelines (GitHub actions, Argo CD, Kargo) to enable safe, frequent, self-service deployments
- Design and operate infrastructure across AWS, GCP, Azure or AliCloud (multi-cloud is a plus), with solid grounding in networking fundamentals — DNS, load balancing, TCP/IP, CDNs
- Build and maintain observability — metrics, logs, and traces via Prometheus, Grafana, Datadog, or OpenTelemetry — to catch issues before they become incidents
- Drive incident response and on-call rotations, including leading postmortems and blameless root-cause analysis, with clear, calm communication during active incidents
- Run resilience testing to validate failover, redundancy, and recovery assumptions
- Own capacity planning, performance tuning, and database reliability (replication, backups, failover) for critical systems
- Analyze incidents to identify systemic issues, and lead initiatives that measurably reduce incident frequency and MTTR
- Design and implement change management processes that satisfy audit and compliance requirements without slowing down engineering velocity
- Apply and enforce security baseline standards across infrastructure and deployment pipelines
- Lead or significantly contribute to the transformation from legacy operational practices to modern SRE workflows (automation-first, self-service, reduced toil)
- Act as a technical lead on cross-functional reliability initiatives
- Partner with development, security, and compliance stakeholders to embed reliability and audit-readiness into the SDLC
- Champion SRE best practices — blameless postmortems, toil reduction, capacity planning, and progressive rollouts
Required Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Systems Engineering, or a related field
- 3-5+ years of hands-on experience in SRE, DevOps, or infrastructure engineering roles
- Demonstrated, hands-on track record designing and owning reliability for production systems at scale
- Solid grounding in SRE principles (SLOs, error budgets, toil, blameless postmortems)
- Experience building or operating change management processes aligned with audit requirements
- Familiarity applying security baseline standards such as CIS benchmarks and NIST