jobs in Confidential Jobs

Kerja Sepenuh Masa, Site Reliability Engineer di Confidential Jobs Federal Territory - Maukerja

Site Reliability Engineer

Confidential Jobs

KL City, Federal Territory

Kongsi
Simpan

Lokasi Kerja

  • Jalan Sultan Mizan Zainal Abidin, Kompleks Kerajaan Kuala Lumpur Federal Territory Malaysia

Penerangan Kerja

Tanggungjawab

About the Role


As a Site Reliability Engineer, you will play a key role in owning the reliability, availability, and performance of production systems at scale. You will work closely with development, security, and compliance teams to design and enforce SLOs/error budgets, drive incident response, and build the automation that keeps the infrastructure resilient and audit-ready. This position is central to leading a modern SRE practice.


Key Responsibilities

  • Own reliability for production systems at scale — defining and tracking SLIs/SLOs, error budgets, and reliability roadmaps
  • Administer and harden Linux/Unix systems, and build automation and tooling in Python, Go, or similar languages to reduce manual toil
  • Manage infrastructure as code using Terraform, Ansible, and operate containerized workloads with Kubernetes and Docker
  • Build and maintain CI/CD pipelines (GitHub actions, Argo CD, Kargo) to enable safe, frequent, self-service deployments
  • Design and operate infrastructure across AWS, GCP, Azure or AliCloud (multi-cloud is a plus), with solid grounding in networking fundamentals — DNS, load balancing, TCP/IP, CDNs
  • Build and maintain observability — metrics, logs, and traces via Prometheus, Grafana, Datadog, or OpenTelemetry — to catch issues before they become incidents
  • Drive incident response and on-call rotations, including leading postmortems and blameless root-cause analysis, with clear, calm communication during active incidents
  • Run resilience testing to validate failover, redundancy, and recovery assumptions
  • Own capacity planning, performance tuning, and database reliability (replication, backups, failover) for critical systems
  • Analyze incidents to identify systemic issues, and lead initiatives that measurably reduce incident frequency and MTTR
  • Design and implement change management processes that satisfy audit and compliance requirements without slowing down engineering velocity
  • Apply and enforce security baseline standards across infrastructure and deployment pipelines
  • Lead or significantly contribute to the transformation from legacy operational practices to modern SRE workflows (automation-first, self-service, reduced toil)
  • Act as a technical lead on cross-functional reliability initiatives
  • Partner with development, security, and compliance stakeholders to embed reliability and audit-readiness into the SDLC
  • Champion SRE best practices — blameless postmortems, toil reduction, capacity planning, and progressive rollouts


Required Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Systems Engineering, or a related field
  • 3-5+ years of hands-on experience in SRE, DevOps, or infrastructure engineering roles
  • Demonstrated, hands-on track record designing and owning reliability for production systems at scale
  • Solid grounding in SRE principles (SLOs, error budgets, toil, blameless postmortems)
  • Experience building or operating change management processes aligned with audit requirements
  • Familiarity applying security baseline standards such as CIS benchmarks and NIST


Peringatan Penting

Jangan pernah kongsikan maklumat bank atau kad kredit anda semasa memohon pekerjaan. Elakkan membuat sebarang pembayaran atau mengisi survey yang tidak berkaitan. Jika ada yang mencurigakan, sila laporkan iklan pekerjaan ini segera.

Lebih Lanjut