Site Reliability Engineering Jobs in Kuala Lumpur - August 2026 - Urgent Hiring

Showing 42 jobs results for "site reliability engineering" in Kuala Lumpur
Never miss any updates for Site Reliability Engineering jobs in Kuala Lumpur
Undisclosed

Bandar Kuala Lumpur, WP Kuala Lumpur

Near Train Station
  • Design and implement secure, scalable infrastructure on Google Cloud Platform (GCP).
  • Lead efforts to build and evolve MOVE’s GCP Landing Zone, including Shared VPC, org structure, IAM, and policy guardrails
  • Build and improve multi-region architectures for high availability and disaster recovery. ...

Be an early applicant!

Posted
11 days ago
Undisclosed

Bandar Kuala Lumpur, WP Kuala Lumpur

Near Train Station
  • Design and implement secure, scalable infrastructure on Google Cloud Platform (GCP)
  • Lead efforts to build and evolve MOVE’s GCP Landing Zone, including Shared VPC, org structure, IAM, and policy guardrails
  • Build and improve multi-region architectures for high availability and disaster recovery ...

Be an early applicant!

Posted
11 days ago
Chat Available
Undisclosed

Bandar Kuala Lumpur, WP Kuala Lumpur

Near Train Station
  • Key Responsibilities
  • • Infrastructure Automation: design and standardize the team's automation scripts (Bash/Python) and Ansible Playbook standards; drive IaC adoption; review automation proposals; significantly reduce manual-ops workload
  • • Multi-Stack Environment Management: plan and maintain Web services and application architectures; support deployment and tuning across Java, Python, Go, PHP, Node.js (Nginx, Tomcat, Docker, K8s) ...
Linux Shell
+3

Be an early applicant!

Posted
14 days ago
Chat Available
MYR7,000 - MYR10,000 Per Month

Bandar Kuala Lumpur, WP Kuala Lumpur

Near Train Station
  • Responsible for Windows and Linux based IT infrastructure maintenance and implementation;
  • Provide technical support and system administration on cloud platform to meet service level objectives.
  • Execution of data backup & recovery strategy, system disaster recovery and business continuity of the core systems; ...
Linux Administration Network Engineering
+1
Posted
a month ago
Undisclosed

KL City

  • Architect and lead the cross-org design, deployment, and ongoing evolution of Kubernetes clusters, database clusters (PostgreSQL/MongoDB/DynamoDB), observability stacks (Prometheus/Grafana/Loki/Tempo), CI/CD platforms (ArgoCD/Github Actions), VPCs, and firewalls using Terraform, GitOps, and policy-as-code to deliver 99.99%+ reliability, proven multi-region resiliency, and cost reduction at scale.
  • Write, update, and simplify all technical documentation (runbooks, RCAs, KB articles, architecture decisions) in Confluence using strict KISS principles immediately after every change to eliminate tribal knowledge and ensure zero ambiguity for team collaborations.
  • Orchestrate safe, fast deployments and database migrations or upgrades across our core services using ArgoCD/Github Actions progressive delivery, automated canaries, and scripted rollbacks. Meanwhile, keeping the CI/CD pipelines simple, reliable, and self-service. So developers ship multiple times per day with zero customer impact. ...
Posted
6 days ago
Undisclosed

KL City

  • Take ownership of one meaningful project end to end — scoped so you can ship it, not so you can survive it
  • Build and refine dashboards and alerts, and learn the difference between a signal and a distraction
  • Write Terraform that goes into our real estate, reviewed like anyone else's ...
Posted
3 days ago
Undisclosed

KL City

  • Lead by example for site reliability engineering execution and manage stages from ideation and development to launch and ongoing maintenance, ensuring timely delivery and high-quality technology infrastructure that meet customer needs, and incorporating feedback loops for iterative improvements.
  • Identify and assist on the stability, scalability, cost optimisation, and security of the technology infrastructure by continuously assessing, upgrading, and implementing best practices, using specific frameworks or standards to support business growth and protect sensitive data.
  • Manage high-performing site reliability engineering teams by recruiting top talent, providing mentorship, and fostering professional growth, with a focus on long-term team development to create a collaborative and innovative work culture. ...
Posted
19 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
14 days ago
Undisclosed

KL City

  • Architect and lead the cross-org design, deployment, and ongoing evolution of Kubernetes clusters, database clusters (PostgreSQL/MongoDB/DynamoDB), observability stacks (Prometheus/Grafana/Loki/Tempo), CI/CD platforms (ArgoCD/Github Actions), VPCs, and firewalls using Terraform, GitOps, and policy-as-code to deliver 99.99%+ reliability, proven multi-region resiliency, and cost reduction at scale.
  • Write, update, and simplify all technical documentation (runbooks, RCAs, KB articles, architecture decisions) in Confluence using strict KISS principles immediately after every change to eliminate tribal knowledge and ensure zero ambiguity for team collaborations.
  • Orchestrate safe, fast deployments and database migrations or upgrades across our core services using ArgoCD/Github Actions progressive delivery, automated canaries, and scripted rollbacks. Meanwhile, keeping the CI/CD pipelines simple, reliable, and self-service. So developers ship multiple times per day with zero customer impact. ...
Posted
a month ago
Undisclosed

KL City

  • Job Description/ Responsibility :
  • • Provides new capabilities for Windows & Linux running SAP workloads on various server hardware platform on either on-prem or cloud
  • • Proactively maintain Windows/Linux infrastructure technology to maintain uptime server ...
Posted
5 days ago
Undisclosed

KL City

  • competitive salary
  • Indefinite contract
  • Health insurance ...
Posted
6 days ago
Undisclosed

KL City

  • Handle SRE role for assigned cloud services owning the KPIs for service reliability, issue to resolution, service deployment, business continuity management, security policy planning, capacity planning, Automation ,etc.
  • Automation: Automate routine and manual operations tasks to reduce "toil" and improve efficiency.
  • Monitoring & Alerting: Implement and use monitoring systems to track system health, set up alerting, and create dashboards. ...
Posted
11 days ago
Undisclosed

KL City

  • Candidates should possess strong skills in Site Reliability Engineering, including reliability-focused design, observability, incident response, and performance optimization.
  • Candidates should possess troubleshooting skills for complex distributed systems, including log analysis, monitoring, and root cause identification.
  • Candidates should possess software development skills, preferably in languages commonly used for infrastructure and tooling (e.g., Python, Go, or similar), and experience with CI/CD pipelines. ...
Posted
11 days ago
Undisclosed

KL City

  • Strong foundation in Site Reliability Engineering, including reliability, scalability, and performance practices.
  • Proficiency in Troubleshooting complex production issues and performing root cause analysis.
  • Experience with Software Development (e.g., scripting or programming languages such as Python, Go, or Java) to build tools and automation. ...
Posted
12 days ago
Undisclosed

KL City

  • VCF Infrastructure Design and Maintenance:
  • Resource Management and Troubleshooting:
  • Backup Services: ...
Posted
15 days ago
Undisclosed

KL City

  • Design, build, and maintain scalable, secure, and highly available cloud infrastructure
  • Own and continuously improve deployment platforms, CI/CD pipelines, and infrastructure automation
  • Partner with engineering teams to provision infrastructure, manage cloud resources, databases, access controls, and platform services ...
Posted
12 hours ago
Undisclosed

KL City

  • Lead the design and implementation of highly available, secure, and scalable banking infrastructure using infrastructure as code (IaC) principles
  • Establish and maintain SLOs/SLIs that define our reliability standards and drive accountability across engineering teams
  • Serve as an incident commander during critical service disruptions, leading cross-functional response teams with calm expertise ...
Posted
12 hours ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
12 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
13 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
14 days ago
Undisclosed

KL City

  • Design and operate the SRE practice for Managed oferings, including on-call processes, SLA frameworks, incident response playbooks, and post-incident review (PIR) processes.
  • Build and maintain observability infrastructure: centralised logging (correlation IDs), metrics dashboards, distributed tracing, and alerting for the Predator/Instinct platform stack.
  • Define and track SLOs (Service Level Objectives) and error budgets for real-time transaction processing pipelines, targeting high TPS and low round-trip latency. ...
Posted
14 days ago
Undisclosed

KL City

  • Minimum 10 years of system administration experience in an (preferably) international setting.
  • Minimum 2 years of experience leading project.
  • Familiarity or experience with data ingestion with big data technologies (Elastic Search, Logstash, Kibana and Kafka). ...
Posted
8 days ago
Undisclosed

KL City

  • Opportunity to be part of the transformational journey
  • Embark on a groundbreaking startup journey set to transform global supply chain
  • Build CI/CD pipelines by introducing automation, reliability controls, and deployment safeguards. Integrate and optimize observability and monitoring tools to strengthen system visibility, detection, and recovery capabilities. ...
Posted
19 days ago
Undisclosed

KL City

  • Build CI/CD pipelines by introducing automation, reliability controls, and deployment safeguards. Integrate and optimize observability and monitoring tools to strengthen system visibility, detection, and recovery capabilities.
  • Build and maintain self-healing and auto-remediation capabilities to minimize manual intervention and accelerate issue resolution.
  • Design, develop, and implement automation solutions to improve system reliability, operational efficiency, and platform resilience. ...
Posted
19 days ago
Undisclosed

KL City

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes. ...
Posted
21 days ago
Undisclosed

KL City

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes. ...
Posted
22 days ago
Undisclosed

KL City

  • Handle SRE role for assigned cloud services owning the KPIs for service reliability, issue to resolution, service deployment, business continuity management, security policy planning, capacity planning, Automation ,etc.
  • Automation: Automate routine and manual operations tasks to reduce "toil" and improve efficiency.
  • Monitoring & Alerting: Implement and use monitoring systems to track system health, set up alerting, and create dashboards. ...
Posted
a month ago
Undisclosed

KL City

  • Bachelor's degree in Computer Science, Network or related field
  • Professional cloud certification
  • Proven 5 experience in a Cloud Network or Cloud Infrastructure role ...
Posted
19 days ago
Undisclosed

KL City

  • Lead the design and implementation of highly available, secure, and scalable banking infrastructure using infrastructure as code (IaC) principles
  • Establish and maintain SLOs/SLIs that define our reliability standards and drive accountability across engineering teams
  • Serve as an incident commander during critical service disruptions, leading cross-functional response teams with calm expertise ...
Posted
a month ago
Undisclosed

KL City

  • Collaborate with global teams to complete the daily ops and alarm handling.
  • Identify and implement solutions on stability, scalability and security of business infrastructure using frameworks and industry best practices.
  • Drive and manage technical and solution architecture discussions between global teams and partners to ensure timely delivery that meet customer needs. ...
Posted
4 days ago