300+ Site Reliability Engineering Jobs in Malaysia | Job Vacancies | September 2026 | Maukerja

Showing 314 jobs results for "site reliability engineering"
Never miss any updates for Site Reliability Engineering jobs
Undisclosed

Singapore

  • Develop deep technical expertise in your assigned product area and tech stack.
  • Own production deployment, configuration, and release processes
  • Drive performance, reliability, and operability through continuous improvement ...
Posted
6 days ago
Undisclosed

KL City

  • Handle SRE role for assigned cloud services owning the KPIs for service reliability, issue to resolution, service deployment, business continuity management, security policy planning, capacity planning, Automation ,etc.
  • Automation: Automate routine and manual operations tasks to reduce "toil" and improve efficiency.
  • Monitoring & Alerting: Implement and use monitoring systems to track system health, set up alerting, and create dashboards. ...
Posted
8 days ago
Undisclosed

Singapore

  • Advance knowledge of core AWS services: EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM.
  • Designing multi-AZ and multi-region highly available architectures.
  • Strong understanding of networking in AWS (subnets, routing tables, NAT, security groups, NACLs, VPC peering, PrivateLink). ...
Posted
9 days ago
Undisclosed

Singapore

  • Strong experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
  • Strong knowledge of Linux, shell scripting, Git, CI/CD, Docker, and Kubernetes.
  • Experience with observability platforms, logging, metrics, tracing, alerting, and dashboarding. ...
Posted
9 days ago
SGD10,000 - SGD10,000 Per Month

Singapore

  • Job Summary
  • SGX is hiring Site Reliability Engineers who treat operations as a software problem. You'll keep production healthy, but more importantly you'll build the automation, tooling, and agentic workflows that make running our systems boring and predictable. This is an engineering role - if your instinct on a recurring issue is to write code that removes it, you'll fit in well.
  • We operate in a regulated capital-markets environment, so the bar for reliability, security, and operational rigor is high. ...
Posted
9 days ago
Undisclosed
  • Role Summary
  • Qualifications
  • Experian is a global data and technology company, powering opportunities for people and businesses around the world. We operate across a range of markets, from financial services to healthcare, automotive, agribusiness, insurance, and many more. Experian invests in people and new advanced technologies to unlock the power of data. We have an amazing team of 25,200 people in 32 countries. ...
Posted
9 days ago
Undisclosed
Posted
9 days ago
Undisclosed

Singapore

  • Advance knowledge of core AWS services: EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM.
  • Designing multi-AZ and multi-region highly available architectures.
  • Strong understanding of networking in AWS (subnets, routing tables, NAT, security groups, NACLs, VPC peering, PrivateLink). ...
Posted
10 days ago
Undisclosed
  • Delivery of high-quality infrastructure as code solutions and CI/CD pipelines
  • Implementation of monitoring solutions for client-facing products and internal data pipelines, including intelligent alarming for quicker incident detection and resolution
  • Track and manage infrastructure & application vulnerabilities ...
Posted
10 days ago
Undisclosed
  • Delivery of high-quality infrastructure as code solutions and CI/CD pipelines
  • Implementation of monitoring solutions for client-facing products and internal data pipelines, including intelligent alarming for quicker incident detection and resolution
  • Track and manage infrastructure & application vulnerabilities ...
Posted
9 days ago
Undisclosed

Singapore

  • Advance knowledge of core AWS services: EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM.
  • Designing multi-AZ and multi-region highly available architectures.
  • Strong understanding of networking in AWS (subnets, routing tables, NAT, security groups, NACLs, VPC peering, PrivateLink). ...
Posted
11 days ago
Undisclosed

KL City

  • Job Description:
  •  Responsible for deployment, change, issues triage and infra management of overseas games and relevant components and system, such as game monitor system, login services.
  •  Responsible for monitoring and dash-boarding for game observability, and ensure the game is reliable, scalable and secure ...
Posted
11 days ago

TalentVibe Business Consultancy

SGD5,000 - SGD6,000 Per Month

Singapore

  • Design, build, and maintain robust operational monitoring, structured logging, performance metrics, distributed tracing, alerting, and reporting systems for the policy engine.
  • Integrate core infrastructure and synchronization tasks with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
  • Develop real-time dashboards to provide visibility into policy engine reliability, data synchronization status, and overall system health. ...
Posted
11 days ago
Undisclosed

Singapore

  • Advance knowledge of core AWS services: EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM.
  • Designing multi-AZ and multi-region highly available architectures.
  • Strong understanding of networking in AWS (subnets, routing tables, NAT, security groups, NACLs, VPC peering, PrivateLink). ...
Posted
11 days ago
SGD7,000 - SGD12,000 Per Month

Outram

Posted
12 days ago
Undisclosed

Singapore

  • Provide L2/L3 support for business-critical production applications
  • Troubleshoot and resolve application, infrastructure, and platform incidents
  • Perform root cause analysis and implement preventative improvements ...
Posted
14 days ago
Undisclosed

KL City

  • Maintain and improve the availability, performance, and resilience of cloud infrastructure across Alibaba Cloud (Alicloud) and AWS through proactive monitoring, incident management, and root cause analysis.
  • Implement automation solutions that reduce operational effort, improve consistency, and strengthen system reliability at scale.
  • Establish and maintain observability practices, including monitoring, alerting, logging, and performance analysis, to identify and resolve issues before they impact users. ...
Posted
14 days ago
Undisclosed

Singapore

  • Own the production environment, driving performance, reliability, and operability through continuous improvement
  • Proactively monitor and troubleshoot large-scale trading systems and exchange connectivity
  • Build and maintain devops toolkit for the production trading system including configuration management, process management, deployment, monitoring, data collection, and analysis ...
Posted
14 days ago
Undisclosed
  • Design, implement, and maintain VMware Cloud Foundation (VCF) infrastructure to support GE’s organizational requirements.
  • Manage and troubleshoot VCF resource availability, including compute, memory, and storage (SAN and vSAN) up to 160 ESXi and more than 1200 VMs across multiple clusters located in both SG and MY, using tools such as NSX-T, vCenter, ESXi 8.x, and VMware vSphere Cluster availability.
  • Experience in integration with backup services using NetBackup (NBU) such as HotAdd for image backup/restore and Media to file level backup/restore to support business application VMs requirements, including full, incremental, and ad-hoc backups. ...
Posted
16 days ago
Undisclosed

Singapore

  • Build and maintain the core infrastructure of the AIOps platform, including the unified monitoring & alerting system and the FinOps cost observability platform.
  • Maintain and continuously optimize internal R&D infrastructure (GitLab, Nexus, Sonar, etc.).
  • Manage monitoring data collection, alert governance, and cost data visualization across multi-cloud environments (Alibaba Cloud / AWS). ...
Posted
3 days ago
Undisclosed

Singapore

  • Monitor and maintain production applications and platforms, ensuring high availability and performance.
  • Provide L2/L3 production support, troubleshoot incidents, and drive root-cause analysis.
  • Develop and maintain Java-based applications, services, and automation scripts. ...
Posted
3 days ago
Undisclosed
  • Looking for someone available locally.
  • Please do not proceed if you are from overseas.
  • Job Summary: ...
Posted
3 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
3 days ago
Undisclosed

Beijing

  • Provide automated solutions for manual tasks and challenges we're facing.
  • Manage cloud infrastructure using Infrastructure as Code solution.
  • Get involved in deep diagnosis of incidents, and engage with multiple accomplished engineering teams on resolutions. ...
Posted
2 days ago
Undisclosed

Singapore

  • Reliability Engineering: Define enterprise reliability strategy and resilience objectives.
  • Observability: Establish enterprise observability strategy and investment roadmap.
  • Incident Management: Define the enterprise incident management framework and resilience standards. ...
Posted
2 days ago
Undisclosed

KL City

  • Ensure the stability, reliability and efficient operation of the Mobile Business Group and overseas sales‑service business, and maintain high availability of services at all times.
  • Provide technical support for overseas businesses, covering core O&M tasks including resource delivery, incident handling, capacity management, resource governance, monitoring management and quality analysis.
  • Own and review technical architecture designs, evaluate the rationality of business architectures, proactively identify and assess risks, and drive or lead risk‑mitigation initiatives. ...
Posted
2 days ago
Undisclosed

KL City

  • Handle SRE role for assigned cloud services owning the KPIs for service reliability, issue to resolution, service deployment, business continuity management, security policy planning, capacity planning, Automation ,etc.
  • Automation: Automate routine and manual operations tasks to reduce "toil" and improve efficiency.
  • Monitoring & Alerting: Implement and use monitoring systems to track system health, set up alerting, and create dashboards. ...
Posted
17 days ago
Undisclosed

Singapore

  • Responsible for monitoring Shopee & Monee’s backbone network, data center network, and related networks. Ensure the rapid detection, notification, location, and resolution of network faults to maintain application stability and availability.
  • Propose and implement standard designs and optimizations for network architecture.
  • Enhance visibility of network infrastructure by participating in and refining standard processes. Collect and establish databases for network devices, links, and configurations. Set up accurate and automated network alerts and notifications using software tools. ...
Posted
2 days ago
Undisclosed

Singapore

  • Managing production Order Management Systems and Market Data Delivery Systems connecting to major electronic exchanges and multiple asset classes.
  • Working closely with trading teams, risk, business management and compliance, understanding their requirements and coordinating the implementation of appropriate production solutions.
  • Troubleshooting complex production issues and providing L1/L2 support across trading systems. ...
Posted
17 days ago