200+ Site Reliability Engineering Jobs in Malaysia | Job Vacancies | July 2026 | Maukerja

Paparan 240 hasil carian kerja kosong untuk "site reliability engineering"
Jangan lepaskan peluang untuk kerja Site Reliability Engineering terkini!
SGD7,800 - SGD8,100 Sebulan

Singapore

  • POSITION GENERAL DUTIES AND TASKS:
  • Mandatory Skills (Must-Have)
  • 7+years strong experience in Production Support / SRE / BizOps (L2 Operations -hands-on troubleshooting, monitoring, and incident handling) ...
Posted
17 days ago
Undisclosed

Singapore

  • Manage a large team of Production Support Personnel across 3 geographical locations
  • Ensure SLAs on Alerts and Incidents are proactively managed and reduce in Mean Time To Recover (MTTR) by 20%
  • Ensure strict adherence to Standard Operating Procedure for recovery ...
Posted
a month ago
Undisclosed

Singapore

  • Manage a large team of Production Support Personnel across 3 geographical locations
  • Ensure SLAs on Alerts and Incidents are proactively managed and reduce in Mean Time To Recover (MTTR) by 20%
  • Ensure strict adherence to Standard Operating Procedure for recovery ...
Posted
a month ago
Undisclosed

KL City

  • Candidates should possess strong skills in Site Reliability Engineering, including reliability-focused design, observability, incident response, and performance optimization.
  • Candidates should possess troubleshooting skills for complex distributed systems, including log analysis, monitoring, and root cause identification.
  • Candidates should possess software development skills, preferably in languages commonly used for infrastructure and tooling (e.g., Python, Go, or similar), and experience with CI/CD pipelines. ...
Posted
3 days ago
Undisclosed

KL City

  • Handle SRE role for assigned cloud services owning the KPIs for service reliability, issue to resolution, service deployment, business continuity management, security policy planning, capacity planning, Automation ,etc.
  • Automation: Automate routine and manual operations tasks to reduce "toil" and improve efficiency.
  • Monitoring & Alerting: Implement and use monitoring systems to track system health, set up alerting, and create dashboards. ...
Posted
2 days ago
Undisclosed

Singapore

  • We are hiring for the Site Reliability Engineer (SRE) role on behalf of our global MNC client in Singapore. This is a 12-month extendable contract with a hybrid work arrangement. If you have a strong background in platform engineering, DevOps, and production operations, this is an exciting opportunity to work on enterprise-scale platforms and ensure their reliability, scalability, and operational excellence.
  • As a Site Reliability Engineer, you will be responsible for deploying and operationalizing enterprise platform solutions across development, testing, and production environments. You will establish production-grade reliability by implementing logging, observability, monitoring, alerting, and operational resilience while integrating platforms with enterprise SIEM and Incident Response workflows. The role also involves managing environment configurations, release processes, deployment automation, and ensuring platforms meet enterprise standards for availability, scalability, security, and recoverability. You will support production readiness, troubleshoot platform and infrastructure issues, and work closely with cybersecurity, infrastructure, application development, and operations teams to deliver secure and highly available services.
  • The ideal candidate has 5–8 years of experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer, with hands-on experience managing production systems across Dev, QA, UAT, and Production environments. You have strong expertise in Linux, shell scripting, Git, CI/CD, Docker, Kubernetes, observability platforms, logging, monitoring, tracing, and alerting. You are experienced in integrating enterprise monitoring, SIEM, and Incident Response workflows, and have a solid understanding of production reliability practices including SLIs, SLOs, incident management, capacity planning, and post-incident reviews. Experience with cloud platforms (AWS, Azure, or Google Cloud), Infrastructure as Code tools such as Terraform, Helm, or Argo CD, identity and access management, secrets management, and automation using Python or Go will be highly valued. Candidates with experience operating enterprise-scale, high-availability platforms in regulated environments will be at an advantage. ...
Posted
2 days ago
Undisclosed

Singapore

  • Develop deep technical expertise in your assigned product area and tech stack.
  • Own production deployment, configuration, and release processes
  • Drive performance, reliability, and operability through continuous improvement ...
Posted
a day ago
Undisclosed

KL City

  • Strong foundation in Site Reliability Engineering, including reliability, scalability, and performance practices.
  • Proficiency in Troubleshooting complex production issues and performing root cause analysis.
  • Experience with Software Development (e.g., scripting or programming languages such as Python, Go, or Java) to build tools and automation. ...
Posted
3 days ago
Undisclosed

Singapore

  • A fast-paced and dynamic environment working with next-gen technology. You’ll be operating at the intersection of sustainability and artificial intelligence – helping to transform an industry.
  • Working with and access to colleagues who are true innovators and leaders in their field.
  • As an emerging company, we work as a close-knit team. Work with the founders, grow a strong network, and witness the impact you make first-hand as we democratise AI tools for everyone – more sustainably, and more affordably. ...
Posted
4 days ago
Undisclosed

Singapore

  • Develop deep technical expertise in your assigned product area and tech stack
  • Own production deployment, configuration, and release processes
  • Drive performance, reliability, and operability through continuous improvement ...
Posted
4 days ago
Undisclosed

KL City

  • VCF Infrastructure Design and Maintenance:
  • Resource Management and Troubleshooting:
  • Backup Services: ...
Posted
6 days ago
Undisclosed

Singapore

  • Supporting the product engineering teams in building highly fault-tolerant, scalable applications by participating in design discussions, engaging in RFCs and code reviews.
  • Contributing to the execution of department strategies such as implementing disaster recovery, backup, redundancy, and capacity planning activities.
  • Participating in a global on-call rotation responsible for identifying and fixing bottlenecks in SaaS customer environments. ...
Posted
10 days ago
Undisclosed

Singapore

  • Own the production environment, driving performance, reliability, and operability through continuous improvement
  • Proactively monitor and troubleshoot large-scale trading systems and exchange connectivity
  • Build and maintain devops toolkit for the production trading system including configuration management, process management, deployment, monitoring, data collection, and analysis ...
Posted
10 days ago
Undisclosed
  • Delivery of high-quality infrastructure as code solutions and CI/CD pipelines
  • Implementation of monitoring solutions for client-facing products and internal data pipelines, including intelligent alarming for quicker incident detection and resolution
  • Track and manage infrastructure & application vulnerabilities ...
Posted
12 days ago
Undisclosed

Singapore

  • Responsibilities:
  • Support public cloud operations and management, including resource planning, cost optimization, account governance, and cloud infrastructure lifecycle management
  • Manage and ensure the reliability of big data platforms (e.g., Hadoop, Spark, Flink) in cloud environments ...
Posted
4 days ago
SGD10,000 - SGD10,000 Sebulan

Singapore

  • Develop and maintain secure, high-performance cloud services
  • Collaborate with architects, designers, engineers and key stakeholders to translate requirements into product features and capabilities
  • Enhance system design and architecture with your cloud expertise throughout the development lifecycle ...
Posted
3 days ago
SGD7,500 - SGD7,500 Sebulan

Singapore

  • Deploy and operationalize vendor security solutions across Dev, QA, UAT, and Production environments.
  • Establish production-grade reliability, including logging, monitoring, observability, alerting, and resilience.
  • Integrate systems with enterprise SIEM and Incident Response workflows for monitoring, alerting, and escalation. ...
Posted
19 hours ago
Undisclosed

Singapore

  • Build and maintain the core infrastructure of the AIOps platform, including the unified monitoring & alerting system and the FinOps cost observability platform.
  • Maintain and continuously optimize internal R&D infrastructure (GitLab, Nexus, Sonar, etc.).
  • Manage monitoring data collection, alert governance, and cost data visualization across multi-cloud environments (Alibaba Cloud / AWS). ...
Posted
3 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
4 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
5 days ago
Undisclosed

Singapore

  • Develop and maintain secure, high-performance cloud services
  • Collaborate with architects, designers, engineers and key stakeholders to translate requirements into product features and capabilities
  • Enhance system design and architecture with your cloud expertise throughout the development lifecycle ...
Posted
5 days ago
Undisclosed

KL City

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence. ...
Posted
5 days ago
Undisclosed

KL City

  • Design and operate the SRE practice for Managed oferings, including on-call processes, SLA frameworks, incident response playbooks, and post-incident review (PIR) processes.
  • Build and maintain observability infrastructure: centralised logging (correlation IDs), metrics dashboards, distributed tracing, and alerting for the Predator/Instinct platform stack.
  • Define and track SLOs (Service Level Objectives) and error budgets for real-time transaction processing pipelines, targeting high TPS and low round-trip latency. ...
Posted
6 days ago
Undisclosed
Kerja di Rumah

Singapore

  • Plaud is building the real-world AI interface for professionals to amplify intelligence, elevate productivity and performance, loved by over 2,000,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud captures, structures, and compounds the intelligence generated in conversations — so humans can think better, decide faster, and execute with clarity.
  • Plaud is building the next generation intelligence infrastructure and interfaces to capture, extract, and utilize intelligence from what people say, hear, see, and think.
  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years. ...
Posted
9 days ago
Undisclosed

Singapore

  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years.
  • Define the next-gen paradigm for human-AI interaction.
  • Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion. ...
Posted
9 days ago
Undisclosed

KL City

  • Opportunity to be part of the transformational journey
  • Embark on a groundbreaking startup journey set to transform global supply chain
  • Build CI/CD pipelines by introducing automation, reliability controls, and deployment safeguards. Integrate and optimize observability and monitoring tools to strengthen system visibility, detection, and recovery capabilities. ...
Posted
10 days ago
Undisclosed

Singapore

  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn every incident into a durable fix in code or automation.
  • Build the platform and tooling that make services easy to deploy, observe, and operate: CI/CD, infrastructure-as-code, observability stacks, runbooks-as-code.
  • Apply AI agentically across operations (triage, root-cause analysis, remediation, change review) and contribute to our internal agentic ecosystem. ...
Posted
18 days ago
Undisclosed
  • Delivery of high-quality infrastructure as code solutions and CI/CD pipelines
  • Implementation of monitoring solutions for client-facing products and internal data pipelines, including intelligent alarming for quicker incident detection and resolution
  • Track and manage infrastructure & application vulnerabilities ...
Posted
18 days ago
Undisclosed
  • Design, deploy, and maintain cloud infrastructure on GCP.
  • Monitor and support production systems, ensuring high availability and service reliability.
  • Respond to and troubleshoot production incidents as part of a 24/7 on-call rotation. ...
Posted
10 days ago
Undisclosed

KL City

  • Build CI/CD pipelines by introducing automation, reliability controls, and deployment safeguards. Integrate and optimize observability and monitoring tools to strengthen system visibility, detection, and recovery capabilities.
  • Build and maintain self-healing and auto-remediation capabilities to minimize manual intervention and accelerate issue resolution.
  • Design, develop, and implement automation solutions to improve system reliability, operational efficiency, and platform resilience. ...
Posted
11 days ago