200+ Site Reliability Engineering Jobs in Malaysia | Job Vacancies | July 2026 | Maukerja

Showing 238 jobs results for "site reliability engineering"
Never miss any updates for Site Reliability Engineering jobs
Undisclosed

KL City

  • Handle SRE role for assigned cloud services owning the KPIs for service reliability, issue to resolution, service deployment, business continuity management, security policy planning, capacity planning, Automation ,etc.
  • Automation: Automate routine and manual operations tasks to reduce "toil" and improve efficiency.
  • Monitoring & Alerting: Implement and use monitoring systems to track system health, set up alerting, and create dashboards. ...
Posted
18 days ago
SGD12,500 - SGD12,500 Per Month

Singapore

  • Own the reliability, availability, and performance of the systems behind k-ID’s platform and public APIs
  • Design and improve scalable infrastructure on AWS and Kubernetes that can support high growth, uneven traffic, and global production workloads
  • Build and maintain strong observability across logs, metrics, tracing, alerting, and service health so issues are caught early and investigated quickly ...
Posted
11 days ago
MYR16,899 - MYR18,589 Per Month
WFH

Malaysia

  • The salary range for the L3 position is RM16,899 - RM18,589.
  • For more junior candidates, we may evaluate you as a L2, with a salary range of RM12,053 - RM13,258.
  • Learn more about our level structure at CoinGecko's Career Progression. ...
Posted
12 days ago
Undisclosed

Singapore

  • We are hiring for the Site Reliability Engineer (SRE) role on behalf of our global MNC client in Singapore. This is a 12-month extendable contract with a hybrid work arrangement. If you have a strong background in platform engineering, DevOps, and production operations, this is an exciting opportunity to work on enterprise-scale platforms and ensure their reliability, scalability, and operational excellence.
  • As a Site Reliability Engineer, you will be responsible for deploying and operationalizing enterprise platform solutions across development, testing, and production environments. You will establish production-grade reliability by implementing logging, observability, monitoring, alerting, and operational resilience while integrating platforms with enterprise SIEM and Incident Response workflows. The role also involves managing environment configurations, release processes, deployment automation, and ensuring platforms meet enterprise standards for availability, scalability, security, and recoverability. You will support production readiness, troubleshoot platform and infrastructure issues, and work closely with cybersecurity, infrastructure, application development, and operations teams to deliver secure and highly available services.
  • The ideal candidate has 5–8 years of experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer, with hands-on experience managing production systems across Dev, QA, UAT, and Production environments. You have strong expertise in Linux, shell scripting, Git, CI/CD, Docker, Kubernetes, observability platforms, logging, monitoring, tracing, and alerting. You are experienced in integrating enterprise monitoring, SIEM, and Incident Response workflows, and have a solid understanding of production reliability practices including SLIs, SLOs, incident management, capacity planning, and post-incident reviews. Experience with cloud platforms (AWS, Azure, or Google Cloud), Infrastructure as Code tools such as Terraform, Helm, or Argo CD, identity and access management, secrets management, and automation using Python or Go will be highly valued. Candidates with experience operating enterprise-scale, high-availability platforms in regulated environments will be at an advantage. ...
Posted
19 days ago
Undisclosed

Singapore

  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years.
  • Define the next-gen paradigm for human-AI interaction.
  • Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion. ...
Posted
12 days ago
Undisclosed

KL City

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes. ...
Posted
12 days ago
Undisclosed

KL City

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes. ...
Posted
13 days ago
Undisclosed

Singapore

  • Design, deploy, and maintain Amazon Connect infrastructure, including Contact Flows, Lambda integrations, Lex Bots, and phone number provisioning.
  • Manage AWS cloud services including EC2, ECS/EKS, Lambda, S3, DynamoDB, IAM, and VPC.
  • Build and maintain CI/CD pipelines using GitLab CI, GitHub Actions, Jenkins, or AWS CodePipeline. ...
Posted
21 days ago
Undisclosed

Singapore

  • Provide 2nd level support for production systems and critical business applications.
  • Investigate, troubleshoot, and resolve incidents and performance issues. Perform root cause analysis (RCA) and document findings in a structured manner.
  • Collaborate closely with development teams to ensure sustainable issue resolution. ...
Posted
21 hours ago
Undisclosed

Singapore

  • Responsible for the operation and maintenance as well as stability assurance of ByteDance's "Network-Traffic Infrastructure".
  • Responsible for the delivery and operation & maintenance of the production system, including the delivery, change and release of facilities, components and products, and improving the efficiency of both delivery and operation & maintenance.
  • Responsible for the design and implementation of the stability assurance system, covering system observability (monitoring/alerts/logging), troubleshooting (root cause analysis/impact assessment), and issue resolution (manual/self-healing). ...
Posted
3 days ago
Undisclosed

Singapore

  • Plaud is a bootstrapped, skyrocketing, profitable company with a $250M revenue run rate achieved in just three years.
  • Define the next-gen paradigm for human-AI interaction.
  • Gain exposure to cutting-edge AI for Pro tools and play a direct role in our global expansion. ...
Posted
16 days ago
Undisclosed

Singapore

  • Act as the first point of contact for external trading connections and play a critical role in ensuring system reliability, performance, and stability.
  • The role involves a dynamic mix of counterparty support, risk monitoring, system triage, and platform optimization.
  • Engineers operate in a high-performance Linux environment, solving complex technical issues while collaborating closely with traders and exchanges to understand market structure and ensure smooth market connectivity. ...
Posted
9 days ago
SGD5,000 - SGD5,000 Per Month

Singapore

  • Develop Automation Solutions: Design, develop, and maintain robust automation solutions using Python and Unix/Linux technologies.
  • Build RPA Workflows: Build and enhance object-based automation workflows leveraging RPA platforms and frameworks (such as Selenium or Playwright).
  • Streamline Operations: Develop reusable automation components to streamline operational processes and significantly reduce manual intervention. ...
Posted
10 days ago
Undisclosed

KL City

  • Bachelor's degree in Computer Science, Network or related field
  • Professional cloud certification
  • Proven 5 experience in a Cloud Network or Cloud Infrastructure role ...
Posted
10 days ago
Undisclosed

Singapore

  • Develop deep technical expertise in your assigned product area and tech stack.
  • Own production deployment, configuration, and release processes
  • Drive performance, reliability, and operability through continuous improvement ...
Posted
23 days ago
Undisclosed

Singapore

Posted
12 days ago
Undisclosed

Singapore

  • Owning production deployment, configuration, and release across your assigned area of the stack
  • Building and maintaining tooling for deployment, orchestration, monitoring, and diagnostics
  • Defining SLIs and SLOs in partnership with product owners, and using them to drive real improvements ...
Posted
19 days ago
Undisclosed

KL City

  • Lead the design and implementation of highly available, secure, and scalable banking infrastructure using infrastructure as code (IaC) principles
  • Establish and maintain SLOs/SLIs that define our reliability standards and drive accountability across engineering teams
  • Serve as an incident commander during critical service disruptions, leading cross-functional response teams with calm expertise ...
Posted
19 days ago
Undisclosed

Singapore

  • We're Hiring
  • DevOps & Site Reliability Engineering (SRE) Lead
  • Location: Singapore (Onsite) ...
Posted
20 days ago
SGD11,250 - SGD11,250 Per Month

Singapore

  • About TikTok
  • TikTok is the leading destination for short-form mobile video. At TikTok, our mission is to inspire creativity and bring joy. TikTok's global headquarters are in Los Angeles and Singapore, and we also have offices in New York City, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.​
  • Why Join Us ...
Posted
3 days ago
SGD6,500 - SGD6,500 Per Month

Singapore

  • About TikTok
  • TikTok is the leading destination for short-form mobile video. At TikTok, our mission is to inspire creativity and bring joy. TikTok's global headquarters are in Los Angeles and Singapore, and we also have offices in New York City, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.​
  • Why Join Us ...
Posted
3 days ago
Undisclosed

Malacca City

  • Set up and maintain applications infrastructure, and define and run code quality checks
  • Deploy application packages along the CI/CD pipeline, including solution reviews, sprint package testing (integration, performance), UAT stages, in collaboration with IT Application Owners and Quality Engineers
  • Collaborate with Platform SRE and Architects to identify and prioritize system improvements, and develop solutions to meet business needs ...
Posted
7 days ago
Undisclosed

Singapore

Posted
16 days ago
Undisclosed

KL City

  • Design and operate the SRE practice for Managed offerings, including on-call processes, SLA frameworks, incident response playbooks, and post-incident review (PIR) processes.
  • Build and maintain observability infrastructure: centralised logging (correlation IDs), metrics dashboards, distributed tracing, and alerting for the Predator/Instinct platform stack.
  • Define and track SLOs (Service Level Objectives) and error budgets for real-time transaction processing pipelines, targeting high TPS and low round-trip latency. ...
Posted
21 days ago
Undisclosed

Singapore

  • A strong believer of automating DevOps & SRE aspects like infrastructure provisioning, deployment, observability, incident lifecycle, uptime SLA etc.
  • Bold to challenge, open to get challenged, curious to learn & grow
  • Using InfrastructureAsCode tooling like Terraform or Ansible to manage AWS resources ...
Posted
a month ago
Undisclosed
  • Design, operate, and continuously optimize cloud infrastructure (AWS preferred), ensuring high availability, scalability, and security.
  • Manage cloud networking components including VPC, Transit Gateway/Peering, Load Balancers, Route Tables, Security Groups, WAF, CDN, DNS, VPN, and Direct Connect.
  • Troubleshoot complex network issues across application, infrastructure, and cloud environments. ...
Posted
14 days ago
SGD11,250 - SGD11,250 Per Month

Singapore

  • About Us
  • Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.​
  • Why Join ByteDance ...
Posted
20 hours ago
SGD6,500 - SGD6,500 Per Month

Singapore

  • About Us
  • Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.​
  • Why Join ByteDance ...
Posted
20 hours ago