Senior DevOps Operation Engineer – CI/CD Monitoring
Location: Petaling Jaya
Employment type : Contract
Contract Duration : 6 months, extendable up to 1 year.
Working arrangement : Onsite
Salary Range : Max RM10k
Open to : Malaysian or PR
Job Summary
We are seeking an experienced DevOps / Operations Engineer to manage, maintain, and optimize our cloud-native infrastructure, AI platforms, and database environments. In this role, you will be responsible for Kubernetes cluster operations, infrastructure automation, system monitoring, and troubleshooting while collaborating with cross-functional teams to ensure the availability, reliability, and performance of our production services.
Key Responsibilities
- Deploy, manage, monitor, and troubleshoot Kubernetes (K8s) clusters and containerized applications to ensure stable and highly available production environments.
- Design and develop automation scripts and internal tools using Python, Java, or Go to streamline operations and reduce manual effort.
- Administer and maintain database systems, including routine maintenance, performance tuning, backup and recovery, and incident resolution.
- Build, manage, and enhance observability platforms, including logging, metrics collection, monitoring, alerting, and performance analysis.
- Collaborate closely with software development teams to improve CI/CD pipelines, deployment automation, and software delivery processes.
- Perform daily operational health checks, respond to production incidents, conduct root cause analysis (RCA), and implement preventive improvements.
- Continuously optimize infrastructure, system reliability, and operational efficiency through automation and best practices.
Required Qualifications
- Minimum 5 years of experience in DevOps, Site Reliability Engineering (SRE), System Operations, or cloud-native environments.
- Hands-on experience deploying, operating, and troubleshooting Kubernetes clusters in production.
- Proficiency in at least one programming language: Python, Java, or Go.
- Experience supporting AI/ML platforms, services, or related infrastructure.
- Strong understanding of Linux systems, networking fundamentals, and cloud-native technologies.
- Excellent analytical, problem-solving, and troubleshooting skills.
Preferred Qualifications
- Experience with OpenSearch, Grafana, ELK Stack, and Application Performance Monitoring (APM) solutions.
- Familiarity with the Argo ecosystem, including Argo CD and Argo Workflows, for GitOps and workflow orchestration.
- Experience designing and maintaining end-to-end observability platforms and enterprise CI/CD pipelines.
- Knowledge of infrastructure automation and DevOps best practices in cloud-native environments.
If you are interested, please send your updated resume to ************* OR WA to +*************