Role Overview
We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems
Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.
Key Responsibilities
Operate, maintain and troubleshoot large-scale Kubernetes-based production environments
Ensure reliability, scalability, availability and performance of critical services.
Investigate complex production issues and perform detailed root-cause analysis.
Participate in incident response and drive permanent corrective actions.
Automate repetitive operational activities and improve platform reliability.
Build and improve monitoring, alerting, logging and observability frameworks.
Define and track SLIs, SLOs and operational reliability metrics
Support Kubernetes upgrades, configuration changes, patching and platform improvements.
Work closely with application engineering, infrastructure, platform, security and DevOps teams.
Perform capacity planning, performance tuning and reliability improvements.
Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
Participate in production readiness reviews and ensure applications meet operational standards.
Mandatory Skills
Strong hands-on experience with
Kubernetes administration and troubleshooting
Strong understanding of Kubernetes architecture, including:
Pods
Deployments
StatefulSets
Services
Ingress
ConfigMaps / Secrets
RBAC
Storage
Networking
Strong
Linux systems administration and troubleshooting skills.
Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
Strong understanding of
Site Reliability Engineering principles
Experience supporting large-scale, high-availability production systems.
Strong incident management and RCA experience.
Hands-on scripting/automation experience using
Python, Bash/Shell or similar
Experience with monitoring and observability tools such as
Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent
Helm or similar Kubernetes package/deployment management tools.
GitOps experience using tools such as Argo CD or Flux.
Knowledge of service mesh concepts.
Experience with container security and Kubernetes security practices.
Experience with cloud or private-cloud infrastructure.
Familiarity with distributed systems and microservices architectures.
Exposure to performance engineering and capacity management.
Experience working in globally distributed engineering environments.