Job Title: Site Reliability Engineer (SRE)
Location: Singapore
Work Mode: Hybrid
Salary Range: SGD 7,500 – SGD 9,100 per month
About the Role
We are seeking a highly experienced Site Reliability Engineer (SRE) to join our engineering team in Singapore. This role requires a hands-on professional who can operate independently from day one, ensuring the reliability, scalability, and security of enterprise-grade platforms.
You will play a key role in deploying, operating, and stabilizing critical systems across multiple environments while collaborating closely with security, infrastructure, and application teams.
Key Responsibilities
- Deploy and operationalize vendor security solutions across Dev, QA, UAT, and Production environments.
- Establish production-grade reliability, including logging, monitoring, observability, alerting, and resilience.
- Integrate systems with enterprise SIEM and Incident Response workflows for monitoring, alerting, and escalation.
- Implement environment and configuration management, CI/CD pipelines, and deployment automation.
- Ensure platforms meet enterprise standards for availability, scalability, security, and recoverability.
- Support production readiness reviews, pilot runs, and transition to support teams.
- Troubleshoot issues across application, infrastructure, networking, and cloud layers.
Requirements (Must Have)
- Proven experience as an SRE, DevOps Engineer, Platform Engineer, or similar role.
- Hands-on experience managing production systems across Dev, QA, UAT, and Production.
- Strong expertise in Linux, shell scripting, Git, CI/CD, Docker, and Kubernetes.
- Experience with observability tools (logging, metrics, tracing, alerting, dashboards).
- Experience integrating with enterprise monitoring, SIEM, or incident response systems.
- Strong troubleshooting skills across infrastructure, containers, networking, and applications.
- Knowledge of SRE practices: SLIs, SLOs, incident management, capacity planning, and postmortems.
- Experience with APIs, microservices, secrets management, and enterprise integrations.
- Ability to collaborate across cybersecurity, engineering, and operations teams.
Good to Have
- Experience supporting AI/LLM or developer platforms in enterprise environments.
- Familiarity with tools like GitHub Copilot, Claude, or similar AI coding assistants.
- Understanding of AI security risks (prompt injection, data leakage, etc.).
- Knowledge of OWASP AI/LLM security and red-teaming concepts.
- Experience with OpenTelemetry or similar frameworks.
- Exposure to cloud platforms (AWS, Azure, GCP).
- Experience with Terraform, Helm, Argo CD, or other IaC tools.
- Knowledge of IAM, RBAC, OAuth/OIDC, and security best practices.
- Experience in regulated or high-availability environments.
Additional Information
- This role requires candidates who can work independently with minimal supervision.
- AI experience is advantageous but not mandatory.
- Only candidates with relevant experience and eligibility to work in Singapore will be considered.