Summary:
- The successful candidate will provide Site Reliability Engineering (SRE) and production support for the Ona platform within Developer platforms.
- Responsibilities include platform monitoring, incident troubleshooting and resolution, root-cause analysis, service reliability improvements, operational automation, runbook maintenance, and on-call support where required.
- The role will help maintain platform availability, stability, and performance while reducing manual operational effort.
Skillset Requirements.
- Minimum 5 years of strong software engineering experience with proficiency in at least one programming language, i.e. JavaScript, Java, Python, or .NET.
- Minimum 2 years of hands-on experience supporting the reliability and availability of production systems in an SRE or production support environment.
- Minimum 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
- Minimum 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous).
- Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
- Proficiency with CI/CD workflows and GitHub Actions.
- Knowledge of artifact repository management systems such as JFrog.
- Strong Linux administration skills and Shell scripting expertise.
- Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
- Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root-cause analysis, and post-incident follow-up.
- Experience in diagnosing and resolving complex system issues across multiple technology layers.
- Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
- Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
- Willingness to participate in an on-call support rotation for critical production services, where required.
- Able to optimize developer workflows and enhance developer experience.
- Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
- Excellent communication skills to work effectively with diverse engineering teams.
- Strong team-player mindset, focused on leveraging experience to help the team succeed.
- Possess positive learning and collaborative mindset.
- Strong analytical, problem-solving and troubleshooting skills.
- Good written and verbal communication skills.
- Agile, fast learner and able to adapt to changes.
SKILLS: SRE, Production support, Terraform, CloudFormation, AWS, Kubernetes, Docker, CI/CD, Python, Java, JavaScript or.NET, Splunk, CloudWatch, Datadog