- Islandwide (Singapore) Singapore
Working Location
Job Description
Responsibilities
Summary:
· The successful candidate will be expected to perform a hands-on engineering role, i.e. solving the problem first prior to escalation.
· The candidate is able to read application code, reproduce defects, and land fixes.
· He /she will work alongside the Site Reliability Engineers, who own the longer-horizon work of stopping the same problem recurring.
Must have:
· Minimum 3 years of hands-on experience in resolving production defects in an application support, L2/L3 or run engineering role on systems with real users and real consequences.
· Able to debug someone else's code (Python primarily, TypeScript usefully). Able to take a stack trace and a vague user report and end up at a specific line or configuration value. This is the core skill of the role.
· Able to perform root cause analysis as a discipline. Forming and eliminating hypotheses, stating what is ruled out and why, and distinguishing the cause from the first thing being noticed.
· Possess fluent log and trace reading - Datadog, Splunk, ELK, CloudWatch or similar starting from" a user says it's broken" and getting to a specific failing component.
· Able to perform practical Kubernetes diagnosis. Comfortable with Kubectl: inspecting pods, events, logs and restarts, and recognizing the common failure shapes. Not required to build clusters.
· Experience in scripting using Python or Bash (a must-have requirement). Able to reduce the volume of work and not only absorbs it.
· Clear written skills, both for an RCA a senior engineer will read and for a status update an investment professional will read. Able to handle senior stakeholders with composure and calm.
· A solve-first, escalate-second mindset. Escalation is what happens when the candidate has genuinely exhausted what he/she can do, with investigation attached (hand over a diagnosis instead of hand over a symptom).
· Possess positive learning and collaborative mindset.
· Strong analytical, problem-solving and troubleshooting skills.
· Good written and verbal communication skills.
· Agile, fast learner and able to adapt to changes.
Good to have:
· Experience in provisioning access in an enterprise directory - AD groups, entitlement or approval workflows.
· Deployment or release verification experience.
Experience and knowledge in AWS fundamentals - to navigate and understand whatsits where.
Experience in building Datadog dashboards and monitors rather than only reading them.
· Experience in supporting an AI or LLM product, where the same input does not always produce the same output and "wrong answer" is a different class of problem from "error".
· Familiarity with databases - SQL and non-SQL.
· Experience in working within a formal change management process.
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.