- Singapore
Working Location
Job Description
Responsibilities
Overview
I'm partnering a rapidly expanding AI infrastructure and cloud computing company to hire a Senior Technical Support Engineer, GPU & Cloud Infrastructure in Singapore. This is a senior technical role focused on resolving complex customer and production issues across GPU compute environments while working closely with Engineering and SRE teams to strengthen platform reliability.
Responsibilities
You will take end-to-end ownership of complex escalations and platform incidents, troubleshooting across GPU compute, Linux, networking/SDN, storage, CUDA and drivers, control plane, and billing systems. Working closely with SRE, Compute, and Engineering teams, you will lead root-cause analysis, identify permanent solutions to recurring issues, and translate operational learnings into stronger monitoring, processes, and runbooks.
You will also play an important role during major incidents and post-incident reviews, while mentoring L1/L2 engineers and improving the quality of technical documentation and escalation practices. The position participates in an on-call escalation rotation.
Requirements
You should have at least 6 years of experience in cloud infrastructure technical support, escalation engineering, production operations, or an SRE-adjacent role, with strong hands-on expertise across Linux, networking, and cloud infrastructure. Experience supporting GPU infrastructure, CUDA, HPC, or large-scale compute environments will be particularly relevant.
You will bring a strong track record of diagnosing difficult production issues, performing structured root-cause analysis, and collaborating effectively with engineering teams to drive issues through to resolution.
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.