jobs in Sourceo

Full Time Lead AI Platform Operations Engineer #AIDA Jobs, in Sourceo - Maukerja

Lead AI Platform Operations Engineer #AIDA

Sourceo

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

Role Summary

  • Responsible for ensuring the high availability, reliability, and performance of our OpenShift on-prem AI platform
  • Lead proactive monitoring, outage detection, and incident response to minimize downtime and operational risk
  • Design and maintain disaster recovery and business continuity processes to safeguard critical AI workloads
  • Oversee cybersecurity operations, including vulnerability management, audits, and compliance with security standards for AIDA’s AI platform
  • Collaborate closely with MLOps, LLMOps, and engineering teams to integrate automation, observability, and security best practices into platform operations

How You will Make An Impact:

  • Lead availability monitoring, outage detection, and performance optimization of our OpenShift on-prem AI platform
  • Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity
  • Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement
  • Handle security audits, compliance reporting, and ensure alignment with Singtel policies, regulatory frameworks and industry best practices
  • Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows
  • Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives
  • Lead AIDA AI platform operations function and coordinate distribution of work within team

Skills for Success:

  • Bachelor’s degree in Computer Science, Engineering, or a related field
  • 10 years of experience in on-prem platform administration and/or operations
  • Deep expertise in OpenShift platform operations and monitoring services including Elastic Search, Grafana and OpenTelemetry
  • Strong background in incident management, SRE practices, and disaster recovery design
  • Hands-on experience with security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection
  • Proficiency in infrastructure-as-code and automation scripting (Ansible)
  • Familiarity with AI/ML infrastructure (GPU, model hosting) and their operational demands
  • Knowledge of security compliance frameworks (ISO 27001, CIS, NIST)
  • Excellent problem-solving, communication, and leadership skills, especially in high-pressure incident scenarios
  • Forward thinking ability to identify possible failure scenarios and formulate effective response plans

Are you ready to say hello to BIG Possibilities?

Join Singtel to shape what's next and accelerate your career through meaningful work, continuous learning, and real impact.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More