You will design, build, and operate the multi-tenant Kubernetes (EKS) platform. This platform runs Grab's workloads, including Spark, Ray, Trino, Starrocks, Airflow, and Michelangelo. Additionally, your responsibilities will include cluster lifecycle, node provisioning, and autoscaling, and scheduling and resource isolation.
You will build Kubernetes operators and custom resources (kubebuilder / controller-runtime, Go) that automate the provisioning and lifecycle of compute engines and their tenants.
You will lead the infrastructure-as-code (Terraform) and CI/CD (GitLab) that provision and change our AWS estate. This estate includes EKS, S3, IAM, RDS, VPC, and networking. You will drive it towards safe, reviewable, automated change.
...
deliver complex machine learning pipelines across data preparation, model development, evaluation, deployment, and monitoring
perform applied research on deep learning models for medical image segmentation, detection, and classification across imaging modalities (e.g. radiographs, CBCT, volumetric data)
design rigorous evaluation frameworks by preparing clinically meaningful metrics and test sets
...
Model & Prompt Engineering: Design, test, and iterate prompts that improve reasoning, factuality, and user experience.
Agentic AI Frameworks: Build, manage and orchestrate agentic workflows involving tool calling, reasoning loops, memory management, skills orchestration, and multi-step task execution.
AI Safety & Reliability: Implement guardrails, monitoring, and validation mechanisms to ensure AI systems remain safe, reliable, and ethically grounded
...
Model & Prompt Engineering: Design, test, and iterate prompts that improve reasoning, factuality, and user experience.
Agentic AI Frameworks: Build, manage and orchestrate agentic workflows involving tool calling, reasoning loops, memory management, skills orchestration, and multi-step task execution.
AI Safety & Reliability: Implement guardrails, monitoring, and validation mechanisms to ensure AI systems remain safe, reliable, and ethically grounded
...
Agentic AI ArchitectureDesign and implement production-grade agentic systems using models such as Claude. Agent orchestration, tool calling, function calling, multi-agent architectures, planning and task decomposition, agent memory, context management, state machines and workflow engines, long-running agents, human-in-the-loop systems, autonomous execution, recovery and retry mechanisms, observability, and evaluation. You understand the difference between LLM → Agent → Workflow → Autonomous System, and when each is appropriate.
Agentic Loops & Self-OptimizationA major part of the role is building closed-loop systems: Goal → Plan → Execute → Observe → Evaluate → Learn → Re-plan → Execute. Systems that evaluate their own outputs, detect failed actions, identify root causes, adjust strategies, optimize prompts and tool selection, keep what works, roll back what does not, and improve over time. Reflection, critique, self-evaluation, feedback loops, reward signals, evaluation frameworks, automated experimentation, memory, retrieval, state management.
Claude / LLM EngineeringDeep practical experience with Claude/Anthropic APIs is highly desirable. Tool use, structured outputs, streaming, context management, prompt engineering, system prompts, long-context workflows, model routing, token optimization, latency and cost optimization, context compression, agent memory, LLM evaluation. Experience with other frontier models (OpenAI, Gemini, Llama or equivalent) is a plus.
...