jobs in NEXadept

Sepenuh Masa LLM Inference Performance Engineer Jobs, in NEXadept - Maukerja

LLM Inference Performance Engineer

NEXadept

Undisclosed

Singapore

Kongsi
Simpan

Lokasi Kerja

  • Singapore

x2_onboarding.experience.fields.job_description.title

Tanggungjawab

About the Role

We are looking for an LLM Inference Performance Engineer to optimize large-scale LLM inference performance across TPU/GPU and other AI accelerators.

You will work across LLM inference, kernels, compilers, and runtime systems, improving latency, throughput, scalability, and overall inference efficiency.


Responsibilities

  • Optimize LLM inference performance on TPU/GPU and other AI accelerators
  • Develop and optimize inference backends, kernels, and runtime components
  • Optimize Attention, GEMM, KV Cache, Sampling, fused kernels and other performance-critical workloads
  • Work with technologies such as JAX, XLA, Pallas, CUDA, Triton or related frameworks
  • Build benchmarking and profiling tools to identify performance bottlenecks
  • Collaborate with model, inference, compiler, and hardware teams to improve production performance


Requirements

  • Bachelor’s degree or equivalent experience in CS, Engineering, ML, Systems, or related fields
  • Experience in LLM inference, ML systems, hardware acceleration, or performance optimization
  • Strong programming skills in C++ and/or Python
  • Understanding of GPU/TPU architecture, memory behavior, kernels, or ML workload performance
  • Experience with performance profiling and benchmarking

job_detail.scamJob.title

job_detail.scamJob.subs

job_detail.scamJob.learnMore