About the Role
We are looking for an LLM Inference Performance Engineer to optimize large-scale LLM inference performance across TPU/GPU and other AI accelerators.
You will work across LLM inference, kernels, compilers, and runtime systems, improving latency, throughput, scalability, and overall inference efficiency.
Responsibilities
- Optimize LLM inference performance on TPU/GPU and other AI accelerators
- Develop and optimize inference backends, kernels, and runtime components
- Optimize Attention, GEMM, KV Cache, Sampling, fused kernels and other performance-critical workloads
- Work with technologies such as JAX, XLA, Pallas, CUDA, Triton or related frameworks
- Build benchmarking and profiling tools to identify performance bottlenecks
- Collaborate with model, inference, compiler, and hardware teams to improve production performance
Requirements
- Bachelor’s degree or equivalent experience in CS, Engineering, ML, Systems, or related fields
- Experience in LLM inference, ML systems, hardware acceleration, or performance optimization
- Strong programming skills in C++ and/or Python
- Understanding of GPU/TPU architecture, memory behavior, kernels, or ML workload performance
- Experience with performance profiling and benchmarking