Job Summary:
We are seeking an Embedded LLM Systems Engineer to design, develop and optimise LLM inference solutions for embedded, mobile and edge devices. The role covers inference-engine development, model compression, heterogeneous hardware optimisation and efficient model deployment.
Duties/ Responsibilities:
On-Device Inference Engine Development
- Design, develop and optimise LLM inference engines for embedded, mobile and edge devices.
- Work on operator development, graph optimisation, memory management and multi-backend adaptation.
- Develop solutions using frameworks such as *************, TensorRT-LLM, MNN, ONNX Runtime or comparable technologies.
- Research and apply quantisation techniques such as INT4, INT8 and FP16.
- Work with relevant technologies such as NEON/SVE, Vulkan Compute, OpenCL or comparable platforms.
- Conduct training-inference consistency validation and support efficient deployment across cloud and edge environments.
- Evaluate practical ways to translate emerging AI capabilities into embedded product applications.
Basic Requirements:
- At least three years of relevant experience in on-device inference, AI infrastructure, embedded systems engineering or a related area.
- Bachelor’s degree or above in Computer Science, Electrical/Electronic Engineering, Mathematics or a related discipline, or equivalent practical experience.
- Proficiency in written and spoken English sufficient to read technical documentation and research papers, participate in technical discussions and prepare clear engineering documentation.
- Strong proficiency in modern C++, with a good understanding of memory models, concurrency and low-level performance optimisation.
- Proficiency in Python for model conversion, evaluation, automation scripts and training-related tooling.
- Experience with CUDA, MediaPipe or related technologies is advantageous.
Preferred Qualifications
Any of the following would be advantageous:
- Contributions to established open-source inference projects such as *************, vLLM, TensorRT-LLM, MLC-LLM or MNN.
- Publications in recognised conferences or journals on efficient inference, model compression or on-device deployment.
- Recognition in competitions such as ACM-ICPC, NOI, Kaggle or on-device AI challenges.
- Experience with prompt engineering, Retrieval-Augmented Generation or AI agent frameworks such as LangChain or LlamaIndex.
- Hands-on experience deploying and optimising inference frameworks such as vLLM, TGI, *************, TensorRT-LLM or MLC-LLM.
- Knowledge of model alignment and fine-tuning techniques, including RLHF, SFT and DPO.
- Practical experience fine-tuning or evaluating models using authorised, organisation-owned or appropriately licensed datasets.
- Strong interest in emerging LLM technologies and the ability to translate new capabilities into practical product value.