Building next-generation generative AI infrastructure for search, advertising, and recommendation businesses.
Through the co-design of large models, multimodal technologies, and system-level innovations, we aim to overcome performance bottlenecks and enable ultra-long context handling, millisecond-level response latency, and high-precision information understanding, thereby driving intelligent upgrades across the business.
Individuals who are completing or have recently completed a PhD degree in Artificial Intelligence, Software Development, Computer Science, Computer Engineering or a related discipline.
...
Conduct novel research on distributed training and inference system optimization for large-scale recommendation models and Large Language Models
Design and implement high-performance GPU kernel architectures and communication primitives to accelerate deep learning workloads
Publish original research findings at top-tier academic conferences and collaborate with cross-functional teams to translate research into production impact
...
Design, build and optimize distributed training infrastructure and low-latency online inference systems for large-scale recommendation models and Large Language Models
Develop high-performance GPU kernel implementations and efficient inter-node communication primitives to improve training and inference efficiency
Build compiler optimization passes and operator fusion technologies for deep learning frameworks to accelerate model execution
...
Build advanced and standardized DevOps/QA tools or platforms to accelerate R&D efficiency, improving the efficiency and quality of our engineering team;
Be responsible for high-quality design, coding, and tackling highly challenging and technical problems;
Cooperate with product manager, participate in product requirement discussion, function definition, etc.
...
Deeply integrate LLM and recommendation technologies; develop, deploy, and optimize LLM training and inference techniques for recommendation foundation models; and solve large-model engineering challenges in recommendation scenarios.
...
Define and drive Service Level Objectives (SLOs) for online machine learning inference systems, ensuring the reliability, availability, and performance of large-scale production inference services;
Ensure the reliability and operational excellence of offline machine learning training pipelines, continuously improving training job success rates;
Drive infrastructure capacity planning and resource management for machine learning workloads, ensuring compute resources meet evolving business demands while continuously improving GPU and CPU utilization through performance optimization;
...
Define and drive Service Level Objectives (SLOs) for online machine learning inference systems, ensuring the reliability, availability, and performance of large-scale production inference services;
Ensure the reliability and operational excellence of offline machine learning training pipelines, continuously improving training job success rates;
Drive infrastructure capacity planning and resource management for machine learning workloads, ensuring compute resources meet evolving business demands while continuously improving GPU and CPU utilization through performance optimization;
...