Design and develop resource scheduling systems for machine learning platforms, supporting model training, evaluation, and inference workloads across domains such as NLP, Computer Vision (CV), and Speech.
Optimize the orchestration and scheduling of heterogeneous computing resources, including GPUs, CPUs, and other specialized accelerators, to maximize the utilization of dedicated, opportunistic (spot), co-located, and multi-cloud resources.
Develop scheduling solutions that optimize the allocation of compute resources, RDMA high-speed networking, and storage resources, enabling large-scale distributed clusters to achieve maximum performance and efficiency.
...
Design and develop resource scheduling systems for machine learning platforms, supporting model training, evaluation, and inference workloads across domains such as NLP, Computer Vision (CV), and Speech.
Optimize the orchestration and scheduling of heterogeneous computing resources, including GPUs, CPUs, and other specialized accelerators, to maximize the utilization of dedicated, opportunistic (spot), co-located, and multi-cloud resources.
Develop scheduling solutions that optimize the allocation of compute resources, RDMA high-speed networking, and storage resources, enabling large-scale distributed clusters to achieve maximum performance and efficiency.
...
Design and develop resource scheduling systems for machine learning platforms, supporting model training, evaluation, and inference workloads across domains such as NLP, Computer Vision (CV), and Speech.
Optimize the orchestration and scheduling of heterogeneous computing resources, including GPUs, CPUs, and other specialized accelerators, to maximize the utilization of dedicated, opportunistic (spot), co-located, and multi-cloud resources.
Develop scheduling solutions that optimize the allocation of compute resources, RDMA high-speed networking, and storage resources, enabling large-scale distributed clusters to achieve maximum performance and efficiency.
...
Responsible for the iteration of the underlying architecture of the large model inference engine and end-to-end GPU performance optimization, through means such as operator fusion and compilation optimization, deeply optimizing GPU memory access, computing pipeline, and Stream asynchronous scheduling, eliminating inference computing bottlenecks, improving single-card inference throughput, and reducing inference latency.
Adapt to all series of GPU/NPU hardware architectures, refine the universality of the inference engine and hardware adaptability, and build a high-performance, low-loss underlying base for large model inference.
Lead the design, development, and optimization of distributed parallel solutions for large model inference scenarios, with a focus on implementing multi-dimensional parallel strategies such as tensor parallelism (TP), pipeline parallelism (PP), sequence parallelism, and MoE expert parallelism, to address core issues such as multi-card splitting and deployment of ultra-large models, high cross-card communication overhead, load imbalance, and low parallel efficiency.
...
For training track, develop the Volcano Ark training platform, enabling both internal and external users to perform serverless post-training (e.g., Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL)) on the Ark platform.
Design elastic training solutions for complex multi-tenant workloads, supporting mixed-tenant training across multiple data centers and heterogeneous hardware while optimizing training throughput, resource utilization, and system stability.
Build next-generation reinforcement learning infrastructure to improve training efficiency, while designing intuitive and developer-friendly APIs for RL training workflows.
...