We are seeking an experienced AI Hardware Engineer to support the design, deployment, validation, and troubleshooting of AI training clusters, GPU servers, networking, and storage infrastructure. The ideal candidate should have strong expertise in server hardware, GPU platforms, high-speed networking, and data center infrastructure to support large-scale AI/HPC environments.
Key Responsibilities
AI Server Hardware Management
Deploy, validate, and maintain AI GPU servers;
Perform hardware diagnostics and component replacement;
Analyze system logs, BMC logs, and hardware alerts;
Manage server hardware lifecycle.
GPU Platform Support
Deploy and validate NVIDIA GPU platforms;
Troubleshoot GPU-related;
Perform GPU benchmarking and stress testing;
Support CUDA, NCCL, and GPU fabric troubleshooting.
AI Cluster Deployment & Validation
Participate in AI/HPC cluster deployment;
Execute cluster hardware qualification testing;
Produce validation reports and documentation.
Network & Storage Support
Configure and maintain high-speed networking:
Support distributed storage systems:
Assist with performance analysis and troubleshooting.
Automation & Tool Development
Develop automation scripts for:
Hardware health checks
Cluster validation
Deployment automation
Log collection
Build tools for testing and operations.
Good communication, teamwork, and ownership mindset.
Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.
Required Qualifications
Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
Hardware
Strong knowledge of x86 server architecture;
Familiar with Intel, AMD, and NVIDIA Grace CPU platforms;
Experience with:
HGX
DGX
GB200 NVL72
GB300 NVL72
Knowledge of BMC/IPMI management.
GPU & AI Platform
Experience with NVIDIA GPU products H100, H200, B200, B300
Familiar with: CUDA ,NCCL ,NV Link ,NV Switch and GPU Direct RDMA
Linux
Strong Linux administration skills (Ubuntu, Rocky Linux);
Proficient in: Shell ,Python , Bash
Capable of independent troubleshooting.
Networking:
Strong understanding of: TCP/IP , VLAN , BGP ,OSPF ,RDMA , InfiniBand and RoCE
Preferred Qualities
Experience operating AI training clusters; Kubernetes experience; Slurm administration
PXE deployment experience; GPU Fabric Manager expertise;
Experience with hyperscale AI datacenter deployments.