Location: Senai, Johor, Malaysia
Work Mode: Onsite
Employment type: Permanent
Our client is a leading specialist in the repair and maintenance of high-end AI computing infrastructure, with a state-of-the-art facility located in Johor Bahru, Malaysia. They are dedicated to providing mission-critical support and have established a reputation for precision and reliability in the Southeast Asian market. With a strong focus on transparency and accountability, they ensure that every repair process is documented and approved by clients, maintaining a high standard of service excellence.
We are seeking an experienced Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.
Responsibilities
- Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
- Monitor server health and review IPMI, BMC, and operating system logs.
- Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
- Manage RAID configurations and monitor SSD and NVMe health.
- Troubleshoot GPU servers and replace faulty hardware components.
- Support GPU cluster performance, network topology, and system stability.
- Collaborate with networking, storage, and virtualisation teams to resolve technical issues.
- Automate routine tasks such as firmware upgrades, inspections, and hardware alert management.
- Prepare technical guides, troubleshooting documentation, and standard operating procedures.
- Coordinate with hardware vendors and manage RMAs and spare parts.
- Monitor rack power, temperature, and air-cooling conditions.
Requirements
Must-have:
- Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
- Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
- At least 5-7 years of experience in server operations.
- Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
- Strong Linux administration and troubleshooting knowledge.
- Good hardware troubleshooting and problem-solving skills.
- Ability to work effectively with internal teams and external vendors.
- Good technical communication skills in English.
Nice-to-have:
- Experience operating large-scale GPU clusters.
- Experience with Shell or Ansible automation.
- Familiarity with InfiniBand, RoCE, RDMA networking and optical modules.
- Exposure to air-cooled server environments.
- Experience with liquid-cooled servers.
- Knowledge of Ceph, KVM, VMware or hyper-converged infrastructure.
- Relevant certifications such as RHCE, RHCA, CompTIA Server+ or server vendor certifications.
Education:
- Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
Why Join Us
- Be part of a dynamic team at the forefront of data centre technology, supporting mission-critical AI infrastructure for leading enterprises.
- Opportunity to work with advanced hardware, collaborate with skilled professionals, and contribute to the reliability of high-performance computing environments across the region.
Apply Now: *************