jobs in Oxydata Software

Full Time Senior Data Centre Operations Engineer Jobs, in Oxydata Software Johor - Maukerja

Senior Data Centre Operations Engineer

Undisclosed
Share
Save

Working Location

  • Senai Johor Malaysia

Job Description

Responsibilities

Senior Data Centre Operations Engineer

Location: Senai, Johor, Malaysia
Work Mode: Onsite

Employment type: Permanent

Our client is a leading specialist in the repair and maintenance of high-end AI computing infrastructure, with a state-of-the-art facility located in Johor Bahru, Malaysia. They are dedicated to providing mission-critical support and have established a reputation for precision and reliability in the Southeast Asian market. With a strong focus on transparency and accountability, they ensure that every repair process is documented and approved by clients, maintaining a high standard of service excellence.

We are seeking an experienced Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.

Responsibilities

  • Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
  • Monitor server health and review IPMI, BMC, and operating system logs.
  • Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
  • Manage RAID configurations and monitor SSD and NVMe health.
  • Troubleshoot GPU servers and replace faulty hardware components.
  • Support GPU cluster performance, network topology, and system stability.
  • Collaborate with networking, storage, and virtualisation teams to resolve technical issues.
  • Automate routine tasks such as firmware upgrades, inspections, and hardware alert management.
  • Prepare technical guides, troubleshooting documentation, and standard operating procedures.
  • Coordinate with hardware vendors and manage RMAs and spare parts.
  • Monitor rack power, temperature, and air-cooling conditions.

Requirements

Must-have:

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
  • Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
  • At least 5-7 years of experience in server operations.
  • Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
  • Strong Linux administration and troubleshooting knowledge.
  • Good hardware troubleshooting and problem-solving skills.
  • Ability to work effectively with internal teams and external vendors.
  • Good technical communication skills in English.

Nice-to-have:

  • Experience operating large-scale GPU clusters.
  • Experience with Shell or Ansible automation.
  • Familiarity with InfiniBand, RoCE, RDMA networking and optical modules.
  • Exposure to air-cooled server environments.
  • Experience with liquid-cooled servers.
  • Knowledge of Ceph, KVM, VMware or hyper-converged infrastructure.
  • Relevant certifications such as RHCE, RHCA, CompTIA Server+ or server vendor certifications.

Education:

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.

Why Join Us

  • Be part of a dynamic team at the forefront of data centre technology, supporting mission-critical AI infrastructure for leading enterprises.
  • Opportunity to work with advanced hardware, collaborate with skilled professionals, and contribute to the reliability of high-performance computing environments across the region.

Apply Now: *************

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More