jobs in INF Tech

Kerja Sepenuh Masa, Datacenter IT Infrastructure Operations Expert di INF Tech - Maukerja

Datacenter IT Infrastructure Operations Expert

INF Tech

Singapore

Kongsi
Simpan

Lokasi Kerja

  • Singapore

Penerangan Kerja

Tanggungjawab

Responsibilities

1. Own day-to-day availability, health inspection and hardware fault closure for GPU servers, CPU servers, network equipment and storage systems in the data hall.

2. Own rack-and-stack, structured cabling and optics management, together with in-rack rPDU and rack-level power distribution and capacity management.

3. Own liquid-cooling CDU and secondary-loop operations: runtime monitoring, coolant and filter maintenance, leak detection and emergency response, and quick-disconnect work procedures.

4. Own firmware and driver baseline management and the full vendor RMA cycle — fault evidence capture, ticket submission, part replacement and repair verification.

5. Own NOC / ECC monitoring operations and incident management: 24×7 coverage, alert tiering and escalation, and command of major incident response and post-incident review.

6. Run operations across multiple sites or projects in parallel, allocating people, spares and budget; manage the frontline team and vendors, and build SOPs, emergency plans and spare-parts strategy.


Requirements

1. Bachelor's degree or above in Computer Science, Electronics, Automation or a related field; 8+ years in datacenter IT equipment operations, including 3+ years in an operations management or team-lead role.

2. Strong hands-on data hall experience: rack-and-stack, cabling, part replacement and on-site fault handling for servers, network and storage equipment.

3. Solid knowledge of server hardware and out-of-band management (BMC / IPMI / Redfish), with command of firmware upgrade and fleet-wide operations methods.

4. Hands-on GPU server hardware operations: fault diagnosis and replacement of GPU boards, NVLink / NVSwitch modules, and power and cooling components.

5. Hands-on field operations for network and storage equipment: switch commissioning and configuration backup, optics and high-speed link (InfiniBand / RoCE) fault isolation.

6. Hands-on liquid-cooling CDU and secondary-loop operations, including coolant and water-quality management, leak detection and emergency response procedures.

7. NOC / ECC build-out or management experience; familiarity with ITIL or an equivalent framework and practical incident, problem and change management.

8. Proven ability to run multiple projects in parallel with strong incident command and cross-team coordination; resilient under pressure and able to support 24×7 escalation.


Preferred

• Field operations and hardware fault-handling experience with large-scale GPU clusters (1,000+ GPUs).

• End-to-end experience with liquid-cooled racks (direct-to-chip / DLC) from deployment and commissioning through to volume operations.

• Linux fundamentals and scripting to build automated inspection or fleet diagnostic tooling.

• ITIL, PMP or a vendor hardware-maintenance certification.

Peringatan Penting

Jangan pernah kongsikan maklumat bank atau kad kredit anda semasa memohon pekerjaan. Elakkan membuat sebarang pembayaran atau mengisi survey yang tidak berkaitan. Jika ada yang mencurigakan, sila laporkan iklan pekerjaan ini segera.

Lebih Lanjut