jobs in Verinon

Full Time System Engineer - Infrastructure (AI - HPC Systems) Jobs, in Verinon Johor - Maukerja

System Engineer - Infrastructure (AI - HPC Systems)

Verinon

Undisclosed
Share
Save

Working Location

  • Kulai Johor Malaysia

Job Description

Responsibilities

Job Role: System Engineer - Infrastructure (AI & HPC Systems)

Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service.

Location: Kulai, Johor, Malaysia

Job Type: Full Time – On Site

Experience: 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.

Applicant: Local Malaysian citizens only

JOB DESCRIPTION

  • Deploy and manage GPU and CPU servers in GPU clusters, ensuring optimized BIOS, firmware, and OS configurations for high-performance AI workloads.
  • Maintain physical and virtual infrastructure, including server health monitoring, firmware upgrades, and hardware-level diagnostics.
  • Support large-scale cluster deployments across multiple racks; coordinate with hardware vendors and integrators on system delivery, RMA, and maintenance.
  • Implement system provisioning processes including automation for OS flashing, GPU/NIC driver installs, and baseline system hardening.
  • Conduct system-level validation, burn-in tests, and workload benchmarking for cluster readiness.
  • Monitor system health, track failure patterns, and drive corrective actions including hardware replacements and root cause analysis.
  • Work closely with networking, storage, and DevOps teams to ensure end-to-end performance and service quality.
  • Write scripts and automation tools to streamline infrastructure setup, monitoring, alerting, and remediation.
  • Maintain technical documentation for rack layouts, cabling diagrams, system configs, and operational procedures.
  • Participate in on-call rotations and provide L2/L3 support for system-related incidents.

JOB REQUIREMENTS

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
  • 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
  • Strong technical expertise in configuring, deploying, and maintaining bare metal servers, GPU nodes, and CPU-based systems.
  • Deep understanding of Linux system internals, kernel tuning, and performance optimization specific to compute-heavy workloads.
  • Familiarity with server provisioning and orchestration tools, such as IPMI, PXE boot, Redfish, or BMC tooling.
  • Basic understanding in monitoring (e.g., Prometheus, Grafana) and centralized logging tools.
  • Basic familiarity with Kubernetes or container-based environments.

Desired Skills

  • Strong interpersonal skills, with a proven ability to develop professional relationships across business and technical teams.
  • Hands-on experience with server vendors and GPU platforms.
  • Hands-on experience with storage systems (e.g., NVMe, SAN, NAS), networking concepts, and protocols (e.g., TCP/IP, RDMA) will be advantageous.
  • Knowledgeable in operating ticketing system and trouble shooting process in CPU/GPU cluster.
  • Excellent documentation skills to effectively articulate technical designs, issues, procedures, and assessments.
  • Strong understanding of GPU architectures, virtualization technologies, and bare metal provisioning is a plus.
  • Strong analytical and troubleshooting skills with a customer-centric approach.

Benefits:

  • Opportunities for promotion
  • Professional development

Application Question(s):

  • Are you a local Malaysian citizen?

Experience:

  • High Performance Computing: 3 years (Required)
  • GPU Cluster: 3 years (Required)
  • AI Cluster: 3 years (Required)
  • CPU Cluster: 3 years (Required)

Work Location: In person

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More