jobs in Techstreet Malaysia

Kerja Sepenuh Masa, Senior AI Network - Security Engineer di Techstreet Malaysia Johor - Maukerja

Senior AI Network - Security Engineer

Techstreet Malaysia

Kongsi
Simpan

Lokasi Kerja

  • Johor Bahru Johor Malaysia

Penerangan Kerja

Tanggungjawab

Responsible for the design, implementation, operation and continuous improvement of network and security infrastructure supporting large-scale AI/GPU computing platforms, data centres and cloud services.


The role combines strong data-center networking and security fundamentals with high-performance AI networking technologies, including NVIDIA InfiniBand, Spectrum Ethernet/Spectrum-X and RoCEv2.


The candidate must demonstrate strong networking fundamentals, hands-on troubleshooting capability and the ability to rapidly learn and develop expertise in AI/GPU networking.


The role will also provide technical leadership during complex incidents, infrastructure deployments and network expansion projects, while mentoring other engineers and working closely with Operations, Systems, Platform, Security, Data Centre teams and external technology partners. The role requires participation in on-call support as needed.


Responsibilities

  • AI & Data Center Network Infrastructure
  • Design, implement, operate and maintain highly available network infrastructure supporting AI/GPU clusters, data centres and AI Cloud services.
  • Support high-performance GPU network fabrics including NVIDIA InfiniBand and Ethernet/RoCE-based architectures.
  • Design and operate data-center networking technologies including routing, switching, VLANs, BGP, EVPN/VXLAN and high-availability architectures.
  • Configure and support network infrastructure across compute, storage, management, out-of-band (OOB), customer and external connectivity networks.
  • Support LAN, WAN, VPN and private interconnect connectivity between data centres, customers, partners and cloud environments.
  • Participate in network architecture, capacity planning and infrastructure expansion activities for new AI/GPU clusters.
  • Network Operations & Performance
  • Monitor network and AI fabric availability, throughput, latency, utilisation, errors and congestion to ensure infrastructure meets performance and SLA requirements.
  • Troubleshoot complex connectivity and performance issues across switches, ConnectX NICs, DPUs and SuperNICs, servers, host networking and GPU workloads.
  • Work with Systems, Platform and Operations teams to identify network-related issues impacting distributed GPU workloads.
  • Develop expertise in AI networking technologies including InfiniBand, RDMA, RoCEv2, lossless Ethernet, PFC, ECN, QoS and congestion management.
  • Support AI fabric monitoring and management platforms such as NVIDIA UFM, NetQ or equivalent tools.
  • Support network validation, commissioning and performance testing for new GPU clusters and infrastructure deployments.
  • Network Security
  • Design, implement and operate network security infrastructure including firewalls, VPNs, ACLs, segmentation, NAT, IPS and secure connectivity.
  • Configure and manage Fortinet/FortiGate or equivalent enterprise firewall platforms.
  • Implement appropriate network segmentation across compute, management, storage, OOB, customer and external-facing environments.
  • Support security hardening, vulnerability remediation and compliance requirements.
  • Work closely with Security and Risk teams during security incidents, assessments and audits.
  • Network Automation
  • Develop and maintain network automation using technologies such as Python, Ansible, APIs, DCIM and Infrastructure as Code.
  • Automate configuration deployment, backups, compliance validation, provisioning and routine operational activities.
  • Support CI/CD and controlled network change processes to improve consistency, reliability and auditability.
  • Work with platform teams to integrate network and fabric telemetry into monitoring platforms.
  • Operations & Reliability
  • Provide technical leadership for complex network incidents, outages and performance degradation.
  • Perform root-cause analysis and drive corrective and preventive actions.
  • Analyse network performance trends, capacity, hardware utilisation and growth requirements.
  • Plan and test redundancy, failover and recovery mechanisms.
  • Participate in change management, maintenance, upgrades and lifecycle management activities.
  • Participate in the operational standby/on-call roster supporting 24×7 AI Cloud services.
  • Develop and maintain HLD/LLD, network diagrams, SOPs, MOPs, EOPs, troubleshooting runbooks and technical documentation.


Requirements

  • Bachelor's degree in Network Engineering, Computer Science, Information Technology or related discipline, or equivalent practical experience.
  • 8+ years of relevant network engineering experience, preferably within data-center, cloud, service-provider or large-scale infrastructure environments.
  • Strong hands-on knowledge of TCP/IP, routing, switching, VLANs, BGP and network redundancy/high availability.
  • Strong experience designing, implementing and troubleshooting production network infrastructure.
  • Hands-on experience with enterprise/data-center switching and routing platforms such as Juniper, Cisco, Arista or equivalent.
  • Experience with enterprise firewall technologies, network segmentation, ACLs, VPNs and security policies.
  • Solid understanding of advanced networking technologies, particularly those related to AI would be highly advantageous.
  • Strong communication skills, both written and verbal.
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Willingness to work site-based in Johor and to participate in a 24×7 escalation roster, if required.
  • Hands-on experience with NVIDIA InfiniBand, Spectrum Ethernet Platform, and/or RDMA over Converged Ethernet (RoCE) preferred.
  • Good understanding of Linux and host networking.
  • Enterprise & AI network delivery (LAN, WAN, WLAN, VPN)
  • Secure connectivity (site-to-site VPN, MPLS, private interconnects)
  • Network automation & NetDevOps (IaC, Ansible, Python, CI/CD)
  • Network security (firewalls, VLANs, ACLs, zero-trust)
  • Operations, monitoring & incident response
  • AI networking tech (InfiniBand, Spectrum, RoCE)
  • Linux networking fundamentals
  • Cross-team collaboration & mentoring
  • Certifications such as NVIDIA-Certified InfiniBand or Networking Professional, JNCIP or JNCIE, CCNP or CCIE, Fortinet NSE 4 and above, CISSP or CISM.


Required Skills

PythonCloud (AWS / Azure / GCP)DevOps (Docker / Kubernetes / CI-CD)


Benefits

  • Health Insurance
  • Performance Bonus
  • Dental Coverage


Peringatan Penting

Jangan pernah kongsikan maklumat bank atau kad kredit anda semasa memohon pekerjaan. Elakkan membuat sebarang pembayaran atau mengisi survey yang tidak berkaitan. Jika ada yang mencurigakan, sila laporkan iklan pekerjaan ini segera.

Lebih Lanjut