jobs in HFG Insurance Recruitment

Kerja Sepenuh Masa, Senior Platform Reliability Engineer di HFG Insurance Recruitment Selangor - Maukerja

Senior Platform Reliability Engineer

HFG Insurance Recruitment

Kongsi
Simpan

Lokasi Kerja

  • Cyberjaya Selangor Malaysia

Penerangan Kerja

Tanggungjawab

Key Responsibilities:

  • As a Senior Platform Reliability Engineer, you will play a key role in maintaining the stability, reliability, and efficiency of the organization's internal container platform and its supporting infrastructure. Your responsibilities will include core operational tasks such as resource provisioning and management, responding to platform and application outages, and monitoring.
  • This includes proactively identifying and resolving reliability issues, analysing product dependencies, pinpointing performance bottlenecks, and implementing optimization strategies to enhance platform availability and cost efficiency.
  • In this role, you will participate in a 24/7 on-call rotation, promptly addressing alerts from the global monitoring team and resolving production incidents to maintain platform and application uptime. Additionally, you will regularly review team workflows to identify manual processes and implement automation solutions that reduce effort and minimize human error.
  • Regularly deploy product updates as required to keep the platform vulnerability-free.
  • Work with open-source technologies, CI/CD, SCM tools as necessary, and source control such as Bitbucket, implement organization containers (e.g., Docker and Kubernetes). Stay current with industry trends and propose new ways for the business to improve.
  • Take accountability in considering business and regulatory compliance risks and take appropriate steps to mitigate the risks.
  • Maintain awareness of industry trends on regulatory compliance, emerging threats and technologies to understand the risk and better safeguard the company.
  • Highlight any potential concerns/risks and proactively share best risk management practices.

We are looking for people with:

  • Bachelor's or Master's degree in Computer Science or a related field.
  • Minimum of 5 to 7 years of overall experience in IT, with at least 3 to 5 years of hands-on experience as a Platform Reliability Engineer or Site Reliability Engineer, specifically managing container orchestration platforms such as Tanzu Application Service, Tanzu Kubernetes Grid Integrated Edition, or other Kubernetes-based platforms.
  • At least 3 years of experience in automation using tools like Ansible and scripting languages such as Python and Bash.
  • 3 years of experience in developing and maintaining Helm charts and Helm repositories.
  • Minimum of 3 years of experience managing NSX-T solutions and integrating them with Tanzu suite products.
  • Possession of one or more of the following certifications:
  • a) Certified Kubernetes Administrator (CKA)
  • b) Certified Kubernetes Application Developer (CKAD)
  • c) Certified Kubernetes Security Specialist (CKS)
  • 3–5 years of experience working in high-demand, fast-paced environments.
  • Strong expertise in platform reliability principles, including scalability, performance optimization, and enterprise platform architecture.
  • Proficiency in designing monitoring dashboards using Grafana and Dynatrace to track SLOs, SLIs, and SLAs of the platform.
  • Solid understanding of DevOps pipelines and automation tools such as Bamboo, Ansible, Bitbucket, Nexus, Jira and Confluence.
  • Strong technical and business acumen with the ability to collaborate across multiple technical teams.
  • Proven experience in diagnosing and resolving infrastructure and networking issues.
  • Extensive experience in CI/CD environments, with a deep understanding of change and version control processes.
  • Hands-on experience with platform upgrades, patching, and buildpack management.
  • Ability to troubleshoot complex network-related problems.
  • Passion for continuous learning and evaluating emerging technologies, with a commitment to knowledge sharing within the team.
  • Ability to document Standard Operating Procedures (SOPs) and contribute to internal knowledge bases.
  • Strong collaboration skills with the ability to work across various stakeholder groups at an organizational level.
  • Excellent communication skills to engage with stakeholders and domain experts in designing and operating enterprise-wide solutions.
  • Self-motivated, disciplined, and proactive with a strong sense of ownership and urgency.
  • High level of integrity, takes accountability of work, and maintains a good attitude toward teamwork.


Peringatan Penting

Jangan pernah kongsikan maklumat bank atau kad kredit anda semasa memohon pekerjaan. Elakkan membuat sebarang pembayaran atau mengisi survey yang tidak berkaitan. Jika ada yang mencurigakan, sila laporkan iklan pekerjaan ini segera.

Lebih Lanjut