Our client is a leading global financial markets and technology organization that operates high-availability infrastructure supporting customers around the world.
They are looking for a Senior Linux Platform Operations Engineer to provide hands-on operational ownership of large-scale Linux environments. This is a highly operational role focused on production support, incident management, maintenance execution, platform reliability, and continuous improvement. The successful candidate will act as a senior escalation point for Linux infrastructure issues and play a key role in ensuring the stability and availability of mission-critical systems.
Key Responsibilities
- Own the operational health, availability, and performance of Linux infrastructure across production and non-production environments.
- Lead and execute weekend maintenance activities, system upgrades, patching, failover testing, and recovery exercises.
- Act as the senior technical escalation point for Linux-related incidents, driving triage, resolution, and root cause analysis.
- Administer and maintain bare-metal Linux servers, storage systems, and configuration management platforms.
- Execute and monitor production changes in accordance with established change management processes.
- Develop and maintain operational automation using Python, Shell scripting, and infrastructure management tools.
- Support monitoring, observability, and capacity management through dashboards, reporting, and system analysis.
- Create and maintain operational procedures, runbooks, and recovery documentation.
- Collaborate with Engineering, SRE, Security, and Infrastructure teams to improve platform reliability and operational efficiency.
Requirements & Qualifications
- At least 8 years of Linux systems administration and infrastructure operations experience in large-scale, 24x7 environments.
- Deep expertise in Linux OS administration, performance tuning, troubleshooting, and kernel fundamentals.
- Strong experience with configuration management tools such as Salt, Puppet, or Ansible.
- Proficiency in Python and Shell scripting for automation and operational tooling.
- Experience supporting bare-metal server environments and enterprise storage technologies (SAN, NAS, RAID, NVMe).
- Hands-on experience with monitoring and observability platforms such as Grafana and Prometheus.
- Familiarity with cloud and container technologies including AWS, GCP, Docker, and Kubernetes.
- Experience with Infrastructure as Code and DevOps tooling, including Terraform, Git, and CI/CD pipelines.
- Strong problem-solving skills with the ability to perform effectively during high-severity incidents.
Work Schedule
- Weekend-focused operational coverage model.
- Early morning starts will be required to support global operations (7AM SGT)