Job Summary:
We are looking for a Cloud Infrastructure Engineer to support the stability, availability, and performance of IT infrastructure across cloud and server environments. The role will focus on infrastructure monitoring, system administration, performance testing, patch validation, troubleshooting, and operational support.
The successful candidate will work with technology teams to maintain reliable and secure infrastructure and ensure systems remain operational and ready for business use.
Key Responsibilities:
- Install, configure, maintain, and support Windows and/or Linux servers.
- Monitor system performance, logs, alerts, and infrastructure health.
- Monitor cloud workloads, services, and resource utilization across AWS and Azure environments.
- Perform system performance testing, validation, and health checks.
- Conduct pre- and post-patch validation to ensure system stability.
- Support User Acceptance Testing (UAT) for infrastructure and system changes.
- Troubleshoot server, network, system, and cloud-related issues and escalate complex incidents when required.
- Support backup, restoration, and disaster recovery testing activities.
- Maintain documentation for system configurations, changes, incidents, and testing activities.
- Follow incident management and escalation procedures to ensure timely resolution of critical issues.
- Participate in infrastructure and operational readiness activities.
- Provide support during critical incidents, outages, weekends, and public holidays when required.
- Participate in 24/7 standby/on-call support for critical infrastructure incidents.
Qualifications & Requirements:
- Bachelor’s degree in Information Technology, Computer Science, Engineering, or a related discipline.
- Hands-on experience in either Windows Server or Linux Server administration is mandatory.
- AWS certification is mandatory, such as AWS Certified Cloud Practitioner, Solutions Architect, SysOps Administrator, or equivalent.
- Practical experience with AWS cloud infrastructure.
- Knowledge or experience with Microsoft Azure is an advantage.
- Understanding of infrastructure monitoring, system performance, patching, backup, and disaster recovery.
- Strong troubleshooting and analytical skills.
- Good understanding of IT infrastructure and cloud operations.
- Strong communication and documentation skills.
- Ability to work effectively in a fast-paced operational environment.
- Willingness to learn new technologies and take on new challenges.
- Willingness to work on standby and provide support outside normal working hours, including weekends and public holidays when required.
Core Skills:
Cloud: AWS, Azure
Operating Systems: Windows Server, Linux
Infrastructure: Server Administration, Monitoring, Performance Management, Patching, Backup & Recovery
Operations: Incident Management, Troubleshooting, UAT, System Validation, Disaster Recovery Testing
Work Requirements:
This position involves operational support for critical IT infrastructure and requires flexibility to participate in 24/7 standby/on-call support, including weekends and public holidays when necessary.