Job Description
IBM Cloud Management Support Engineer
Position Title
IBM Cloud Management Support Engineer
Department
Managed Services / Cloud Operations
Reporting To
Technical Consultant
Position Summary
The IBM Cloud Management Support Engineer is responsible for the administration, monitoring, support, maintenance, and optimization of customer IBM Cloud environments. The role encompasses operational support across infrastructure, networking, security, and platform services, ensuring service availability, performance, security, and compliance with agreed Service Level Agreements (SLAs).
The Engineer will perform incident management, service request fulfilment, technical troubleshooting, cloud administration, change implementation, and continuous service improvement while collaborating with customers, vendors, IBM Support, and internal technical teams.
Key Responsibilities
1. Incident Monitoring and First-Level Analysis
The Engineer shall:
- Continuously monitor IBM Cloud production environments.
- Monitor service health, availability, and infrastructure performance using IBM Cloud monitoring tools.
- Detect and assess infrastructure-related incidents.
- Perform routine health verification activities.
- Identify and investigate abnormalities including:
- Virtual Server (VSI) unavailability
- High CPU, memory, or disk utilization
- Network latency or connectivity degradation
- Storage utilization and capacity issues
- Backup or recovery failures
- Kubernetes/OpenShift cluster alerts
- Security and access anomalies
- Other IBM Cloud infrastructure-related alerts
2. Initial Troubleshooting
The Engineer shall perform first-level diagnostics and troubleshooting before escalation, including:
- Verifying resource availability and cloud service health.
- Performing connectivity checks and network diagnostics.
- Reviewing monitoring alerts and event logs.
- Collecting incident-related information from customers.
- Determining whether issues are infrastructure-related, platform-related, or application-related.
- Documenting findings and troubleshooting results.
- Executing predefined operational runbooks and standard operating procedures.
Exclusions:
- Application debugging.
- Source code troubleshooting.
- Database query optimization.
- Application performance tuning unless specifically stated in the contract.
3. Incident Escalation and Coordination
The Engineer shall:
- Coordinate and manage incident resolution activities across internal teams, vendors, and IBM Support.
- Raise and manage support cases with IBM Support where required.
- Engage appropriate technical resources to assist in the investigation and resolution of complex incidents.
- Track incident progress through to resolution and closure.
- Provide timely updates and communication to stakeholders throughout the incident lifecycle.
- Ensure accurate documentation of incident activities, findings, and resolutions within the ticketing system.
- Participate in major incident management and service restoration activities.
- Follow established incident management and escalation procedures to ensure timely resolution.
4. Operational Service Requests
Support operational activities including:
- User access verification and administration requests
- IAM access requests and modifications
- Basic virtual server operational requests
- Monitoring configuration updates
- Basic storage administration requests
- Backup verification activities
- Approved network configuration requests
- Other approved IBM Cloud administrative requests
5. Communication and Service Management
The Engineer shall:
- Provide incident and request status updates.
- Notify designated customer contacts of major service disruptions.
- Maintain detailed incident records and ticket updates.
- Participate in Root Cause Analysis (RCA) discussions when required.
- Recommend improvements to reduce recurring incidents.
- Support audit and compliance reporting activities.
6. Service Level Requirements
Support coverage may include:
Business Hours Support
- Monday to Friday
- 8:00 AM – 6:00 PM
Extended Support (Where Applicable)
- After-hours support via On-Call arrangement
- Weekend and Public Holiday standby support
Service Objectives
- Adherence to SLA response and resolution targets.
- Accurate and timely incident logging and updates.
- Compliance with customer operational procedures.
Required Qualifications
Education
Bachelor's Degree in:
- Information Technology
- Computer Science
- Information Systems
- Engineering
- Related disciplines
Required Certifications
Candidates should possess at least one (1) or more relevant certifications:
IBM Cloud Certifications
- IBM Certified Professional Cloud Architect
- IBM Certified Solution Advisor - Cloud
- IBM Cloud Professional Administrator
- IBM Cloud Technical Advocate
Infrastructure Certifications
- Red Hat Certified System Administrator (RHCSA)
- Red Hat Certified Engineer (RHCE)
- Certified Kubernetes Administrator (CKA)
IT Service Management
- ITIL Foundation Certification
Required Technical Skills
- IBM Cloud Infrastructure (VPC & Classic)
- Linux Administration (RHEL, CentOS, Ubuntu)
- Windows Server Administration
- IBM Cloud Networking
- VPN Management
- Load Balancer Administration
- DNS Administration
- Identity and Access Management (IAM)
- Backup and Recovery Solutions
- Monitoring and Alerting Platforms
- Basic Shell / PowerShell Scripting
- Experience Requirements
- Junior Support Engineer
- 1-3 years of IT Infrastructure or Cloud Operations experience.
- Senior Support Engineer
- 3-5 years of IBM Cloud, Data Center, or Managed Services experience.
- Experience handling enterprise incidents and operational support environments.
- Key Competencies
- Strong analytical and troubleshooting skills
- Customer service orientation
- Effective verbal and written communication
- Incident management and prioritization skills
- Ability to work under pressure during major incidents
- Team collaboration and escalation management
- Documentation and reporting discipline
- Continuous learning mindset
- Deliverables / Expected Outcomes
- Timely incident detection and response.
- Accurate ticket documentation and updates.
- SLA compliance achievement.
- Proper escalation and coordination management.
- High service availability and operational stability.
- Positive customer satisfaction outcomes.
Location
Klang, Selangor
Onsite