L2/L3 Major Incident & Service Operations Manager
Location: Kuala Lumpur, Malaysia
Employment Type: Contract – 12 Months
Contract Duration: 12 Months, Extendable Annually
Experience: 8+ Years
Work Arrangement: Onsite / Client Location
Job Overview
We are seeking an experienced L2/L3 Major Incident & Service Operations professional to join our team and support one of our key clients in Kuala Lumpur.
The role is responsible for managing and coordinating L2/L3 technical incidents, with a strong focus on Major Incident Management, service restoration, escalation management, SLA adherence, root cause analysis, and stakeholder communication.
The ideal candidate should have strong experience working in enterprise IT environments and be capable of coordinating multiple technical teams during high-priority incidents while maintaining clear communication with business and technical stakeholders.
This is a senior operational role requiring strong ITIL, Incident Management, problem-solving, communication, and technical coordination skills.
Key ResponsibilitiesMajor Incident Management
- Own and manage P1/P2 and other high-priority incidents from identification through service restoration and closure.
- Coordinate L2/L3 technical teams during critical incidents and ensure timely resolution.
- Initiate and lead Major Incident / War Room calls when required.
- Establish clear incident ownership, technical action plans, timelines, and escalation paths.
- Ensure appropriate technical teams and vendors are engaged based on the nature of the incident.
- Drive incidents toward service restoration within agreed SLA and business expectations.
- Maintain accurate incident timelines, actions, decisions, and communications.
- Provide regular updates to technical and business stakeholders during major incidents.
- Ensure proper handover and follow-up after service restoration.
L2/L3 Incident Coordination
- Coordinate with infrastructure, network, server, database, application, cloud, security, and other technical teams.
- Analyse incident information and determine the appropriate technical escalation path.
- Challenge delays and drive technical teams toward timely resolution.
- Ensure escalated incidents receive appropriate technical attention and ownership.
- Monitor ageing incidents and proactively escalate potential SLA breaches.
- Support L2/L3 teams in identifying recurring or complex technical issues.
Incident Investigation & RCA
- Coordinate technical investigations for major and recurring incidents.
- Ensure appropriate Root Cause Analysis (RCA) is completed following major incidents.
- Review RCA reports for completeness, accuracy, and corrective actions.
- Track corrective and preventive actions until completion.
- Identify recurring incidents and work with Problem Management teams to eliminate underlying causes.
- Recommend improvements to prevent repeat incidents and improve service stability.
ITIL & Service Management
- Follow established ITIL-based Incident, Major Incident, Problem, Change, and Service Management processes.
- Ensure incidents are properly categorised, prioritised, assigned, documented, and closed.
- Ensure incidents meet defined SLA, OLA, and operational requirements.
- Identify process gaps and recommend improvements.
- Support continual service improvement initiatives.
Stakeholder & Communication Management
- Act as the key coordination point during major technical incidents.
- Provide clear and timely communications to business stakeholders, management, technical teams, and vendors.
- Prepare incident updates, executive summaries, and post-incident reports.
- Communicate technical issues in a clear and business-friendly manner.
- Maintain professional communication during high-pressure and critical situations.
Vendor & Third-Party Coordination
- Coordinate with external vendors and managed service providers during major incidents.
- Ensure vendors provide timely technical investigation, updates, and resolution.
- Escalate vendor issues when SLA or resolution timelines are at risk.
- Track vendor-related actions through to completion.
- Participate in vendor service reviews and operational improvement discussions.
Service Monitoring & Operational Excellence
- Monitor incident trends, recurring issues, SLA performance, and service availability.
- Identify operational risks and proactively escalate potential service-impacting issues.
- Analyse incident data and recommend improvements to service reliability.
- Support service review meetings and operational performance reporting.
- Contribute to improving incident response processes and operational procedures.
Reporting & Documentation
- Prepare regular incident and service management reports.
- Track key metrics such as:
- P1/P2 incident volumes
- MTTR
- SLA compliance
- Incident ageing
- Recurring incidents
- Major incident trends
- RCA completion
- Corrective action closure
- Maintain accurate incident records and operational documentation.
- Prepare post-incident reviews and management summaries.
Technical Understanding
The candidate does not need to be a hands-on specialist across every infrastructure technology, but must have sufficient technical understanding to effectively coordinate L2/L3 teams.
Good understanding of areas such as:
- Windows / Linux environments
- Servers and enterprise infrastructure
- Networking – TCP/IP, DNS, DHCP, VPN, routing, switching
- Databases
- Cloud environments – AWS / Azure / GCP
- Virtualisation technologies
- Storage and backup
- Enterprise applications
- Cybersecurity and infrastructure services
- Monitoring and observability tools
Required Skills & Experience
- 8+ years of experience in IT Service Management, Incident Management, Infrastructure Operations, Technical Support, or a related IT operations role.
- Strong experience in L2/L3 Incident Management or Major Incident Management.
- Proven experience managing P1/P2 critical incidents in an enterprise environment.
- Strong experience coordinating multiple L2/L3 technical teams.
- Hands-on experience with ITIL-based Incident Management processes.
- Strong knowledge of SLA, OLA, escalation, prioritisation, and service restoration processes.
- Experience conducting or coordinating RCA and post-incident reviews.
- Strong stakeholder and vendor management experience.
- Experience leading conference bridges / war rooms during critical incidents.
- Strong understanding of enterprise infrastructure and IT operations.
- Excellent communication and coordination skills.
- Ability to remain calm and structured during high-severity incidents.
- Strong analytical, problem-solving, and decision-making skills.
Preferred Skills
Experience with the following will be an advantage:
- ServiceNow or other enterprise ITSM platforms
- ITIL Foundation / ITIL 4 certification
- Major Incident Management certification or training
- Problem Management
- Change Management
- Service Management reporting and KPI dashboards
- Enterprise monitoring tools
- Managed Services / IT Operations environments
- 24x7 production support environments
Candidate Profile
We are looking for someone who can take ownership of critical incidents and drive teams toward resolution, rather than simply forwarding tickets between support groups.
The successful candidate should be:
- Strong in Major Incident Management
- Technically aware and comfortable working with L2/L3 engineers
- Excellent at stakeholder communication
- Strong in escalation and coordination
- Experienced in managing high-pressure production incidents
- Process-oriented with good ITIL knowledge
- Comfortable working with senior management and technical teams
- Proactive in identifying recurring issues and service improvements
Contract Details
Position: L2/L3 Major Incident & Service Operations Manager
Location: Kuala Lumpur, Malaysia
Contract: 12 Months
Extension: Extendable annually based on performance and business requirements
Experience: 8+ Years
Work Arrangement: Onsite / Client Location
Interested candidates are invited to apply with an updated CV highlighting relevant Major Incident Management, L2/L3 support, ITIL, ServiceNow/ITSM, RCA, SLA management, stakeholder management, and enterprise IT operations experience.
Please mention your current location, total years of experience, notice period, and relevant ITSM/technical experience in your application.
Pay: RM6,000.00 - RM12,000.00 per month
Work Location: In person