Responsibilities
Data Centre Operations & Service Delivery
- Take overall responsibility for the day-to-day operations and performance of the data centre.
- Ensure high availability, reliability, security, and continuity of data centre services.
- Oversee incident management, service recovery, and root cause analysis for major operational issues.
- Establish and maintain operational procedures, escalation processes, and emergency response plans.
- Monitor data centre performance and operational KPIs, and drive timely corrective actions.
Infrastructure & Capacity Management
- Oversee the availability and performance of critical data centre infrastructure, including power, cooling, UPS, generators, fire protection, and monitoring systems.
- Ensure preventive and corrective maintenance activities are properly planned and executed with minimal impact to operations.
- Oversee the lifecycle of data centre assets, including installation, upgrades, relocation, and decommissioning.
- Monitor capacity across power, cooling, space, and rack utilisation to support business growth and future expansion.
- Identify infrastructure risks and ensure appropriate mitigation plans are in place.
People & Team Management
- Lead, manage, and develop the data centre operations team to achieve service and operational objectives.
- Plan manpower requirements, shift coverage, resource allocation, and succession planning.
- Set team objectives and performance expectations and conduct regular performance reviews.
- Ensure staff are adequately trained on operational procedures, safety requirements, and emergency response.
- Build a culture of accountability, safety, operational discipline, and continuous improvement.
Vendor & Stakeholder Management
- Manage key vendors, contractors, and service providers to ensure performance, service quality, and SLA compliance.
- Lead vendor performance reviews, contract-related operational matters, and service escalations.
- Work closely with IT, Network, Security, Facilities, and other business stakeholders to support operational and project requirements.
- Ensure all third-party personnel comply with site access, security, safety, and operational requirements.
- Provide regular updates to management on operational performance, risks, incidents, and improvement initiatives.
Governance, Risk & Continuous Improvement
- Ensure data centre operations comply with company policies, regulatory requirements, ISO standards, and relevant industry best practices.
- Lead operational audits, risk assessments, business continuity, and disaster recovery activities.
- Ensure BCP and DR plans are regularly reviewed, tested, and updated.
- Review operational processes and identify opportunities to improve reliability, efficiency, service quality, and cost effectiveness.
- Prepare and report key operational metrics, risks, incidents, and improvement plans to management.
Requirements
- Bachelor's Degree in Electrical Engineering, Mechanical Engineering, Facilities Management, or a related discipline.
- 5–10 years of experience in Data Centre Operations, Critical Facilities, Infrastructure Management, or a similar environment.
- Proven experience managing a data centre or other mission-critical facility with high availability requirements.
- Strong understanding of data centre infrastructure, including electrical and mechanical systems, UPS, generators, cooling systems, fire protection, DCIM, and BMS.
- Experience in managing data centre operations, maintenance, capacity planning, incident management, risk management, and business continuity.
- Proven experience in people management, including leading technical teams and managing manpower and performance.
- Strong experience in vendor and stakeholder management, including managing contractors and service providers.
- Strong leadership, decision-making, problem-solving, communication, and organisational skills.
- Ability to communicate effectively with both technical and non-technical stakeholders.