Oversee day-to-day operations across all data centers in the cluster, ensuring all systems and infrastructure maintain 100% uptime and full compliance with customer SLAs.
Manage incident response, troubleshooting, and resolution — minimizing downtime and service disruption across the cluster.
Track operational metrics, analyze performance data, and identify risks before they affect operations or client commitments.
...
Direct and oversee 24/7 site operations across the regional facility cluster, ensuring maximum infrastructure availability (100% uptime) and full compliance with client SLAs.
Enforce, update, and optimize operational SOPs, EOPs (Emergency Operating Procedures), and industry best practices.
Lead incident management, root-cause analysis (RCA), and emergency response efforts to minimize downtime and mitigate operational risk.
...