Proactively monitor and support the performance and availability of IT systems, including windows/servers, virtualization platforms, applications and core infrastructure services. Provide first-level support and troubleshooting for detected issues.
Proactively monitor and support network infrastructure including routers, switches, firewalls, SD-WAN, and wireless solutions. Provide first-level support and troubleshooting for detected issues.
Respond to and resolve IT incidents in a timely manner, adhering to established procedures and escalation protocols. Accurately document incidents and resolutions within the incident management system.
...
Operate proactive monitoring, outage detection, and incident response to minimize downtime and operational risk
Support disaster recovery and business continuity processes to safeguard critical AI workloads
Collaborate closely with MLOps, LLMOps, and engineering teams to integrate automation, observability, and security best practices into platform operations
...
Monitor and manage critical Data Centre (DC) facilities infrastructure in a 24x7 Centralized Network Operations Centre (CNOC) environment.
Perform real-time monitoring of electrical, mechanical and environmental systems including UPS, generators, CRAC/CRAH units, power distribution systems, temperature and humidity monitoring, leak detection systems and access control systems.
Respond promptly to facility alarms, incidents and abnormalities, and escalate issues to customer success team, technical specialists and customers based on severity and operational impact.
...
Design, Implement and maintain Openshift Virtualization and Container infrastructure in covering on-premises datacentres. Support also VMware, Nutanix or HyperV platforms
Drive Process, Quality, Productivity and Security improvements within the domain
Works with other technology groups in resolving production issues, including security related issues
...