We are looking for an experienced and hands-on, Vice President of Site Reliability Engineering to lead the reliability, availability, and resilience strategy across critical platforms and services. This individual will establish SRE as a core engineering discipline, driving automation, observability, incident management, and reliability-by-design practices across the organization.
Key Responsibilities
- Define and execute the enterprise SRE strategy, embedding reliability engineering across products, platforms, and services.
- Own and govern SLIs, SLOs, and error budgets, ensuring reliability decisions are data-driven.
- Drive improvements in service availability, resilience, recoverability, and performance.
- Lead the strategy for observability, automation, self-healing capabilities, and resilience engineering platforms.
- Oversee incident management, major incident response, post-incident reviews, and chaos engineering initiatives.
- Drive operational excellence through automation, toil reduction, and platform standardization.
- Partner with Engineering, Infrastructure, Security, and Product teams to embed reliability requirements early in the development lifecycle.
- Build, mentor, and lead a high-performing team of SRE leaders and engineers.
- Influence technical decisions across teams and stakeholders, driving reliability outcomes in a matrixed environment.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
- At least 12 years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Infrastructure, DevOps, or related disciplines.
- Proven experience building, scaling, or transforming SRE functions within complex enterprise environments.
- Strong expertise in:
- SRE principles and automation-first operations
- Observability and monitoring frameworks
- Resilience engineering and incident management
- Capacity planning, performance optimization, and disaster recovery
- Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budget management
- Solid software engineering background with experience in Java and developing or supporting large-scale distributed systems.
- Experience designing and implementing observability, automation, self-healing, or reliability platforms.
- Strong stakeholder management skills with the ability to influence engineering teams and senior leaders without direct authority.