Job Summary
As an OpenStack Platform Engineer, you will design, implement, automate, and maintain scalable OpenStack infrastructure. Collaborate with cross-functional teams to ensure platform availability, security, and performance while leveraging AI-assisted automation to enhance operational efficiency.
Responsibilities
Design, implement, and maintain Red Hat OpenStack-based cloud infrastructure solutions to ensure reliability and scalability
Develop and optimize automation scripts for deployment and management of OpenStack environments to improve operational efficiency
Collaborate with cross-functional teams to integrate OpenStack with Kubernetes, OpenShift, and other systems for seamless platform interoperability
Troubleshoot and resolve complex OpenStack and underlying infrastructure issues to maintain platform stability
Implement security best practices and ensure compliance with organizational policies to safeguard infrastructure
Contribute to capacity planning, performance tuning, and scalability enhancements of OpenStack environments
Assist in change management activities including upgrades, hotfixes, and SysAdmin tasks to maintain platform currency
Explore and implement Agentic AI and AIOps solutions to automate OpenStack administration and improve operational workflows
Leverage AI-assisted automation for monitoring, troubleshooting, incident analysis, and remediation of OpenStack environments
Develop and integrate AI agents with OpenStack APIs and automation tools to streamline infrastructure operations
Conduct root cause analysis of issues, review code, and perform unit testing to ensure system reliability
Author and edit system documentation and playbooks to support operational consistency
Mentor less experienced internal and third-party team members on procedural and technical matters
Required competencies and certifications
Expertise in Red Hat or any OpenStack platform administration
Expertise in Kubernetes platform management
Expertise in Red Hat Linux administration
Experience provisioning infrastructure with VMware ESX and monitoring using Prometheus and Ops
Experience integrating with Elastic Stack for logging and monitoring
Familiarity with Generative AI, LLMs, Agentic AI, and AIOps concepts for infrastructure operations
Experience or exposure to AI agents and AI-assisted automation for cloud/OpenStack administration
Strong analytical and problem-solving skills applied to complex infrastructure issues
Excellent oral and written communication skills to maintain clear communication with stakeholders
Proven leadership experience in technical environments
Ability to succeed in fast-paced, high-demand environments
Preferred competencies and qualifications
Understanding of integrating AI/LLM capabilities with APIs, Python, Ansible, or Terraform
Red Hat/OpenStack certifications in Linux and OpenStack administration
Experience with infrastructure-as-code technologies such as Terraform
Understanding of application development lifecycles and experience with CI/CD tools in OpenStack platform lifecycle
Excellent project management and organizational skills
Strong customer service orientation and ability to establish new standards for quality, performance, or productivity