Required Senior Principal DevOps Engineer (Cortex Cloud)
Your Impact:
As a Senior Principal DevOps Engineer, you will serve as a visionary technical leader within the Cortex Cloud DevOps group. You will define the technical strategy and architecture that ensures our massive-scale production services remain highly reliable, exceptionally secure, and performant. You will pioneer the integration of AI-driven capabilities into our daily operations, establishing elite engineering standards and fundamentally transforming the workflows of hundreds of developers through autonomous agents and intelligent procedures.
Your Career:
Architectural Vision & Scalability: Design and scale massive, resilient distributed systems and global Kubernetes infrastructure, implementing robust observability and monitoring frameworks.
AI-Driven Transformation: Revolutionize the SDLC by integrating Generative AI, autonomous agents, and LLM-powered workflows into CI/CD and self-healing systems to accelerate developer velocity.
Technical Leadership & IaC: Define architectural standards, lead GitOps/IaC (FluxCD/Terraform) strategies, and mentor Senior/Staff engineers across the R&D organization.
Developer Experience & Efficiency: Build and champion AI-powered platforms that automate troubleshooting and eliminate friction. Partner directly with Engineering Directors, Principal Architects, and Product Management to align infrastructure initiatives with business goals, optimizing for scale, high availability, and multi-million-dollar cost-efficiencies.
Security & Compliance: Embed "Security by Design" principles into the platform architecture to ensure platform integrity without sacrificing delivery speed.
Requirements: Your Experience:
10+ years of progressive experience in DevOps, SRE, Platform, or Infrastructure Engineering roles, with a significant portion at the principal/ tech leadership/ staff, or architectural level.
System Design from Scratch: A proven track record of designing, building, and deploying large-scale, highly available distributed systems and cloud platforms from the ground up.
AI-Powered Automation: Proven experience designing and integrating AI-driven systems, autonomous agents, and LLM-based tools into engineering workflows to optimize development processes, procedures, and overall organizational efficiency.
Communication: Exceptional interpersonal skills, capable of articulating complex architectural and AI workflow concepts clearly to both deeply technical peers and executive leadership.
Cloud & IaC Mastery: Expert-level proficiency with GCP (or equivalent major cloud providers) and deep architectural experience with Terraform.
Advanced Container Orchestration: Deep, internal knowledge of virtualized and containerized environments, with architectural-level expertise in scaling Kubernetes, extending it via custom operators, and automating complex operational logic.
Software Engineering Approach: Advanced coding and automation skills in Python or Go. You treat infrastructure as a software engineering discipline and can build custom tooling/services when off-the-shelf solutions fall short.
Proven Leadership: Demonstrated ability to lead complex, cross-team technical initiatives from conception to delivery, including setting technical roadmaps and driving consensus among stakeholders.
OS/Systems Expertise: Mastery of Linux systems, including kernel tuning, advanced networking, and performance troubleshooting.
Nice to Have:
Deep expertise in managing and scaling stateful workloads and distributed databases (e.g., Cassandra, ScyllaDB, MemSQL, or MySQL) in containerized environments.
Experience contributing to open-source infrastructure projects, or presenting at major tech conferences.
This position is open to all candidates.