Required DevOps Team Leader - Infra and DevEx
The role:
You will lead the DevOps Infrastructure & Developer Experience team - the team that owns the foundational layer every engineering team builds on. EKS clusters, VPCs, IAM, security boundaries, monitoring, CI/CD, and the developer workflows that tie it all together. When your infrastructure works well, every team ships faster. When it doesnt, everyone feels it.
Youll manage a team of DevOps engineers, set technical direction, and own the infrastructure platform end-to-end - from networking and compute through security and observability to the developer experience layer on top. You are accountable for the reliability, security, and usability of the shared platform, and you measure success by the outcomes it enables across the organization.
This role sits at the intersection of infrastructure depth and organizational breadth. Youll partner with Development, Platform, AI/ML, Security, and Product - ensuring the underlying systems evolve to meet their needs while maintaining the standards and guardrails that keep production safe.
Were looking for someone who has already done this - operated production infrastructure at scale, shipped developer tooling, driven AI adoption in engineering workflows, and measured the impact of
The day-to-day
Team Leadership:
Lead and grow a team of DevOps engineers - set direction, remove blockers, own the roadmap
Balance operational needs with platform investment. Represent infrastructure tradeoffs in cross-org planning
Drive hiring, onboarding, and professional growth
Infrastructure Ownership:
Own the shared layer all teams depend on - EKS, VPC, IAM, networking, security, compute (including GPU), monitoring, data services
Design and operate highly available distributed systems at scale on AWS and GCP. IaC with Terraform, multi-account, multi-region
Own security posture (IAM, network boundaries, secrets, vulnerability scanning) and observability (Prometheus, Grafana, CloudWatch, alerting).
Requirements: 4+ years in DevOps / Platform / SRE roles, managing a team, operating in a SaaS production environment with multi-region deployment within an R&D organization
Deep hands-on expertise with AWS in production - EKS, VPC, IAM, EC2, S3, CloudFront, MSK, Lambda, CloudWatch, Security Hub. Youve built and owned multi-account infrastructure, not just used it
Experience with GCP in production - GKE, Cloud Run, IAM, VPC, Cloud Build, or similar. Comfortable operating across cloud providers
Deep expertise with Kubernetes - cluster operations, networking (CNI, service mesh), RBAC, node lifecycle, Helm chart architecture. You understand the internals, not just the YAML
Strong networking and OS fundamentals - TCP/IP, DNS, load balancing, Linux internals, troubleshooting at every layer
Proven experience owning developer experience - youve built internal platforms, CI/CD systems, or self-service tooling that developers actually adopted and that measurably improved their velocity
Strong security mindset - IAM design, network segmentation. Security is part of how you build, not an afterthought
Experience with CI/CD at scale (GitHub Actions), monitoring and observability (Prometheus, Grafana), and Infrastructure as Code (Terraform)
High proficiency with AI coding tools - you use them daily, understand how they work, and can drive their adoption across a team
These would also be nice:
Experience building or operating GPU infrastructure for AI/ML inference at scale
Experience with AI agents - building them, deploying them in production, or enabling teams to use them
Experience consolidating or migrating infrastructure across acquisitions or organizational changes
Experience defining and reporting on engineering metrics (DORA, platform KPIs, cost models)
Experience with multi-cloud (AWS + GCP) infrastructure.
This position is open to all candidates.