Required DevOps Infrastructure Team Lead
As a DevOps Infrastructure Team Lead, you will:
Lead, mentor and grow a team of DevOps and Infrastructure engineers, providing technical direction and fostering ownership and continuous improvement
Own the core architecture, development and operation of the dev and production infrastructure and self-hosted AI environments, including on-premises GPU infrastructure
Lead Kubernetes-based infrastructure end to end, including system design, development, deployment, operations, troubleshooting and lifecycle management
Define infrastructure platform standards, best practices and engineering methodologies
Own the DevOps toolchain, including DevOps tools and monitoring\logging platforms
Partner closely with engineering teams to enable delivery and solve complex platform and infrastructure challenges
Evaluate and introduce new technologies through POCs and technical assessments
Drive incident resolution, root-cause analysis and continuous improvement of platforms and infrastructure services.
Requirements: If you have:
At least 5 years of hands-on experience in DevOps engineering
At least 2 years of experience in managing, leading or mentoring engineers
Proven ability to design, manage and maintain highly scalable production systems
Hands-on production experience with Kubernetes, including deployment, operations, troubleshooting and platform management
Strong experience with Terraform, Ansible and Helm
Prometheus, Grafana, and ELK or similar monitoring, observability and logging stacks
Strong programming/scripting skills in Python, Go, Ruby, Java, PowerShell or Bash
Experience with GitOps practices and tools such as Argo CD
Strong technical leadership, problem-solving and communication skills
A proactive, independent and hands-on approach, with a passion for learning and tackling complex challenges
It would be great if you also have:
Experience managing on-prem GPU infrastructure for AI/ML environments
Strong knowledge of networking fundamentals
Experience defining technical strategy, architecture and infrastructure roadmaps
Experience with highly available, large-scale or security-sensitive environments.
This position is open to all candidates.