As a Senior DevOps Engineer, you will join our data & engineering team and design and manage scalable, cloud-based infrastructure focused on big data and machine learning. You will ensure high-scale production environments are reliable, secure, and efficient, supporting seamless ML/AI deployments.
Location: Ramat Gan, Israel (hybrid model)
Reporting to: DevOps Team Lead
What will you do?
Design, build, and maintain scalable and reliable infrastructure for our big data applications.
Automate the deployment and scaling of our applications using Kubernetes and related technologies.
Set up, operate, and maintain Kubernetes-based environments, including GKE clusters and Helm-based deployments.
Support deployment workflows for PyTorch and TensorFlow-based ML/AI workloads.
Support and optimize GPU cluster infrastructure for machine learning workloads.
Build and maintain CI/CD pipelines using tools such as GitHub Actions.
Set up and maintain deployment environments for ML services and infrastructure components.
Work with MLOps platforms, including ClearML, to support experiment tracking, orchestration, and production workflows.
Manage cloud infrastructure on GCP with attention to reliability, scalability, cost awareness, and security.
Build and maintain cloud networking and security configurations.
Implement observability and monitoring using tools such as Grafana, Loki, and Prometheus.
Collaborate with DevOps, ML, data, and engineering teams to improve platform reliability and developer experience.
Requirements: Required qualifications:
+6 years of experience in DevOps, with a focus on big data technologies such as BigQuery.
Strong experience with cloud platform (preferably GCP / AWS).
Strong hands-on experience with Kubernetes, including deploying, scaling, and managing containerized workloads in production.
Experience designing, building, and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or similar tools.
Solid experience with Docker, including building, securing, and optimizing production-ready container images.
Experience implementing and maintaining observability solutions using tools such as Datadog, Grafana, CloudWatch, or similar platforms.
Hands-on experience with GitOps practices and tools, particularly Argo CD, for automated and reliable application deployments.
Experience leveraging AI agents and Model Context Protocol (MCP) to automate and improve DevOps workflows, operational processes, and infrastructure management.
Preferred qualifications:
Experience in cloud financial management and cost optimization.
Experience with Terraform or Terragrunt, managing infrastructure as code across multiple environments.
Strong understanding of security and compliance protocols.
Experience with Agile methodologies and Scrum.
Desired personal traits:
Passion for staying up-to-date with the latest industry trends and technologies.
Strong ownership mindset and comfort operating production systems.
Practical problem-solving approach and attention to reliability.
Clear communication with technical stakeholders across infrastructure, ML, and engineering.
Curiosity about ML infrastructure and the systems that help scientific teams move faster.
Commitment to building inclusive, collaborative engineering environments.
This position is open to all candidates.