You will join our Media & Telecom division, building and operating mission-critical TV and streaming platforms used by global telecom operators.
This is a hands-on DevOps + SRE hybrid role, focused on engineering, reliability, and large-scale production systems. The team is responsible for designing, improving, and operating highly available services in a fast-moving, production-driven environment.
You will work closely with R&D, NOC, and Product teams, taking ownership of system behavior in production and driving continuous improvements in scalability, performance, and automation.
The day-to-day
Design, build, and improve scalable, highly available systems running on AWS
Work deeply with Kubernetes (K8s), Docker, and Helm charts in production environments
Build and maintain CI/CD pipelines (Jenkins, Argo Workflows) to support rapid and reliable delivery
Develop and manage infrastructure using Terraform (Infrastructure as Code)
Troubleshoot and resolve complex issues in distributed systems across application, infrastructure, network, data, and observability layers
Improve system reliability, monitoring, and observability using tools such as Prometheus, Coralogix (or similar)
Build, integrate, and operate AI-based systems and agents as part of production services
Collaborate with R&D, NOC, and Product teams to drive system improvements and resolve production challenges
Take part in on-call rotations as part of shared production ownership
Requirements: Ideally, were looking for:
3+ years of hands-on experience in DevOps / SRE roles
Proven experience in production environments with high-availability systems and on-call responsibility
Strong hands-on experience with: AWS (production environments), Kubernetes (K8s) and Helm charts, Docker / containerized systems
Experience building and maintaining CI/CD pipelines (Jenkins, Argo Workflows)
Hands-on experience with Terraform or similar IaC tools
Experience working with Linux and Windows systems
Strong understanding of networking concepts and troubleshooting (TCP/IP, DNS, load balancing, connectivity debugging)
Hands-on experience with monitoring and observability tools (e.g., Prometheus, Coralogix, or similar)
Familiarity with data layer technologies such as MSSQL, MongoDB, and Redis
These would also be nice:
Experience with AI/ML technologies in production environments
Experience with DevSecOps or DevFinOps practices
Experience with configuration management tools (e.g., Chef) or large-scale distributed systems (media / telecom)
This position is open to all candidates.