דרושים » ניהול ביניים » Senior Site Reliability Engineer (SRE)

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Required Senior Site Reliability Engineer (SRE)
Realize your potential by joining the leading performance-driven advertising company
As a Senior Site Reliability Engineer on the R&D Infrastructure team in our Tel Aviv Office, youll play a vital role in building, scaling and maintaining high-scale infrastructure across on-premise cloud, public cloud and rapidly growing AI/ML Kubernetes environments. You will push Linux to its limits, writing software and automation to eliminate manual tasks and solving performance bottlenecks across the stack.
How youll make an impact:
Keep our hybrid infrastructure (on-prem, public cloud and AI/ML clusters) highly available, performant and cost-efficient.
Build internal software tooling and manage IaC pipelines in Go, Python or Rust to eliminate repetitive operations.
Perform deep-dive troubleshooting across the full stack-from CDN edge configurations down to Linux kernel tuning and network layer bottlenecks.
Design and maintain monitoring and alerting setups to spot and address system health issues before they impact users.
Participate in on-call rotations, lead incident resolution and conduct blameless post-mortems to ensure system resilience.
Requirements:
To thrive in this role, youll need:
7+ years of experience managing, scaling and troubleshooting large-scale distributed Linux environments in production.
Deep understanding of Linux system internals and network protocols (TCP/IP, DNS, HTTP, gRPC), together with hands-on experience of edge/CDN services such as Fastly, Cloudflare, Akamai or CloudFront.
Hands-on experience with Infrastructure as Code (IaC) and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD or Jenkins.
Production experience managing containerized environments using Kubernetes and Docker.
Solid programming skills in at least one modern language (Go, Python or Rust).
Bonus points if you have:
Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK/logging).
Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8841079
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the Technology team might be right for you. We're looking for a Senior DevOps Engineer with a strong security orientation to join our DevOps team. our platform runs at significant scale, and our DevOps team sits at the core of keeping it fast, reliable, and secure. As we grow, so does the scope of what we own - we need an engineer who can step in as a strong pillar of our production ecosystem. This is a hands-on role for someone who takes security seriously, moves fast, and knows how to get things done in a complex, high-scale environment. Our Technology Stack: AWS, GCP, CLoudFlare,Kubernetes, Terragrunt, Ansible, Jenkins, ArgoCD, Argo Workflows, Kong & Nginx, HashiCorp Vault, Kafka, RabbitMQ, Mongodb, Aurora Postgresql & Mysql, Prometheus, Grafana, VictoriaMetrics Programming languages: Python, NodeJS, Go, Kotlin


What am I going to do?:

* Full Ownership: Drive infrastructure initiatives through their entire lifecycle, taking accountability from initial design to delivery and long-term operations.
* Kubernetes Orchestration: Architect, implement, and maintain production-grade Kubernetes clusters, ensuring they remain scalable and resilient under high-scale demand.
* Cloud Architecture: Design and manage robust AWS environments, including VPC, IAM, and EKS, to support a highly available platform architecture.
* Infrastructure as Code: Utilize Terraform to build and evolve our environment, applying configuration management principles to all IaC workflows.
* CI/CD Excellence: Support and improve our deployment pipelines using Jenkins and GitHub Actions to maintain a fast development velocity.
* Observability: Implement comprehensive monitoring solutions with Prometheus and Grafana to ensure deep visibility into platform health.
* Operational Resilience: Join the DevOps on-call rotation, taking responsibility for mitigating production issues and maintaining site reliability.
* Tooling & Innovation: Continuously evaluate and adopt tools - security and otherwise - that raise the bar on engineering efficiency and security posture.
Equal opportunities:
We're not about checklists. If you don't meet 100% of the requirements for this role but still feel passionate about the position and think you have the right skills and qualifications to excel at it, we want to hear from you. We prioritize diversity. We celebrate difference and embed it into every aspect of our workplace and product, as well as our community. We are proud and committed to providing equal opportunity employment to all individuals regardless of race, color, religion, sex, sexual orientation, citizenship, national origin, disability, Veteran status, or any other characteristic protected by law. In addition, we will provide accommodation to individuals with disabilities or a special need.
Requirements:
* 6+ years of hands-on DevOps / Platform Engineering experience in large-scale production environments on a public cloud (AWS preferred).
* Proven leadership mindset - able to own projects end to end and be accountable for outcomes.
* Strong, production-grade Kubernetes experience across design, deployment, scaling, and troubleshooting.
* Solid AWS experience with VPC, IAM, EC2, EKS, Load Balancers, and DNS.
* Experience designing and operating highly available, scalable infrastructure systems.
* Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis).
* Hands-on
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8799862
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do at our company and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the company Technology team might be right for you.
We're looking for a Senior DevOps Engineer with a strong security orientation to join our company's DevOps team.
our company's platform runs at significant scale, and our DevOps team sits at the core of keeping it fast, reliable, and secure. As we grow, so does the scope of what we own - we need an engineer who can step in as a strong pillar of our company's production ecosystem.
This is a hands-on role for someone who takes security seriously, moves fast, and knows how to get things done in a complex, high-scale environment.
our company's Technology Stack:
AWS, GCP, CLoudFlare ,Kubernetes, Terragrunt, Ansible, Jenkins, ArgoCD, Argo Workflows, Kong & Nginx, HashiCorp Vault, Kafka, RabbitMQ, Mongodb, Aurora Postgresql & Mysql, Prometheus, Grafana, VictoriaMetrics
Programming languages: Python, NodeJS, Go, Kotlin
What am I going to do?
Full Ownership: Drive infrastructure initiatives through their entire lifecycle, taking accountability from initial design to delivery and long-term operations.
Kubernetes Orchestration: Architect, implement, and maintain production-grade Kubernetes clusters, ensuring they remain scalable and resilient under high-scale demand.
Cloud Architecture: Design and manage robust AWS environments, including VPC, IAM, and EKS, to support a highly available platform architecture.
Infrastructure as Code: Utilize Terraform to build and evolve our environment, applying configuration management principles to all IaC workflows.
CI/CD Excellence: Support and improve our deployment pipelines using Jenkins and GitHub Actions to maintain a fast development velocity.
Observability: Implement comprehensive monitoring solutions with Prometheus and Grafana to ensure deep visibility into platform health.
Operational Resilience: Join the DevOps on-call rotation, taking responsibility for mitigating production issues and maintaining site reliability.
Tooling & Innovation: Continuously evaluate and adopt tools - security and otherwise - that raise the bar on engineering efficiency and security posture at our company.
Requirements:
6+ years of hands-on DevOps / Platform Engineering experience in large-scale production environments on a public cloud (AWS preferred).
Proven leadership mindset - able to own projects end to end and be accountable for outcomes.
Strong, production-grade Kubernetes experience across design, deployment, scaling, and troubleshooting.
Solid AWS experience with VPC, IAM, EC2, EKS, Load Balancers, and DNS.
Experience designing and operating highly available, scalable infrastructure systems.
Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis).
Hands-on experience with Infrastructure as Code and configuration management (Terraform required; Terragrunt and Ansible a plus).
Experience with Docker and containerized workloads.
2+ years building and maintaining CI/CD pipelines (Jenkins, GitHub Actions)
Experience with monitoring and observability tools (Prometheus, Grafana).
Experience with GenAI platforms (AWS Bedrock, Vertex AI, OpenAI).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8806869
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
28/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Senior Platform Engineer to join our CTI Group. We're a startup that builds a lot, so you'll work across a wide range of systems and tools, build new platforms and services, and have real influence over where the platform goes next.
You'll work on the platform behind our CTI products: the microservices running on Kubernetes, the cloud and on-prem environments around them, and the infrastructure behind the systems. There's a lot still to build, and you'll have a real say in what comes next.
Responsibilities
Own what you build end to end, from design through implementation to running it in production.
Design and build the internal platforms, tools, and services other engineering teams rely on, and add new ones as we grow.
Make architectural decisions across infrastructure, applications, networking, security, and data.
Design, build, and maintain our cloud and on-prem environments.
Build and operate the infrastructure behind our AI/ML systems at scale.
Manage and evolve our Kubernetes environments.
Build and maintain CI/CD pipelines, deployment automation, and Infrastructure-as-Code with Terraform.
Build security into the platform, from IAM and secrets management to network policies.
Set up the monitoring and observability that show us how the platform is doing.
Design for scalability, availability, performance, and resilience in high-scale production systems.
Apply Infrastructure-as-Code practices (Terraform) end to end.
Requirements:
5+ years in DevOps or Platform Engineering roles.
Strong software engineering skills, including production-grade Python development using OOP, and proficiency in Bash scripting.
Experience designing and building platforms and infrastructure from scratch.
A strong architectural sense: you understand how infrastructure, applications, networking, security, and data fit together.
Hands-on AWS experience.
Solid Kubernetes experience, including deploying, operating, and troubleshooting clusters.
Experience with Terraform and Docker, and comfort with Linux, networking, and security fundamentals.
Experience with high-scale production systems and with the infrastructure behind AI/ML workloads.
A startup mindset: hands-on, curious, and comfortable with broad scope and a lot of room to shape your own work.
Bonus Points:
Experience in the cyber security field.
On-prem, customer private cloud, or air-gapped deployment experience.
Model serving and ML tooling such as MLFlow, and GPU scheduling on Kubernetes.
Observability stacks such as Prometheus, Grafana, or Loki.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836113
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an independent Senior DevOps Engineer who thrives in a fast-paced environment. You will join our centralized DevOps team, serving as a pillar of reliability and innovation for the entire engineering organization.

This is a 50/50 "Build vs. Run" role. Splitting your time between building modern automation and improving our developer experience and operational excellence, ensuring our production environments are rock-solid, secure, and highly available while managing the live environment from server infrastructure all the way to the application.

Tasks you will take part in:
Design and implement next-generation CI/CD automation flows to streamline the path from "Code" to "Production."
Lead the evolution of our Infrastructure-as-Code (Terraform) to make our systems more modular and self-service for developers.
Collaborate with backend teams to architect scalable microservices and optimize datastore performance (RDS, MongoDB).
Embrace AI-powered development tools (Cursor, Claude Code) to automate toil and accelerate infrastructure delivery.
Full ownership of our AWS accounts and multi-cluster Kubernetes environments.
Manage and fine-tune live production environments, ensuring 99.9% uptime for our mission-critical fintech services.
Drive our Observability strategy (Datadog/Grafana) to proactively catch issues before they impact our customers.
Participate in a weekly on-call rotation, serving as the first line of defense for platform reliability.
is used for creating requisitions 'from scratch'
Requirements:
Requirements:
5+ years of experience managing high-scale production systems.
Mastery of AWS: Deep knowledge of AWS services, security best practices, and account management.
Orchestration & Containers: Expert-level experience with Kubernetes (EKS) and Docker.
Automation-First Mindset: Proficiency in Terraform and at least one high-level language (Python, Go, or Bash).
Database Ops: Solid understanding of the DevOps aspects of MySQL/RDS and MongoDB (scaling, backups, performance tuning).
SRE Discipline: Strong experience with monitoring and logging stacks like Datadog, Grafana, or Prometheus.
Team Player: Strong communication skills and a "can-do" attitude-you enjoy solving problems as part of a collaborative squad.
Experience with AI tools and a strong interest in continuously exploring and applying them in everyday work are highly valued.

Advantages:
Security Knowledge: Experience with FW, WAF, IPS/IDS, and SELinux.
Data Infrastructure: Familiarity with Data tools like Kafka, Spark, Snowflake or Data Lakes.
Multi-Cloud: Familiarity with hybrid or multi-cloud architectures.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8831065
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Required Senior Site Reliability Engineer (Cortex)
Your Career:
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834292
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
23/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Required Senior Principal DevOps Engineer (Cortex Cloud)
Your Impact:
As a Senior Principal DevOps Engineer, you will serve as a visionary technical leader within the Cortex Cloud DevOps group. You will define the technical strategy and architecture that ensures our massive-scale production services remain highly reliable, exceptionally secure, and performant. You will pioneer the integration of AI-driven capabilities into our daily operations, establishing elite engineering standards and fundamentally transforming the workflows of hundreds of developers through autonomous agents and intelligent procedures.
Your Career:
Architectural Vision & Scalability: Design and scale massive, resilient distributed systems and global Kubernetes infrastructure, implementing robust observability and monitoring frameworks.
AI-Driven Transformation: Revolutionize the SDLC by integrating Generative AI, autonomous agents, and LLM-powered workflows into CI/CD and self-healing systems to accelerate developer velocity.
Technical Leadership & IaC: Define architectural standards, lead GitOps/IaC (FluxCD/Terraform) strategies, and mentor Senior/Staff engineers across the R&D organization.
Developer Experience & Efficiency: Build and champion AI-powered platforms that automate troubleshooting and eliminate friction. Partner directly with Engineering Directors, Principal Architects, and Product Management to align infrastructure initiatives with business goals, optimizing for scale, high availability, and multi-million-dollar cost-efficiencies.
Security & Compliance: Embed "Security by Design" principles into the platform architecture to ensure platform integrity without sacrificing delivery speed.
Requirements:
Your Experience:
10+ years of progressive experience in DevOps, SRE, Platform, or Infrastructure Engineering roles, with a significant portion at the principal/ tech leadership/ staff, or architectural level.
System Design from Scratch: A proven track record of designing, building, and deploying large-scale, highly available distributed systems and cloud platforms from the ground up.
AI-Powered Automation: Proven experience designing and integrating AI-driven systems, autonomous agents, and LLM-based tools into engineering workflows to optimize development processes, procedures, and overall organizational efficiency.
Communication: Exceptional interpersonal skills, capable of articulating complex architectural and AI workflow concepts clearly to both deeply technical peers and executive leadership.
Cloud & IaC Mastery: Expert-level proficiency with GCP (or equivalent major cloud providers) and deep architectural experience with Terraform.
Advanced Container Orchestration: Deep, internal knowledge of virtualized and containerized environments, with architectural-level expertise in scaling Kubernetes, extending it via custom operators, and automating complex operational logic.
Software Engineering Approach: Advanced coding and automation skills in Python or Go. You treat infrastructure as a software engineering discipline and can build custom tooling/services when off-the-shelf solutions fall short.
Proven Leadership: Demonstrated ability to lead complex, cross-team technical initiatives from conception to delivery, including setting technical roadmaps and driving consensus among stakeholders.
OS/Systems Expertise: Mastery of Linux systems, including kernel tuning, advanced networking, and performance troubleshooting.
Nice to Have:
Deep expertise in managing and scaling stateful workloads and distributed databases (e.g., Cassandra, ScyllaDB, MemSQL, or MySQL) in containerized environments.
Experience contributing to open-source infrastructure projects, or presenting at major tech conferences.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8831043
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo and Ra'anana
Job Type: Full Time
As a Senior DevOps Engineer, youll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIAs AI infrastructure.


What You'll Be Doing:
Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations.
Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints.
Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory.
Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity.
Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations.
Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows.
Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening.
Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes.
Requirements:
What We Need to See:
Bachelors degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices.
Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting.
Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible.
Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices.
Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting.
Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting.
Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment.


Ways To Stand Out From the Crowd:
Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost.
Familiarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage.
Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana.
Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers.
Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837920
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
2 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company Software Technologies, we secure the world - and now we're securing the AI revolution. We're building a Workforce AI Security Platform: a cloud-native, multi-tenant SaaS platform that governs, protects, and enables safe AI adoption across global enterprises. This is a greenfield opportunity to help architect the infrastructure backbone of a platform that will define how organizations adopt AI at work securely.
We're looking for a Senior DevOps Engineer who owns infrastructure end-to-end, ships with confidence, and raises the reliability bar without being asked. You will work closely with backend, full-stack, and security teams, as well as DevOps teams across other our company organizations, to build a highly available, reliable, and secure production environment. If you get energized by building systems that scale, pipelines that teams love, and platforms that never sleep - this role is for you.
Key Responsibilities
Own and evolve our cloud infrastructure across multi-region production environments, end-to-end.
Lead our GitOps deployment model - designing and maintaining declarative, automated deployment workflows with zero manual gates.
Build, maintain, and optimize CI/CD pipelines with a strong focus on developer experience, reliability, and speed.
Initiate, implement, and champion an AI-first DevOps & SRE ecosystem - identifying opportunities, building AI agents and intelligent automation, and driving their adoption across engineering operations.
Develop automation frameworks for provisioning, scaling, observability, and incident response, leveraging AI-powered tooling and agentic workflows to reduce toil.
Operate and improve our observability platform: metrics, logs, alerting, dashboards, SLOs/SLIs, and on-call tooling.
Champion zero-trust secrets management and credential-less authentication patterns across the stack.
Partner with architects and engineering leadership on cloud cost optimization, availability, and performance.
Build internal tooling and automation that multiplies engineering velocity across the organization.
דרישות:
5+ years of hands-on DevOps experience in a SaaS product environment - Must.
Demonstrated initiative in applying AI to engineering operations - designing and building AI agents, agentic workflows, LLM-powered automation, or Model Context Protocol (MCP) integrations that reduced operational toil and improved production reliability, quality, or velocity - Must.
Strong scripting and programming skills - Python and Bash for automation, tooling, and AI agent development; Go is a plus.
Strong motivation to continuously learn and adopt emerging technologies, and to share that knowledge across the team.
Deep, hands-on AWS expertise; multi-cloud (AWS, GCP, Azure) experience is a strong plus - Must.
Strong understanding of containers and orchestration - Docker, Kubernetes, including workloads, networking, service mesh (Istio), Helm/Kustomize, and autoscaling (KEDA, HPA, VPA).
Strong experience with:
Infrastructure-as-Code - Terraform, Crossplane, and/or cloud-native declarative tooling.
GitOps principles and tooling (ArgoCD or equivalent).
CI/CD platforms - building reusable, scalable, security-hardened pipeline templates (GitHub Actions or equivalent).
Secrets management - dynamic injection, IRSA/Workload Identity, avoiding long-lived credentials.
Experience embedding security into CI/CD: vulnerability scanning, SBOM generation, and supply chain security (Trivy, Grype, Syft, JFrog Xray).
Solid observability knowledge - OpenTelemetry, Prometheus, Grafana, Datadog, ELK/OpenSearch, distributed tracing.
Hands-on experience with AI/ML workloads or LLMOps infrastructure - a significant advantage.
Cost-awareness (FinOps) - treating cloud spend as a core engineering metric.
Clear communication skills - able to align engineers, security teams, and leadership around infrastructure decisions.
A strong sense of ownership - proactively identifying gaps and driving improvemen המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8841149
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Responsibilities
- Design, build, and operate the internal engineering platform powering our company's build, test, deployment, and security validation workflows at scale
- Write and maintain production-grade Python and shell tooling that drives platform automation - this is a hands-on coding role, not just pipeline configuration
- Architect and manage hybrid cloud/on-prem execution infrastructure, including large-scale Kubernetes runner pools across multiple AWS regions
- Own and evolve CI/CD pipelines at scale using GitHub Actions, including reusable workflows, ARC-based runner orchestration, and build caching strategies (BuildKit, sccache, Valkey)
- Operate and tune DinD environments (Sysbox, EBS/NVMe, overlay storage, MTU/networking) for build, test, and release workloads
- Connect and manage self-hosted and on-prem runners, routing physical device (wbox) test jobs by site and device type
- Implement DevSecOps controls including least-privilege IAM, OIDC, isolated runner groups, container signing, and automated security scans
- Drive platform observability, cost optimization, and reliability improvements across the engineering infrastructure
- Collaborate cross-functionally with hundreds of engineers to improve engineering velocity and release confidence
- Take end-to-end ownership of complex infrastructure problems and drive them to resolution.
Requirements:
Technical Skills
- 5+ years of hands-on DevOps experience with a strong software development background - prior development experience is a must
- B.Sc. in Computer Science or equivalent practical experience
- Strong programming skills in Python (or a similar high-level language); ability to write and own production tooling
- Proven experience designing and building scalable systems, automation frameworks, and infrastructure as code using Terraform and Helm
- Solid understanding of Linux, containers (Docker), and Git-based workflows
- Hands-on experience with CI/CD at scale using GitHub Actions or similar - including reusable actions, workflow design, and automation frameworks
- Deep experience with hybrid cloud infrastructure (AWS and on-prem), including EKS, ARC, Karpenter, ECR, S3, Direct Connect, VPC endpoints, IAM/OIDC, and Secrets Manager
- Experience operating spot and on-demand runner pools for builds, DinD tests, releases, and security scans across multiple AWS regions
- Experience with DinD environments (Sysbox, EBS/NVMe, memory limits, overlay storage, MTU/networking) and build caching (BuildKit, sccache, Valkey)
- Experience connecting on-prem/self-hosted runners and routing physical device (wbox) test jobs by site and device type
- Experience implementing DevSecOps controls and improving platform observability, cost efficiency, and reliability
- Platform & tooling familiarity: Kubernetes (EKS, on-prem) GitHub Actions ARC Karpenter Terraform Helm Docker/DinD Sysbox containerd BuildKit ECR S3 ElastiCache (Valkey) sccache Direct Connect VPC endpoints IAM/OIDC Secrets Manager self-hosted runners
Soft Skills
- Strong system-level thinking and troubleshooting skills; able to diagnose and resolve complex infrastructure issues independently
- Takes end-to-end ownership and drives problems to resolution without hand-holding
- Excellent communication and cross-team collaboration skills; comfortable working alongside large engineering organizations
Nice to Have / Advantage
- Experience with Jenkins
- Familiarity with GitHub merge queue
- Experience with MinIO or on-prem S3 caching
- Hardware-in-the-loop CI experience
- MTU/VPC networking tuning expertise
- Monorepo CI optimization experience.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8828046
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Senior DevOps Engineer who thrives in high-scale, high-performance environments. In this role, you will design and manage mission-critical infrastructure, develop automation tools, and ensure the reliability and scalability of our platform. You will work closely with engineering teams to optimize performance, enhance observability, and prevent incidents before they happen.
If you love solving complex infrastructure challenges and building tools that empower engineers, this role is for you!
What Will You Do?
As a Senior DevOps Engineer, you will:
Design, build, and maintain a highly available, scalable, and resilient production infrastructure handling millions of requests per second.
Manage and optimize tens of AWS accounts in a multi-account cloud environment.
Developing AI-driven automation frameworks and tools that more than 400 R&D engineers use while enhancing system reliability and engineering efficiency.
Design, implement, and support the LLM and agentic workflow infrastructure, ensuring its scalability and reliability.
Manage thousands of servers and containers, ensuring seamless infrastructure operations.
Improve and maintain our logging, monitoring, and alerting stacks for enhanced observability.
Troubleshoot and mitigate production incidents, participating in on-call rotations to ensure system health and stability.
Collaborate with engineering teams to optimize performance, reduce latency, and improve deployment processes.
Requirements:
What Will You Bring To The Team?
5+ years of experience in building and maintaining production infrastructure at scale.
5+ years of experience in developing automation tools and server-side applications using Python, Go, Ruby, Java, or Node.js.
Strong cloud experience with AWS, Google Cloud, or similar platforms.
Experience with the deployment, management, and leveraging of Large Language Models (LLMs) and agentic infrastructure.
Deep knowledge of Linux systems, including troubleshooting, architecture, and system internals.
Solid understanding of web servers, load balancers, caching systems, relational databases, and networking.
A proactive, problem-solving mindset with a passion for optimizing infrastructure and AI-driven automation.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8838344
סגור
שירות זה פתוח ללקוחות VIP בלבד