דרושים » הנדסה » Senior Site Reliability Engineer (DevTools)

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
4 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Senior Site Reliability Engineer (DevTools)
Your responsibilities will include:

Improving services based on user feedback

Building fault-tolerant, self-healing architecture

Finding ways to speed up our systems and reduce user friction

Modifying well-known closed- and open-source solutions Supporting our users
Requirements:
We expect you to have:

A combination of SRE and SWE experience (for us that's a 50/50 split). Our code is in Java/Kotlin, Go, Python, and Ruby

An understanding of what's happening under the hood in Unix-like systems and the JVM

A passion for improving the user experience

The ability to adapt quickly on the fly in a fast-changing environment
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761256
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 4 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an exceptional Platform Engineer who combines deep software engineering with a robust DevOps approach to help propel development infrastructure to the next level. , SPEED is integral to our DNA. AI transforms the way we develop at speed, and our development infrastructure is key to the scale and hypergrowth.



The DevX team is a multi-disciplinary group of DevOps and development experts focused on building and operating internal tools, infrastructure, and services to make engineering easy and intuitive SPEED.



Your mission is to build a top-tier Continuous Integration and Delivery (CI/CD) platform. You will own both the application services and the infrastructure that supports them, ensuring speed, stability, and reliability for our entire R&D organization.



Key Focus Areas & What You'll Do



Improve CI/CD pipelines and CI workflows (e.g., GitHub Actions) with a focus on speed and reliability.
Reduce CI flakiness and improve overall pipeline stability through systematic triage and root-cause analysis.
Shorten developer feedback loops by optimizing test strategy, pipeline consistency, and local development workflows.
Strengthen CD and release processes for secure, repeatable, and fast deployments.
Promote best practices such as GitOps and progressive delivery where they fit.
Track and improve delivery metrics, with emphasis on Lead Time for Changes and Deployment Frequency.
Partner with engineering teams to identify friction and deliver scalable automation and paved paths.
Requirements:
3+ years of experience as a Platform Engineer or a strong Backend Engineer with a DevOps focus.
Strong knowledge of CI/CD pipelines, versioning, and release management.
Proven experience with modern build tools (e.g., Bazel, SWC) and managing CI workflows (e.g., GitHub Actions, GitLab CI).
Good skills with infrastructure-as-code tools like Terraform and Helm.
Hands-on experience with Kubernetes, including containers (Docker), Kafka and cloud providers (AWS, GCP, or Azure).
Strong programming skills in Python, Go, TypeScript, or another modern backend language.
Experience with monorepo tools (Lerna, NX, PNPM) is a big plus.
Familiarity with GitOps workflows using tools like ArgoCD or Flux.
A strong sense of ownership and a passion for making systems reliable, scalable, and improving the developer experience.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8765802
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
15/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. we are building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep our company reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.
This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.
This position is based out of our new Tel Aviv office.
What You'll Do:
Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739673
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 4 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an exceptional Platform Engineer who combines deep software engineering with a robust DevOps approach to help propel development infrastructure to the next level. , SPEED is integral to our DNA. AI transforms the way we develop at speed, and our development infrastructure is key to the scale and hypergrowth.



The DevX team is a multi-disciplinary group of DevOps and development experts focused on building and operating internal tools, infrastructure, and services to make engineering easy and intuitive SPEED.



Your mission is to build a top-tier Continuous Integration and Delivery (CI/CD) platform. You will own both the application services and the infrastructure that supports them, ensuring speed, stability, and reliability for our entire R&D organization.



Key Focus Areas & What You'll Do



Improve CI/CD pipelines and CI workflows (e.g., GitHub Actions) with a focus on speed and reliability.
Reduce CI flakiness and improve overall pipeline stability through systematic triage and root-cause analysis.
Shorten developer feedback loops by optimizing test strategy, pipeline consistency, and local development workflows.
Strengthen CD and release processes for secure, repeatable, and fast deployments.
Promote best practices such as GitOps and progressive delivery where they fit.
Track and improve delivery metrics, with emphasis on Lead Time for Changes and Deployment Frequency.
Partner with engineering teams to identify friction and deliver scalable automation and paved paths.
Requirements:
What Were Looking For:



3+ years of experience as a Platform Engineer or a strong Backend Engineer with a DevOps focus.
Strong knowledge of CI/CD pipelines, versioning, and release management.
Proven experience with modern build tools (e.g., Bazel, SWC) and managing CI workflows (e.g., GitHub Actions, GitLab CI).
Good skills with infrastructure-as-code tools like Terraform and Helm.
Hands-on experience with Kubernetes, including containers (Docker), Kafka and cloud providers (AWS, GCP, or Azure).
Strong programming skills in Python, Go, TypeScript, or another modern backend language.
Experience with monorepo tools (Lerna, NX, PNPM) is a big plus.
Familiarity with GitOps workflows using tools like ArgoCD or Flux.
A strong sense of ownership and a passion for making systems reliable, scalable, and improving the developer experience.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8765844
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 5 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Senior DevOps Engineer to join our fantastic team!



Responsibilities:

Develop and architect our SaaS platform
Working with an extensive range of cloud infrastructure tools (multi-cloud: AWS, Azure & GCP), both for our cloud-based infrastructure, but mainly due to our comprehensive support of cloud technologies for our clients
As our customers include Fortune 500 companies, we face the challenge of working with immense amounts of data - Petabytes over Petabytes. As a result, our scale of operation is immense and requires us to use highly scalable infrastructures for our system.
Monitor, troubleshoot, and resolve production-grade issues. Configure the system and application aspects.
Develop architectural POCs
Work on a microservices architecture
Automate everything, keeping it DRY
Assist and work closely with the team across R&D
Work on our next-generation cloud-native applications
Requirements:
Requirements


At least 4 years experience focused on DevOps Engineering and Site Reliability - Must.
Linux/Unix and bash scripting - Must.
Experience in AWS / GCP / Azure cloud infrastructure production - Must.
In-depth knowledge of build/release systems CI/CD pipelines.
In-depth understanding of public cloud and application architectures and technologies - Must.
Hands-on experience with Docker (K8s orchestrators) - Must.
Experience in infrastructure as code tools, Terraform, Ansible, or CloudFormation - Must.
Experience with monitoring & logging tools.
Experience building/supporting a distributed production environment.
Multi DB expertise (MongoDB, PostgreSQL) - Advantage.
Proficiency in networking fundamentals and concepts (Load Balancing, Protocols).
Experience with NodeJS - Advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8765575
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
29/06/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Senior Infrastructure Engineer to work on our dedicated engineering team building processing pipelines, information storage systems, and presentation layers in support of intelligence analysts.

The collection, processing, and exploration of malware samples and other information at a large scale is at the core of the Intelligence mission The Intelligence Automation team is responsible for prototyping, building, and operating the systems that enable this mission, and we'd like you to join us!

Your job will be to build, maintain, and improve infrastructure to support the entire breadth of our teams activities. You will work on classical datacenter and cloud infrastructure as well as environments for malware sandboxing or world-wide threat monitoring and hunting.

Advancing our fast-paced intelligence mission, requirements sometimes shift rapidly, and projects can live anything from weeks to years depending on changes in the surrounding ecosystem. It will be your responsibility to provide an infrastructure that can keep up with and adapt to these changes.

Occasionally things inside or outside of your control break and you will use your debugging skills to pinpoint the issue no matter whether it is on a hardware, network, cloud, kernel, or user space level.

You will be responsible for all aspects of the infrastructure you design, build, and maintain. This includes gathering requirements, making technical choices, creating documentation, securing workloads, upskilling colleagues, proactively monitoring operations, and gathering feedback from stakeholders.

You will join a team of very experienced infrastructure engineers who will always have your back. However, as a remote employee on a team distributed across many regions and time zones, you will not have direct access to all of your co-workers for the entire workday. Thus, the ability to work unsupervised, communicate asynchronously, and take the initiative in maintaining lines of communication is crucial. Additionally, we are looking for someone who would like to be part of a team who are passionate about their work and go the extra mile to exceed expectations. We love enthusiastic individuals who bring a positive attitude to their work and really care about what they produce for our stakeholders.

What You'll Do:

Maintain a can-do attitude and be solution-oriented

Deliver on ambiguous assignments and quickly evolving requirements in a fast-paced environment

Design, implement, document, and maintain our multi-cloud infrastructure

Be a consultant to development teams to ensure smooth deployment, monitoring, and maintenance of applications and services

Develop and maintain infrastructure-as-code (IaC) and management automation tools

Secure traditional and AI workloads

Ensure high availability, scalability, and performance of our systems

Troubleshoot and resolve complex infrastructure issues

Deploy, monitor, and troubleshoot relational and NoSQL databases

Mentor junior engineers and contribute to knowledge sharing within the team

Stay up-to-date with industry trends, best practices, and emerging technologies

Judge security and compliance risk
Requirements:
5+ years of experience in a DevOps or SRE role, with a focus on cloud-native technologies

Experience running Kubernetes clusters

Solid understanding of the risks and limitations of AI-based tooling

Ability to work on a geographically distributed and diverse team

Ability to independently make sound, justifiable decisions and take action

Proficiency in Go or Python, with experience developing automation scripts and tools

Proficiency in Linux, networking fundamentals and Kubernetes

Experience with monitoring and logging tools like Prometheus, Grafana, Splunk, LogScale

Experience with cloud providers such as AWS, GCP, or Azure

Experience with infrastructure-as-code tools like Helm, terraform
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8715577
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior IT SRE Engineer, you will be a key player in ensuring the reliability, scalability, and performance of our critical IT infrastructure. You will leverage SRE principles and an automation-first mindset to build and maintain resilient hybrid cloud environments. This role is ideal for a candidate who thrives in a fast-paced, innovative setting and is passionate about solving complex challenges with cutting-edge technology.
Key Responsibilities
Provision, configure, and support resilient hybrid cloud deployment architectures using an Infrastructure-as-Code framework.
Proactively collaborate with development teams to ensure new applications are production-ready, scalable, and reliable from inception.
Develop and maintain tools and frameworks to automate operational tasks, including deployment, monitoring, and recovery.
Conduct thorough root cause analysis of production issues and implement preventative measures to improve system resilience, demonstrating strong problem-solving skills.
Manage CI/CD platforms, Linux infrastructure, and contribute to capacity planning and operational runbooks.
Design and implement proactive service monitoring, alerting, and trend analysis to maintain service availability and performance SLAs.
Participate in an on-call rotation to support critical applications and services, responding to and resolving incidents efficiently.
Contribute to comprehensive documentation related to infrastructure design, deployment, and operational procedures.
Requirements:
Your Expereience:
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
6+ years of Devops engineering experience on mission-critical, enterprise-level systems in a hybrid (both cloud and on-prem) environment.
3+ years of hands-on experience with cloud environments, preferably Google Cloud Platform (GCP).
Expertise in configuration management and Infrastructure-as-Code using frameworks such as Terraform and Ansible.
Strong programming/scripting knowledge in languages like Python, Bash, or Go for infrastructure automation.
Demonstrated experience with CI/CD pipelines (e.g., GitHub, Jenkins, Artifactory) and a strong foundation in Linux/Unix administration.
Preferred Qualifications
Experience with containerization and orchestration technologies, particularly Kubernetes.
Hands-on experience with monitoring and observability tools such as Datadog, Grafana, or Prometheus.
Understanding of networking principles including firewalls, load balancers, and complex network designs.
A curious and positive mindset with a passion for applied learning and challenging existing processes for continuous improvement.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8713868
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Software Engineer - Verification and Reliability.
In this role as a SDET (Software Development Engineer in Test), you are a developer first. You will join a high-impact team of engineers who write production-grade code to build a massive-scale validation ecosystem. Your job is to act as "The Breaker"-designing the infrastructure, chaos experiments, and AI-driven tools that push our platform to its theoretical limits.
What Youll Build:
Adversarial Engineering: Design and implement Python-based distributed frameworks capable of orchestrating millions of concurrent IO operations to hunt down race conditions and memory leaks.
AI-Augmented Validation: Be at the forefront of the AI-Native transformation. You will leverage LLMs and Generative AI to automate complex scenario generation, build intelligent agents for root-cause analysis, and multiply your engineering velocity.
Simulation & Chaos: Build the "Entropy Engine." You will develop tools that inject real-world failures - latency, packet loss, and hardware crashes - to prove the resilience of our Raft and RDMA implementations.
Deep-System Observability: Move beyond "Pass/Fail." You will build telemetry pipelines to track P99 latency and jitter, providing critical architectural feedback to the Core Kernel teams.
Collaborative Architecture: You will operate with the same rigorous standards as the Core R&D team: design docs, production-grade code reviews, and high-level architectural planning.
Requirements:
Extensive Coding Experience: 5+ years of hands-on Python development experience is required. You are a Python expert who understands the language "under the hood" and are comfortable reading and debugging C++, Rust, or Go to understand how the core system works.
Systems Engineering Mindset: You have a background in distributed systems, networking (TCP/IP, RDMA), or storage protocols. You understand the complexities of consistency and metadata at scale.
AI Enthusiast: You are an early adopter of AI tools (Copilot, LLMs) and are excited about using them to automate the most tedious parts of the engineering lifecycle.
The "SRE" Lens: You approach quality through the lens of Site Reliability Engineering. You care about observability, MTTD (Mean Time to Detection), and building self-healing testing loops.
Problem Hunter: You have a "hacker" instinct. You dont just find a bug; you find the architectural flaw that allowed it to exist.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8757535
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are seeking a promising and talented Senior DevFinOps Engineer to join our DevOps group. If you thrive in a fast-paced, dynamic environment, can handle multiple requests simultaneously, and enjoy working independently as part of a cutting-edge DevOps team, this is your opportunity to help make the world a safer place!
Key Responsibilities
Act as a DevFinOps Engineer within a highly skilled team, bridging engineering and finance to drive cloud cost efficiency across large-scale operations from development to production.
Design, develop, and maintain Avanan's cloud cost visibility, allocation, and optimization solutions - including tagging strategies, cost dashboards, budgets, and anomaly detection across accounts and services.
Implement tools and procedures for cost monitoring, forecasting, and alerting across our SaaS multi-tenant product family.
Embed FinOps practices into the CI/CD lifecycle - surfacing the cost impact of changes early, and enforcing cost guardrails as part of deployment automation.
build AI-based FinOps agents for cloud services at the infrastructure and application levels
Continuously identify and execute cost-optimization opportunities (right-sizing, reserved capacity/savings plans, spot usage, storage tiering, idle-resource cleanup) without compromising performance, reliability, or security.
Partner with engineering teams to drive cost accountability - providing unit-economics insight (cost per tenant/feature/service) and making cost a first-class engineering metric.
Plan capacity and model the financial impact of scaling decisions, balancing cost efficiency with fault tolerance and growth.
Design and shape our cost reporting, chargeback/showback, and budgeting solutions.
Execute all tasks with cloud infrastructure security and financial governance as guiding principles.
Requirements:
Hands-on mindset - we all write code daily!
3+ years of relevant DevOps/Cloud experience building and operating CI/CD pipelines for both development and production - must.
2+ years of AWS Cloud experience working with high-traffic systems and multiple services, with a strong grasp of AWS pricing models and cost management tooling (Cost Explorer, CUR, Budgets, Compute Optimizer) - must.
Strong scripting skills, with fluency in Python - must.
Experience working with AI tools to achieve cloud or application cost control, identify and investigate the root cause of cost changes (differentiating between organic growth / infrastructure change / config change / application change).
Demonstrated experience driving measurable cloud cost reductions and building cost-optimization tooling or automation - must.
Experience with containers and orchestration tools (Docker, Kubernetes, or ECS) and understanding of their cost drivers - must.
Familiarity with FinOps principles and practices (FinOps Foundation framework, showback/chargeback, unit economics) - an advantage.
Experience with CI integration tools such as Jenkins.
Familiarity with AWS CloudFormation and infrastructure-as-code - an advantage.
Exposure to a wide range of open-source technologies (Redis, Nagios, Grafana, Prometheus, etc.) and cost-analytics tooling.
Knowledge of best practices in security, performance, monitoring, and cost governance.
Proven ability to research, evaluate, and implement new technologies, including running proofs of concept and cost analysis.
Huge advantage: measuring and controlling AI cost.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8723227
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
Your experience:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8714862
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a technically strong and AI-savvy SRE Team Lead & Escalation Manager to own production reliability, incident management, and cross-functional prioritization. This role leads our AI-driven automation strategy, drives self-healing infrastructure development, and sets a new standard for modern reliability engineering.
Key Responsibilities
Lead and mentor the SRE team; improve monitoring, alerting, and observability.
Own production incidents and escalations end-to-end - from mitigation to RCA and corrective action.
Lead the design and development of self-healing systems capable of detecting, diagnosing, and remediating incidents autonomously.
Drive automation of repetitive operational workflows using AI/ML-based solutions to reduce toil and MTTR.
Manage the cross-functional Squad handling customer and production issues; align priorities across Support, QA, R&D, and Sources.
Track key operational metrics and lead long-term reliability improvements.
Requirements:
3-5 years in SRE or Incident Management.
Mandatory: Hands-on experience applied to operational challenges (AIOps, anomaly detection, LLM-based automation, or auto-remediation).
Proven track record of automating workflows and reducing manual toil at scale.
Strong cloud background (AWS/Azure/GCP) and experience with Kubernetes, Docker, and CI/CD.
Proficiency with observability tools (Grafana, Prometheus, ELK) and scripting (Python, Bash).
Demonstrated leadership in high-pressure, cross-functional environments.
Advantages
Background in cybersecurity or SaaS platforms.
Experience with LLMOps, AI agents, or orchestration platforms (e.g., n8n, Temporal).
Key Attributes
Strong ownership, accountability, and composure under pressure.
Passionate about leveraging AI to automate workflows, reduce toil, and accelerate incident resolution.
Visionary about self-healing operations - able to both define the strategy and drive its implementation.
Collaborative leader with the ability to align cross-functional stakeholders.
Technically hands-on systems-level thinker with the drive to engineer scalable, long-term solutions.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8720950
סגור
שירות זה פתוח ללקוחות VIP בלבד