דרושים » מחשבים ורשתות » Site Reliability Team Leader

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
3 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an experienced SRE Team Lead to drive the reliability, observability, and automation practices across our private cloud infrastructure and operations. In this role, you will lead a team of site reliability engineers, own the engineering roadmap for monitoring and automation, and act as a key liaison between development, operations, and platform teams. You bring at least 3-4 years of hands-on people management experience and a deep technical background in SRE or DevOps disciplines.


What will you do?

Leadership & Team Management

Lead, mentor, and grow a team of SREs, providing technical direction, career development guidance, and day-to-day management.

Own the team roadmap for reliability, observability, and automation initiatives - prioritizing work, removing blockers, and driving delivery.

Conduct regular 1:1s, performance reviews, and hiring processes to build and sustain a high-performing team.

Foster a culture of operational excellence, blameless post-mortems, and continuous improvement.

Act as an escalation point for complex incidents and reliability issues, leading post-incident reviews and ensuring follow-through on action items.


Automation & Infrastructure

Design, develop, and maintain automation tools to support infrastructure and operations teams at scale.

Manage pipelines and infrastructure workflows using Jenkins, Ansible, Python, and Bash.

Drive the adoption of infrastructure-as-code practices across the organization.

Collaborate with system engineers to improve scalability, performance, and fault tolerance of critical systems.


Monitoring & Observability

Build and extend monitoring and alerting systems using Grafana, the ELK (Elastic) stack, Zabbix, and custom scripts.

Implement and enforce observability best practices to ensure full visibility into systems, applications, and infrastructure.

Define and track SLIs, SLOs, and error budgets across key services.

Partner with development teams to embed observability earlier in the software development lifecycle.


Database & Platform Support

Support monitoring and infrastructure integration for databases including MongoDB and PostgreSQL.

Maintain documentation and champion knowledge sharing around automation, monitoring, and reliability practices.
Requirements:
Experience & Leadership:

3-4+ years of experience in a people management or team lead capacity within SRE, DevOps, or infrastructure engineering.

5-8+ years of overall experience in SRE, DevOps, or infrastructure automation roles.

Proven track record of building, coaching, and retaining high-performing engineering teams.

Experience owning an engineering roadmap and driving cross-functional reliability initiatives.


Technical Skills :

Strong scripting skills in Python and Bash; comfortable building and maintaining production-grade automation.

Hands-on experience with infrastructure automation tools, particularly Ansible.

Solid experience with monitoring and observability platforms - ELK stack, Grafana, and Zabbix.

Good understanding of CI/CD pipelines and related tooling, including Jenkins.

Familiarity with managing and monitoring MongoDB and PostgreSQL in a production environment.

Comfortable working in Linux-based environments.

Excellent problem-solving skills and strong written and verbal communication.


Ability to support the following:

Experience with cloud providers - AWS, GCP, or Azure.

Exposure to containerization technologies such as Docker and Kubernetes.

Familiarity with infrastructure provisioning using Terraform.

Experience introducing SRE practices (SLOs, error budgets, chaos engineering) at an organizational level.

Exposure and experience with migrating/ building AI tools to improve process.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8760168
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced DevOps Engineer to join our Engineering team and play a key role in building and operating our cloud-native platform. The ideal candidate will have hands-on experience managing production environments at scale, driving cloud transformation initiatives, and supporting the of enterprise systems from on-premises deployments to modern SaaS and cloud-native architectures. You will be responsible for designing, automating, and maintaining scalable infrastructure, CI/CD pipelines, and deployment processes that support both our core products and emerging AI-driven capabilities. Working closely with Engineering, QA, Product, and AI teams, you will help ensure the reliability, security, and performance of our services while driving operational excellence and continuous improvement across our technology stack.
Responsibilities:
Design, implement, and maintain CI/CD pipelines.
Manage and optimize cloud infrastructure across AWS, Azure, and/or GCP.
Develop and maintain Infrastructure as Code using Terraform.
Manage Kubernetes-based environments and GitOps deployment workflows using Argo CD and Kustomize.
Lead and support the migration of enterprise applications and infrastructure from on-premises
environments to scalable SaaS and cloud-native architectures.
Establish, maintain, and continuously improve production environments, ensuring high availability, security, scalability, and operational excellence.
Demonstrate strong production ownership, including incident management, root cause analysis, capacity planning, and performance optimization.
Collaborate with Engineering, QA, Product, and AI teams.
Support the deployment, operation, and monitoring of AI and Generative AI services.
Build and maintain monitoring, logging, and alerting systems.
Troubleshoot and resolve infrastructure, deployment, and production issues.
Requirements:
5+ years of experience as a DevOps Engineer or similar
infrastructure-focused role.
Hands-on experience with Azure, Aws, or GCP.
Experience with CI/CD tools such as Jenkins, GitHub Actions, or similar platforms.
Strong knowledge of Terraform and Infrastructure as Code practices.
Experience with Docker, Kubernetes, Argo CD, and Kustomize.
Experience designing and operating production-grade Kubernetes saas platforms or enterprise environments.
Experience with monitoring and observability tools such as Prometheus, Grafana, and ELK.
Strong troubleshooting, analytical, and communication skills.
B.Sc. in Computer Science, Computer Engineering, Information Systems, or a related field (or equivalent practical experience).
Nice to have
Experience supporting AI, Machine Learning, or Generative AI workloads, including familiarity with MLOps concepts, AI deployment platforms, or cloud-based AI services.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739903
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an experienced DevOps Manager to lead and grow our DevOps function. This role combines people leadership, technical direction, and ownership of the infrastructure, tooling, automation, and operational practices that power Stamplis production environment.

You will manage DevOps Engineers, hire and onboard an additional team member, and drive the strategy, execution, and evolution of Stamplis internal DevOps platform. You will work closely with Engineering, Data, AI, Product, and Security teams to improve developer experience, enable fast and safe delivery, and keep production stable.

This role requires a strong hands-on DevOps / Platform Engineering background, combined with proven leadership capabilities. If you believe DevOps should operate as a self-service platform, love automation, and think in systems and end-to-end flows, keep reading.


What You Will Do
Lead, mentor, and manage a DevOps team, fostering ownership, excellence, collaboration, and continuous improvement.
Own the DevOps roadmap, priorities, execution, and delivery, aligned with Engineering, Data, AI, Security, and business goals.
Provide technical and architectural guidance across infrastructure, CI/CD, cloud operations, automation, observability, security, and platform engineering initiatives.
Build and evolve our internal DevOps platform, creating self-service capabilities, internal services, and golden paths that scale across teams.
Own CI/CD end-to-end, including Jenkins, GitHub, and GitHub Actions pipelines from commit to production.
Oversee and evolve our AWS stack, including ECS, EKS, Lambda, DynamoDB, Redshift, S3, DocumentDB, networking, IAM, observability, and deployment patterns.
Enable MLOps and data workflows using tools such as Airflow, MLflow, and Jupyter Notebooks.
Drive an automation-first mindset through Infrastructure-as-Code, scripting, internal tooling, and reusable components.
Lead cost optimization efforts with a FinOps mindset, including visibility, budgets, rightsizing, and workload efficiency.
Ensure security is embedded into DevOps practices, including least privilege, secrets management, vulnerability scanning, secure SDLC, and incident readiness.
Leverage AI-assisted development tools such as Cursor, GitHub Copilot, Claude Code, and ChatGPT Enterprise to improve team productivity and delivery speed.
Collaborate closely with cross-functional stakeholders to unblock delivery, improve developer experience, and maintain production stability.
Requirements:
7+ years of experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering, ideally in a SaaS production environment.
2+ years of managerial or team leadership experience, including mentoring engineers, driving execution, and owning team delivery.
Strong hands-on technical background with the ability to guide architecture, review technical decisions, and stay close to execution when needed.
Strong development background: you write code comfortably, build internal tools, and approach infrastructure work with software engineering discipline.
Proven experience with AWS services such as ECS, EKS, Lambda, DynamoDB, Redshift, S3, and DocumentDB.
Strong CI/CD experience with Jenkins, GitHub, and GitHub Actions.
Experience with Infrastructure-as-Code, automation, observability, cloud networking, IAM, and production operations.
Experience with ML/data tooling such as Airflow, MLflow, and Jupyter Notebooks - an advantage.
Hands-on experience with AI-assisted development tools such as Cursor, GitHub Copilot, Claude Code, or similar.
Demonstrated experience in cost optimization, cloud security, secure SDLC, and operational security.
A wide-angle thinker who sees the whole system, understands dependencies, and builds solutions that scale across teams.
Strong people leadership, communication, collaboration, and prioritization skills.
Strong communication skills in English. Hebrew is an advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8709491
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior IT SRE Engineer, you will be a key player in ensuring the reliability, scalability, and performance of our critical IT infrastructure. You will leverage SRE principles and an automation-first mindset to build and maintain resilient hybrid cloud environments. This role is ideal for a candidate who thrives in a fast-paced, innovative setting and is passionate about solving complex challenges with cutting-edge technology.
Key Responsibilities
Provision, configure, and support resilient hybrid cloud deployment architectures using an Infrastructure-as-Code framework.
Proactively collaborate with development teams to ensure new applications are production-ready, scalable, and reliable from inception.
Develop and maintain tools and frameworks to automate operational tasks, including deployment, monitoring, and recovery.
Conduct thorough root cause analysis of production issues and implement preventative measures to improve system resilience, demonstrating strong problem-solving skills.
Manage CI/CD platforms, Linux infrastructure, and contribute to capacity planning and operational runbooks.
Design and implement proactive service monitoring, alerting, and trend analysis to maintain service availability and performance SLAs.
Participate in an on-call rotation to support critical applications and services, responding to and resolving incidents efficiently.
Contribute to comprehensive documentation related to infrastructure design, deployment, and operational procedures.
Requirements:
Your Expereience:
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
6+ years of Devops engineering experience on mission-critical, enterprise-level systems in a hybrid (both cloud and on-prem) environment.
3+ years of hands-on experience with cloud environments, preferably Google Cloud Platform (GCP).
Expertise in configuration management and Infrastructure-as-Code using frameworks such as Terraform and Ansible.
Strong programming/scripting knowledge in languages like Python, Bash, or Go for infrastructure automation.
Demonstrated experience with CI/CD pipelines (e.g., GitHub, Jenkins, Artifactory) and a strong foundation in Linux/Unix administration.
Preferred Qualifications
Experience with containerization and orchestration technologies, particularly Kubernetes.
Hands-on experience with monitoring and observability tools such as Datadog, Grafana, or Prometheus.
Understanding of networking principles including firewalls, load balancers, and complex network designs.
A curious and positive mindset with a passion for applied learning and challenging existing processes for continuous improvement.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8713868
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a technically strong and AI-savvy SRE Team Lead & Escalation Manager to own production reliability, incident management, and cross-functional prioritization. This role leads our AI-driven automation strategy, drives self-healing infrastructure development, and sets a new standard for modern reliability engineering.
Key Responsibilities
Lead and mentor the SRE team; improve monitoring, alerting, and observability.
Own production incidents and escalations end-to-end - from mitigation to RCA and corrective action.
Lead the design and development of self-healing systems capable of detecting, diagnosing, and remediating incidents autonomously.
Drive automation of repetitive operational workflows using AI/ML-based solutions to reduce toil and MTTR.
Manage the cross-functional Squad handling customer and production issues; align priorities across Support, QA, R&D, and Sources.
Track key operational metrics and lead long-term reliability improvements.
Requirements:
3-5 years in SRE or Incident Management.
Mandatory: Hands-on experience applied to operational challenges (AIOps, anomaly detection, LLM-based automation, or auto-remediation).
Proven track record of automating workflows and reducing manual toil at scale.
Strong cloud background (AWS/Azure/GCP) and experience with Kubernetes, Docker, and CI/CD.
Proficiency with observability tools (Grafana, Prometheus, ELK) and scripting (Python, Bash).
Demonstrated leadership in high-pressure, cross-functional environments.
Advantages
Background in cybersecurity or SaaS platforms.
Experience with LLMOps, AI agents, or orchestration platforms (e.g., n8n, Temporal).
Key Attributes
Strong ownership, accountability, and composure under pressure.
Passionate about leveraging AI to automate workflows, reduce toil, and accelerate incident resolution.
Visionary about self-healing operations - able to both define the strategy and drive its implementation.
Collaborative leader with the ability to align cross-functional stakeholders.
Technically hands-on systems-level thinker with the drive to engineer scalable, long-term solutions.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8720950
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
15/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. we are building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep our company reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.
This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.
This position is based out of our new Tel Aviv office.
What You'll Do:
Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739673
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
29/06/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Senior Infrastructure Engineer to work on our dedicated engineering team building processing pipelines, information storage systems, and presentation layers in support of intelligence analysts.

The collection, processing, and exploration of malware samples and other information at a large scale is at the core of the Intelligence mission The Intelligence Automation team is responsible for prototyping, building, and operating the systems that enable this mission, and we'd like you to join us!

Your job will be to build, maintain, and improve infrastructure to support the entire breadth of our teams activities. You will work on classical datacenter and cloud infrastructure as well as environments for malware sandboxing or world-wide threat monitoring and hunting.

Advancing our fast-paced intelligence mission, requirements sometimes shift rapidly, and projects can live anything from weeks to years depending on changes in the surrounding ecosystem. It will be your responsibility to provide an infrastructure that can keep up with and adapt to these changes.

Occasionally things inside or outside of your control break and you will use your debugging skills to pinpoint the issue no matter whether it is on a hardware, network, cloud, kernel, or user space level.

You will be responsible for all aspects of the infrastructure you design, build, and maintain. This includes gathering requirements, making technical choices, creating documentation, securing workloads, upskilling colleagues, proactively monitoring operations, and gathering feedback from stakeholders.

You will join a team of very experienced infrastructure engineers who will always have your back. However, as a remote employee on a team distributed across many regions and time zones, you will not have direct access to all of your co-workers for the entire workday. Thus, the ability to work unsupervised, communicate asynchronously, and take the initiative in maintaining lines of communication is crucial. Additionally, we are looking for someone who would like to be part of a team who are passionate about their work and go the extra mile to exceed expectations. We love enthusiastic individuals who bring a positive attitude to their work and really care about what they produce for our stakeholders.

What You'll Do:

Maintain a can-do attitude and be solution-oriented

Deliver on ambiguous assignments and quickly evolving requirements in a fast-paced environment

Design, implement, document, and maintain our multi-cloud infrastructure

Be a consultant to development teams to ensure smooth deployment, monitoring, and maintenance of applications and services

Develop and maintain infrastructure-as-code (IaC) and management automation tools

Secure traditional and AI workloads

Ensure high availability, scalability, and performance of our systems

Troubleshoot and resolve complex infrastructure issues

Deploy, monitor, and troubleshoot relational and NoSQL databases

Mentor junior engineers and contribute to knowledge sharing within the team

Stay up-to-date with industry trends, best practices, and emerging technologies

Judge security and compliance risk
Requirements:
5+ years of experience in a DevOps or SRE role, with a focus on cloud-native technologies

Experience running Kubernetes clusters

Solid understanding of the risks and limitations of AI-based tooling

Ability to work on a geographically distributed and diverse team

Ability to independently make sound, justifiable decisions and take action

Proficiency in Go or Python, with experience developing automation scripts and tools

Proficiency in Linux, networking fundamentals and Kubernetes

Experience with monitoring and logging tools like Prometheus, Grafana, Splunk, LogScale

Experience with cloud providers such as AWS, GCP, or Azure

Experience with infrastructure-as-code tools like Helm, terraform
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8715577
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Required NOC Team Leader
Description:
The NOC Team Lead is responsible for ensuring the high availability, reliability, and performance of a company's production systems, infrastructure, and applications. This hybrid role bridges traditional reactive monitoring (Network Operations Center - NOC) with proactive, automated, and software-centric engineering practices (Site Reliability Engineering - SRE).
Key Responsibilities:
Operational Leadership: Directs 24/7 global NOC teams, managing incident response, service uptime, and operational excellence.
Proactive Monitoring & Observability: Implements monitoring, alerting, and observability tools (e.g., Prometheus, Grafana, Datadog) to track system health via golden signals (latency, traffic, errors, saturation).
Automation and Toil Reduction: Leads efforts to automate repetitive operational tasks and manual troubleshooting to improve system reliability and reduce human error.
Incident Management & Root Cause Analysis (RCA): Oversees the management of critical incidents, ensures timely communication, and performs post-mortem analysis to prevent recurrence.
Team Management & Development: Coaches, mentors, and develops NOC engineers and SREs, fostering a culture of high performance and continuous improvement.
Stakeholder Collaboration: Partners with engineering, development, and IT teams to align system performance with business goals and SLA requirements.
Key Differences in Focus:
NOC Focus: Primarily monitoring, detecting, and responding to incidents, ensuring connectivity and managing alerts.
SRE Focus: Focuses on engineering improvements, reducing technical debt, automation, and system resilience.
The Transformation: Modern managers are transforming traditional, manual NOCs into automated, SRE-driven environments.
Requirements:
Experience: 3+ years in managing technical teams within NOC, SRE, or infrastructure domains.
Technical Skills: Proficiency in cloud platforms (AWS, Azure, GCP), Kubernetes, Linux/Unix, and scripting languages (Python, Bash, Golang).
Tools: Experience with monitoring platforms (Datadog) and CI/CD tools (e.g., Jenkins, GitLab).
Soft Skills: Strong communication, leadership, and crisis management skills.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8728013
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
Your experience:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8704900
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time and Hybrid work
This role requires working out of the Tel Aviv office three days per week.

About the Role:
As a SRE Technical Lead at our company, you will play a critical role in ensuring the reliability, scalability, and performance of the platform that powers our customers operations. Youll operate at the intersection of software engineering and production operations, taking full ownership of the systems you build and run.
This is a hands-on individual contributor role - not a managerial position. You'll be deep in the technical work, driving impact through engineering excellence rather than people management.
This role is not just about responding to incidents - its about fundamentally improving how our platform behaves under real-world conditions. You will drive reliability initiatives end-to-end: defining measurable service goals, shaping engineering priorities through error budgets, and implementing solutions that prevent issues before they occur.
Youll work closely with teams across the organization, embedding reliability and observability into every layer of the stack. At the same time, youll leverage automation, modern infrastructure practices, and emerging AI capabilities to continuously evolve how we operate and scale.

What you will do:
Develop deep product knowledge across our platform - understanding its internals, failure modes, and operational behavior well enough to own incident resolution end-to-end.
Define and track SLAs/SLOs/SLIs across critical platform services, and use error budgets to drive engineering decisions.
Own production reliability - including on-call rotations, incident response, and post-mortems - with a focus on minimizing MTTR and preventing recurrence through systemic fixes, not just firefighting.
Work hand-in-hand with engineering teams across the stack - infrastructure, application, and business layers - to embed reliability requirements everywhere.
Requirements:
What skills and experience youll bring to our company:
5+ years of experience as an SRE or Platform Developer (or similar) in a high-scale production environment, with hands-on ownership across the full stack - infrastructure and application layers.
Experience introducing or scaling AI-powered systems in real-world products (ML, LLMs, agents, or decision systems)
Strong coding skills and a software engineering mindset - you build your own tools rather than waiting for someone else to.
A true owner - you take responsibility for systems end-to-end and proactively drive improvements without waiting for direction.
Business-level reliability experience is a strong advantage.
Experience with infrastructure-as-code and modern container orchestration platforms.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8743767
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Senior DevOps Engineer to join our R&D team in developing the next rising product in the health tech landscape. If you are looking for a challenging, influential position and are passionate about making an impact, this might be the role for you.

As a Senior DevOps Engineer , youll play a key role in the design, development, testing, deployment, and monitoring of our infrastructure and products. In this position, you'll make significant contributions to our observability stack, helping build and maintain robust systems for logs, metrics, traces, and alerting.

Our ideal candidate is passionate about DevOps and observability, has strong communication skills, and thrives on constant improvement for both technology and processes. If you enjoy working on multiple projects in parallel and are a proactive team player, youll fit right in.

This is a unique opportunity to join the core team of a fast-growing startup, where your contributions will have a direct impact on our product and success.

Responsibilities

Support and collaborate with cross-functional engineering teams using cutting-edge technologies.
Contribute to the design, implementation, and maintenance of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Loki)
Secure, scale, and manage our cloud environments (AWS and GCP)
Design and implement automation solutions for both development and production
Manage and improve our CI/CD pipelines for fast and safe delivery
Lead best practices in infrastructure, observability, configuration management, and system hardening
Continuously assess and improve existing infrastructure in line with industry standards
Requirements:
BSc in Computer Science, Engineering, or equivalent experience
5+ years of experience as a DevOps Engineer or similar software engineering role
Proven experience with Docker and Kubernetes (EKS preferred)
Hands-on experience with monitoring and observability tools, including Prometheus, Grafana, Datadog, or similar.
Expertise in Terraform for AWS infrastructure-as-code deployments
Strong collaboration and interpersonal communication skills
Excellent analytical thinking and problem-solving mindset
Proficiency with relational databases
Solid knowledge of Python and Bash scripting
Experience with test automation - an advantage
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8729405
סגור
שירות זה פתוח ללקוחות VIP בלבד