דרושים » ניהול ביניים » Principal DevOps Engineer (Cortex Agentix Endpoint Security)

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 16 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779592
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
7 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769987
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Improve observability and incident response - raise signal quality, cut alert noise, and own the monitoring the triage agents depend on.
Eliminate toil - relentlessly identify manual, repetitive operational work and remove it through automation and agents, protecting engineering time for reliability work that only humans can do.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, capacity risks, and automation opportunities.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Requirements:
Your Experience:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8771789
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
15/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. we are building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep our company reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.
This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.
This position is based out of our new Tel Aviv office.
What You'll Do:
Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739673
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
03/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Youll be building infrastructure that supports work with frontier AI labs such as Google, Anthropic, and OpenAI, and meets the scale, security, and reliability standards this ecosystem requires. This includes architecting ephemeral, on-demand environments, and spinning up and tearing down multi-service deployments (VPCs, compute, serverless components, and more) as needed.

You will work closely with research, engineering, security, and leadership to design and maintain scalable infrastructure, improve system reliability, and enable fast, safe, and secure product delivery.

Key Responsibilities

Own and evolve Irregulars cloud infrastructure, with a strong focus on scalability, security, and reliability

Design, build, and maintain end-to-end CI/CD pipelines and deployment workflows

Lead and continuously improve containerized environments (Docker, Kubernetes)

Build and maintain Infrastructure as Code and automation frameworks (e.g., Terraform)

Own production readiness: availability, monitoring, logging, alerting, and incident response

Work closely with engineering and research teams to support development, testing, and production needs

Drive improvements in system performance, reliability, and security posture

Take part in - and often lead architectural decisions and long-term infrastructure planning

Act as a key operational owner in a fast-growing startup, setting best practices and standards
Requirements:
5+ years of hands-on experience as a DevOps / Infrastructure Engineer in production environments

Strong, hands-on experience with AWS (architecture, networking, security, and cost awareness)

Proven experience with containerization and orchestration (Docker, Kubernetes) in production

Hands-on experience designing and maintaining CI/CD pipelines (e.g., GitHub Actions or similar)

Strong experience with Infrastructure as Code and automation (e.g., Terraform)

Ability to take end-to-end ownership and operate independently in an early-stage, high-ambiguity environment
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8766476
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a Senior Infrastructure Engineer who views "Infrastructure as Software." In 2026, we dont just manage servers; we build high-performance environments that allow multi-agent systems to operate at scale.
You will be a core member of the R&D team, blending deep DevOps expertise with the coding rigor of a Backend Engineer. You arent just "configuring" AWS; you are architecting the distributed systems and data pipelines that power our autonomous security brain. Your mission is to ensure that while our agents are evolving and taking actions, our underlying platform remains immutable, observable, and infinitely scalable.
What You'll Do
Design, build, and operate our company's cloud infrastructure using AWS, Kubernetes, and Infrastructure as Code.
Build internal tools and platform services using Python and Go to improve developer productivity and system reliability.
Own infrastructure automation with Terraform, Pulumi, and modern cloud-native tooling.
Partner closely with Backend, Data Science, and Security Engineering teams to build scalable, reliable platforms.
Improve observability, monitoring, and incident response across distributed production systems.
Design and optimize infrastructure for performance, scalability, security, and cost efficiency.
Help shape engineering best practices, platform architecture, and developer experience as our company continues to grow.
Requirements:
5+ years of experience in Infrastructure, DevOps, Platform Engineering, or Backend Engineering.
Strong software engineering skills with hands-on experience building production systems in Python or Go.
Deep hands-on experience with AWS, including services such as EKS, RDS, VPC, and IAM.
Strong experience designing, operating, and scaling production Kubernetes environments.
Experience with Infrastructure as Code, CI/CD, GitOps, and modern cloud-native development practices.
A systems mindset with the ability to solve architectural challenges across infrastructure and application layers.
Comfortable using modern AI-powered developer tools and agentic workflows to improve engineering productivity.
The company Mindset: You take ownership, act with accountability, collaborate openly, and focus on delivering meaningful impact. You thrive in fast-moving environments, embrace ambiguity, and enjoy solving hard problems together.
Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.
Full professional fluency (written and verbal) in both Hebrew and English.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8764502
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a DevOps Engineer.
our DevOps team operates the infrastructure that powers our AI and Computer Vision platform across construction sites in 15+ countries. From data pipelines and ML workloads to backend services - you'll work with a diverse, modern, Kubernetes-based stack and have real influence on how we build, deploy, and operate.
What you'll do:
Own Multi-Cloud Infrastructure: Work alongside the team to design, scale, and operate our high-scale, multi-region production infrastructure across AWS and GCP, powering construction sites globally.
Drive Kubernetes at Scale: Manage and evolve our Kubernetes platform on EKS and leveraging GitOps practices with ArgoCD and Helm to enable safe, fast, and reliable deployments.
Build Robust CI/CD: Design and maintain CI/CD pipelines that empower dozens of engineers to ship confidently - with automation, testing, and progressive delivery built in.
Tackle Diverse Infrastructure Challenges: Work hands-on with a wide variety of workloads - from heavy data processing and Computer Vision pipelines to backend services and ML inference - each with unique scaling, performance, and reliability requirements.
Ensure Reliability & Observability: Build and maintain world-class observability (metrics, logs, tracing, alerting) so that issues are caught early and resolved fast. Performance, reliability, and scalability are at the core of what you do.
Security & Cost: Partner with the team to strengthen our security posture, identity and access management, compliance, and cloud cost optimization across both clouds.
Ownership from 0 to 1: You will have real influence over our architecture and tooling. We want engineers who care about shaping what we build and how we build it, ensuring performance, security, and observability are baked in from day one.
Requirements:
A seasoned DevOps / Infrastructure engineer (5+ years) with strong hands-on experience in production cloud environments.
Proven expertise operating large-scale, distributed systems - with deep understanding of Kubernetes, networking, and cloud-native architecture.
Strong experience with multi-cloud environments (AWS and/or GCP), Infrastructure-as-Code (Terraform), and GitOps workflows (ArgoCD, Flux, or similar).
Hands-on experience with CI/CD systems (Jenkins, GitHub Actions, etc.).
Solid scripting and automation skills (Python, Bash, or Go).
Proven track record of being a collaborative team player who partners closely with developers, ML engineers, and cross-functional stakeholders across the organization.
Experience with observability stacks (Prometheus, Grafana, OpenTelemetry, Logz.io, or similar).
Experience with databases (relational and/or NoSQL) - including operational aspects like backups, migrations, and performance tuning.
AI-Native Engineering: You are an AI-native engineer who leverages LLMs and agentic tools (like Cursor, Copilot, or Claude) not just for command completion, but as a core operational partner - automating diagnostics, runbooks, and infrastructure workflows so you can focus on the critical things
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8729164
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a hands-on Senior DevOps Engineer with a strong cloud-native mindset to build, maintain, and evolve our highly scalable, highly-available cloud infrastructure. This role is pivotal in driving operational excellence, security, and automation across our entire engineering organization. You will promote communication, integration, and collaboration to significantly enhance our software development productivity and reliability. You'll work closely with engineering and product teams to streamline delivery, enforce platform standards, and enable a high-velocity development environment-all while keeping reliability and security top of mind.
Responsibilities:
Design, Automate, and Manage complex cloud infrastructure on AWS using best-in-class Infrastructure as Code (IaC) practices.
Lead the operation and enhancement of our production Kubernetes environments (EKS), focusing on automation, security, observability, and seamless CI/CD integration.
Drive continuous improvement across platform tooling, developer experience, and operational processes to meet our ambitious performance and uptime goals.
Implement and enforce security-first infrastructure patterns, including strong IAM, network segmentation, and secure secrets management.
Actively contribute to high-level technical design discussions and cross-functional architectural decision-making, ensuring solutions align with long-term platform strategy.
Requirements:
7+ years of experience as a DevOps Engineer, Platform Engineer, or in a similar infrastructure-focused role.
Strong hands-on expertise across the AWS Stack (e.g. EC2, EKS, RDS, VPC, IAM, S3, Lambda).
Mastery of Infrastructure as Code - Terraform or equivalent.
Deep operational knowledge of Kubernetes, including architecture, cluster management, networking, and advanced debugging in production environments.
Strong expertise in designing and managing CI/CD methodologies and platforms (e.g. Jenkins, Github Actions).
Experience with monitoring tools such as Prometheus, DataDog, Coralogix (OTEL), Grafana etc.
Proven prior experience building and maintaining highly-available, production-grade, and service-oriented systems.
Strong scripting and automation background in languages such as Python or Bash.
Exceptional communication and collaboration skills with the ability to articulate complex technical needs and influence cross-functional teams.
Strong knowledge of AWS Networking - an advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8748461
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
7 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
Your experience:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769584
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Who We Are: Incredibuild is the leading platform empowering developers and enterprises to radically accelerate their development cycles. We help world-leading brands like Microsoft, Citi Bank, GM, Amazon, and Adobe shorten build times, allowing for more iterations, faster product releases, and significant resource savings. Our technology streamlines and accelerates everything from compilation to release automation, dramatically expediting time to market for our clients. You aren’t just "fixing CI/CD pipelines" - you are building a robust Internal Developer Platform (IDP) and creating the tools that enable our engineers to ship high-performance code at scale. If you love writing clean code and get equally excited about a perfectly orchestrated Kubernetes cluster or an elegant Terra grunt module, you’ll fit right in.
What you'll do: Architect & Build: Take full ownership of our infrastructure, from backend services to cloud resources. Empower Developers: Develop and maintain our homegrown IDP , ensuring a seamless "golden path" for engineering teams. Scale the Stack: Design and manage scalable, reliable environments using AWS, Argo CD, and Helm. Automate Everything: Build sophisticated CI/CD workflows using GitHub Actions and manage stateful infrastructure with Terra grunt. Innovate: Evaluate and integrate new CNCF tools and internal utilities to keep our platform at the cutting edge.
Requirements:
* 3-5 years of hands-on experience in DevOps, Platform Engineering, Site Reliability Engineering (SRE), or Cloud Infrastructure roles within modern SaaS environments.
* Hands-on experience deploying, operating, and troubleshooting Kubernetes environments in production.
* Strong experience working with AWS services and containerized workloads.
* Proven experience using Terraform or Terragrunt to provision, manage, and maintain cloud infrastructure at scale.
* Experience building and maintaining CI/CD pipelines using tools such as GitHub Actions.
* Familiarity with GitOps practices and technologies including Argo CD and Helm.
* Strong scripting and automation skills using Python, Bash, Go, or similar languages.
* Ability to develop internal tools and operational automation that improve developer productivity and platform reliability.
* Passionate about solving infrastructure challenges related to scalability, reliability, security, and operational efficiency.
* Committed to reducing developer friction and improving the overall engineering experience.
How We Work We are a small, highly collaborative team with a strong sense of ownership. Engineers are encouraged to take initiative, contribute to technical decisions, and have a meaningful impact on the platform's evolution. We move quickly while maintaining high engineering standards, balancing pragmatism with long-term maintainability.
Our Tech Stack AWS | Kubernetes | Terraform | Terragrunt | GitHub Actions | Argo CD | Helm | Homegrown Internal Developer Platform (IDP).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8693323
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a skilled and motivated DevOps engineer with deep familiarity in the streaming ecosystem to join our elite infrastructure team. If you're excited by the challenge of operating mission-critical systems at scale and optimizing the developer experience through automation and tooling, wed love to hear from you.
What you will do:
Automate Deployment and Operation:
Oversee deployment of Kafka and RabbitMQ clusters (including Confluent Cloud & CFK). Build automation pipelines to ensure repeatability and resiliency across environments.
Monitor and Support Production Systems:
Own production stability of global Kafka clusters. Handle on-call rotations, incident management, troubleshooting, and scaling challenges.
Improve Infrastructure Observability
Build and maintain observability systems: dashboards, alerting pipelines, metrics collection (Prometheus, Grafana, etc.).
Optimize System Performance:
Collaborate with peers on benchmarking and optimization initiatives. Work on tuning Kafka brokers, cluster configurations, and runtime parameters.
Provide Developer Support and Training (Infra-focused)
Help developers configure topics, quotas, and consumers appropriately. Train service owners to interpret monitoring data and avoid pitfalls.
Develop and Maintain Infrastructure:
Contribute to building infrastructure tools and scripts (IaC, Helm charts, etc.) that make provisioning and managing clusters reliable and efficient.
Secure Infrastructure Access:
Configure and maintain secure access patterns across streaming infrastructure, ensuring proper authentication and role-based access controls are enforced for both developers and services.
Requirements:
8+ years of experience in DevOps, SRE, or Infrastructure Engineering roles.
Deep hands-on Kafka experience, including deploying, maintaining, scaling, and monitoring clusters.
Experience with RabbitMQ.
Extensive experience with Docker, Kubernetes, Helm, and GitOps-style deployments.
Infrastructure as Code experience (Terraform, Pulumi, etc.).
Strong skills in scripting and automation (Python, Bash, etc.).
Familiarity with Confluent Cloud, Confluent for Kubernetes, and similar tools.
Solid understanding of authentication and authorization mechanisms in distributed systems.
Production support mindset - with proven troubleshooting and incident resolution history.
Collaboration and communication skills - especially with dev teams depending on platform support.
Experience with Istio Service Mesh (bonus).
Experience with GovCloud (bonus).
Bonus Qualities:
Mentorship and leadership experience in infrastructure or SRE teams.
Contributions to automation or monitoring open-source tooling.
Active participant in SRE or DevOps communities.
Conference speaker or internal tech trainer.
Technical writing about infrastructure automation or reliability.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8754250
סגור
שירות זה פתוח ללקוחות VIP בלבד