דרושים » מחשבים ורשתות » Senior DevOps Engineer

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
משרה זו סומנה ע"י המעסיק כלא אקטואלית יותר
שם חברה חסוי
מיקום המשרה: תל אביב יפו
סוג משרה: משרה מלאה
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a Senior Infrastructure Engineer who views "Infrastructure as Software." In 2026, we dont just manage servers; we build high-performance environments that allow multi-agent systems to operate at scale.
You will be a core member of the R&D team, blending deep DevOps expertise with the coding rigor of a Backend Engineer. You arent just "configuring" AWS; you are architecting the distributed systems and data pipelines that power our autonomous security brain. Your mission is to ensure that while our agents are evolving and taking actions, our underlying platform remains immutable, observable, and infinitely scalable.
What You'll Do
Design, build, and operate our company's cloud infrastructure using AWS, Kubernetes, and Infrastructure as Code.
Build internal tools and platform services using Python and Go to improve developer productivity and system reliability.
Own infrastructure automation with Terraform, Pulumi, and modern cloud-native tooling.
Partner closely with Backend, Data Science, and Security Engineering teams to build scalable, reliable platforms.
Improve observability, monitoring, and incident response across distributed production systems.
Design and optimize infrastructure for performance, scalability, security, and cost efficiency.
Help shape engineering best practices, platform architecture, and developer experience as our company continues to grow.
Requirements:
5+ years of experience in Infrastructure, DevOps, Platform Engineering, or Backend Engineering.
Strong software engineering skills with hands-on experience building production systems in Python or Go.
Deep hands-on experience with AWS, including services such as EKS, RDS, VPC, and IAM.
Strong experience designing, operating, and scaling production Kubernetes environments.
Experience with Infrastructure as Code, CI/CD, GitOps, and modern cloud-native development practices.
A systems mindset with the ability to solve architectural challenges across infrastructure and application layers.
Comfortable using modern AI-powered developer tools and agentic workflows to improve engineering productivity.
The company Mindset: You take ownership, act with accountability, collaborate openly, and focus on delivering meaningful impact. You thrive in fast-moving environments, embrace ambiguity, and enjoy solving hard problems together.
Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.
Full professional fluency (written and verbal) in both Hebrew and English.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8764502
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769987
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Improve observability and incident response - raise signal quality, cut alert noise, and own the monitoring the triage agents depend on.
Eliminate toil - relentlessly identify manual, repetitive operational work and remove it through automation and agents, protecting engineering time for reliability work that only humans can do.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, capacity risks, and automation opportunities.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Requirements:
Your Experience:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8781551
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are now hiring a talented, self-driven and passionate Senior Data Engineer to build and maintain optimized and highly available data pipelines that facilitate deeper analysis and reporting.
What Youll Do:
Design, develop, and maintain scalable data pipelines that integrate data from multiple sources, including APIs, databases, streaming platforms, edge computing devices, cloud services, and on-premise systems.
Design reliable and business-critical data processing workflows for batch and real-time use cases.
Develop data access services and APIs that enable efficient communication between edge devices, cloud infrastructure, and on-premise environments.
Design and optimize storage solutions for structured, semi-structured, and high-dimensional sensor data to support AI model training, inference, and analytics.
Strong understanding of distributed systems and scalable data processing architectures.
Build and maintain scalable streaming data pipelines for low-latency processing and event-driven architectures.
Analyze existing data architecture, storage models, and processing workflows, and continuously improve performance, scalability, reliability, and maintainability.
Optimize cloud infrastructure and data storage costs while maintaining high availability and low-latency access.
Collaborate closely with Software, AI, Algorithms, DevOps, and Product teams to translate business requirements into scalable technical solutions.
Design monitoring, observability, and operational processes for data platforms.
Requirements:
Bachelors degree in Computer Science, Engineering, Mathematics, or a related quantitative field.
5+ years of professional experience in Data Engineering or a related role.
Strong experience designing and implementing large-scale data pipelines using orchestration frameworks such as Apache Airflow, Prefect, or similar.
5+ years of software development experience, including at least 2 years of Python development.
Strong knowledge of relational and NoSQL databases such as PostgreSQL, MySQL, MongoDB, Elasticsearch/OpenSearch, ClickHouse, or similar technologies.
Experience designing and implementing streaming and event-driven data architectures using technologies such as AWS Kinesis, Amazon SQS, RabbitMQ, Kafka, or similar messaging systems.
Experience designing REST APIs and backend services (FastAPI or similar frameworks).
Experience working with AWS cloud services (S3, EC2, Lambda, CloudWatch, IAM, etc.).
Experience with Git, Docker, CI/CD pipelines, and modern software engineering practices.
Excellent communication and collaboration skills with engineering, AI, and Product teams.
Self-driven, innovative, and continuously looking for ways to improve systems and processes.
Great to Have:
Experience with Kubernetes and container orchestration.
Experience with distributed computing platforms and distributed data processing systems.
Experience building ML data pipelines supporting training and inference workloads.
Experience working with large-scale sensor, IoT, or time-series data.
Experience with monitoring and observability tools such as Grafana, Prometheus, ELK, Kibana, or OpenSearch.
Experience working in edge computing or hybrid cloud environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8773889
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
15/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a Senior AI Engineer to join our AI team. This is a greenfield opportunity to shape how we build and deploy intelligent systems at Aura, designing LLM powered agents, building production AI infrastructure, and embedding agentic workflows into the engineering culture. The ideal candidate is a builder at heart: someone who ships fast, operates with urgency, and sees AI not just as a tool but as a platform for rethinking how software gets built.
What you'll be doing:
Build Autonomous Agents: Design and develop autonomous agents that accelerate Aura's engineering lifecycle - from AI powered grooming, to coding, testing and shipping to production. Reduce cycle time end to end.
Tackle complex engineering challenges: Contribute to the evolution of our AI capabilities. Build full-stack products and platforms that teams rely on for decision making.
Scale AI-Powered Operations: Build and scale AI driven tooling that reduces production downtime and cuts developer overhead in incident research and response.
Production AI Ownership: Own the full lifecycle of AI features - from prototype to production deployment, monitoring, and continuous iteration.
Evaluate \\& Improve: Build evaluation pipelines, observability tooling, and feedback loops to measure and improve AI system quality in production.
Cross-Functional Collaboration: Partner with Data Scientists, Product, and Engineering teams to identify high-value AI use cases and ship them end to end.
Requirements:
Experience: 5+ years as a Software Engineer, with 1-2 years hands-on building and shipping AI/LLM-powered systems to production.
Engineering Fundamentals: Strong backend engineering skills with sound software engineering principles - APIs, testing, and clean architecture.
Agentic Systems: Hands-on experience designing and building LLM-powered agents using modern agentic frameworks (e.g., LangChain, LangGraph, Claude/OpenAI Agents SDK), including experience building MCP servers.
Production LLM: Proven track record deploying agentic applications to production - managing latency, cost, reliability, and failure modes.
Evaluation \\& Observability: Experience building evaluation frameworks and LLM observability tooling (e.g., Langfuse or similar).
Cloud \\& Infrastructure: Hands-on with cloud platforms (GCP/AWS), containerization (Docker, Kubernetes), and CI/CD pipelines.
AI Tooling: Fluency with modern AI coding tools (Cursor, Claude Code, Copilot) and agentic workflows - leveraging AI to accelerate the full software development lifecycle.
Cross-Functional Collaboration: Ability to partner with Data Scientists, ML Engineers, Product Managers, and Analysts to translate requirements into production systems.
Ownership \\& Urgency: Strong sense of ownership and urgency, comfort with ambiguity, ability to operate independently and move fast.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739974
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Senior AI Engineer to design and build production-grade, LLM-powered systems. You'll work at the intersection of software engineering and applied AI - shipping agents, RAG pipelines, and tool-using systems that solve real problems at scale. This is a hands-on, high-ownership role for someone who thrives at the frontier of what's possible with modern LLMs and isn't afraid to write the glue, the infrastructure, and the prompts that make it all work.
This is a **cross-functional, company-wide role**. You won't be embedded in a single product team - instead, you'll partner with every department to identify high-leverage opportunities and build AI-powered tools and workflows that boost productivity and efficiency across the entire organization.
This is a great opportunity to be part of one of the fastest-growing infrastructure companies in history, an organization that is in the center of the hurricane being created by the revolution in artificial intelligence.
"our company's data management vision is the future of the market."- Forbes
we are the data platform company for the AI era. We are building the enterprise software infrastructure to capture, catalog, refine, enrich, and protect massive datasets and make them available for real-time data analysis and AI training and inference. Designed from the ground up to make AI simple to deploy and manage, our company takes the cost and complexity out of deploying enterprise and AI infrastructure across data center, edge, and cloud.
Our success has been built through intense innovation, a customer-first mentality and a team of fearless workers who leverage their skills & experiences to make real market impact. This is an opportunity to be a key contributor at a pivotal time in our companys growth and at a pivotal point in computing history.
What You'll Do:
- Design, build, and operate LLM-powered applications, agents, and workflows end-to-end - from prototype to production.
- Architect retrieval, context engineering, and tool-use strategies that make models reliable, accurate, and cost-efficient.
- Integrate LLMs with internal services, third-party APIs, and data stores to automate complex business and engineering workflows.
- Build, evaluate, and continuously improve evaluation harnesses for non-deterministic systems.
- Collaborate closely with product, research, and platform teams to translate ambiguous problems into shipped capabilities.
- Stay ahead of the rapidly evolving LLM ecosystem (models, frameworks, agentic patterns) and bring the best ideas into our stack.
Requirements:
Engineering Foundations:
- Strong Python skills- you write clean, idiomatic, well-tested code and understand the language deeply.
- Hands-on experience using coding agents(Cursor, Claude Code, GitHub Copilot, or similar) to build complex software systems. You know how to delegate effectively to AI assistants and review their output critically.
- Experience with multiple database paradigms- both SQL (PostgreSQL, MySQL) and NoSQL (MongoDB, Redis, DynamoDB, or similar). You can choose the right tool for the job.
- Experience designing and integrating with third-party APIs- REST and gRPC. Comfortable building robust clients, handling auth, retries, rate limits, and schema evolution.
- Production experience with Docker and Kubernetes- containerizing services, writing manifests, and debugging deployments.
- Strong Linux fundamentals- confident in bash and the terminal; you can navigate, script, and troubleshoot a server without reaching for a GUI.
- Experience building cloud-native tools on AWS, GCP, or Azure (compute, storage, queues, serverless, IAM).
AI / LLM Expertise:
- Solid understanding of what an LLM is and how it works- tokenization, attention, context windows, sampling, and the practical implications of each for system design.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8744445
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
06/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Improve observability and incident response - raise signal quality, cut alert noise, and own the monitoring the triage agents depend on.
Eliminate toil - relentlessly identify manual, repetitive operational work and remove it through automation and agents, protecting engineering time for reliability work that only humans can do.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, capacity risks, and automation opportunities.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Requirements:
Your Experience:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8771789
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
Your experience:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769584
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779592
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for people who are relentlessly curious and committed to continuous learning. AI is reshaping every function across our business, and we enable every team member, regardless of role or level, to build fluency in AI tools and concepts. Those who thrive here actively seek out new solutions, experiment thoughtfully, and apply what they learn to drive better, faster, smarter outcomes.
As a Senior Staff Software Engineer in the Detection Platform group, you will be tasked with being the technical authority responsible for defining and evolving the architecture of the cloud-native systems that power our AI SIEM detection, hunting, and response capabilities, including large-scale real-time detection engines, stateful detection engines, anomaly detections, ML pipelines, agentic SOC and threat-hunting capabilities. You will lead the design and execution of backend systems that process billions of events and several petabytes of data daily and serve tens of thousands of security specialists at enterprise and government customers worldwide. Your technical leadership will bridge long-term architectural strategy and high-velocity product delivery, and you will drive cross-team initiatives that shape how detection and response are built and operated across the group.
Requirements:
10+ years of software engineering experience with deep production-level mastery of Go and/or Java (Python a plus), and a strong track record of building and operating high-scale distributed backend services.
A track record of being a recognized subject-matter expert others seek out to review and elevate their designs, with a passion for building high-scale distributed systems.
Platform thinking: proven experience building and evolving platforms, not just features, with a focus on API design (gRPC, REST), service boundaries, multi-tenancy, and shared infrastructure in a high-scale SaaS environment.
Strong background in distributed data processing and microservices, building high-quality, scalable data products that handle millions of events per second.
Deep experience with AWS and/or GCP, Kubernetes, Docker, Postgres, Redis, Kafka, Cassandra, and ClickHouse.
Hands-on experience leveraging AI in the development process (e.g. AI coding assistants and agentic dev tools such as Claude Code, Cursor, or Copilot) and a desire to reshape how a team builds software to better utilize AI.
Experience embedding AI into production services, building agentic and LLM-powered capabilities. Familiarity with modern techniques such as agentic frameworks and orchestration, retrieval-augmented generation (RAG), the Model Context Protocol (MCP), vector databases, prompt engineering, and evaluation and guardrail frameworks for reliable AI systems is a strong advantage.
The ability to turn vaguely specified, complex requirements into efficient, future-proof end-to-end designs, and to drive multi-team initiatives and influence the engineering roadmap.
Strategic communication: able to articulate complex technical trade-offs to both technical and non-technical stakeholders, including Product Management, Directors, and VPs.
Ability to swiftly delve into new products, and to collaborate effectively with local and remote teams across time zones.
Customer focus: you care about delivering value and want to hear directly from customers on how to evolve your systems.
Previous experience developing security-related products is a strong advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774178
סגור
שירות זה פתוח ללקוחות VIP בלבד