דרושים » מחשבים ורשתות » Senior DevOps Engineer - Data Platform (Cortex)

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 4 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Senior DevOps Engineer with a strong Developer Experience (DevEx) mindset, someone who genuinely enjoys making developers more productive by building and improving the tools, infrastructure, automation, and workflows they use every day.
You will work across CI/CD, developer tooling, automation, infrastructure, observability and productivity. Your primary focus will be making development and delivery faster, simpler, more reliable, and more cost-effective.
We need someone who can independently identify developer pain points, proactively discover opportunities for improvement, and turn them into scalable engineering solutions. You will be expected to measure the impact of your work and continuously improve critical metrics such as CI feedback time, pipeline reliability, deployment time, infrastructure cost, and overall developer productivity.
This role is a perfect fit for someone who thinks far beyond just "keeping the infrastructure running" and is excited about revolutionizing the entire engineering experience.
Your Impact:
You will tackle high-impact engineering challenges across our platform. Representative areas you will own and work on include:
Optimize CI/CD Pipeline Velocity & Stability
Improve CI Stability: Leverage AI to automatically detect recurring test failures, dynamically disable problematic tests, and open Jira tickets for follow-up.
Automate Workflows: Utilize AI to automate merge conflict resolutions and keep the merge train moving smoothly.
Accelerate CI Velocity: Architect workflows to parallelize work both across jobs and within jobs. Introduce highly effective caching strategies and local Artifactory mirrors to drastically reduce build and dependency-fetch times.
Optimize CD: Reduce container image sizes and accelerate deployment times. Partner with QA to introduce and streamline automated regression testing against tenants.
Drive Repository Security & Maintenance
Improve Security Posture: Automate the identification and remediation of Go and Python dependencies with known CVEs.
Automate Maintenance: Design systems to automatically handle Go and Python version upgrades across the codebase.
Ensure Performance: Automate performance testing to proactively detect and alert on performance regressions before they hit production.
Cloud Infrastructure & Cost Engineering
Reduce Infrastructure Costs: Identify and implement innovative opportunities to reduce overall infrastructure and resource consumption.
Resource Efficiency: Introduce mechanisms to make our "slim tenants" even more resource-efficient for development and testing.
Smart Observability: Re-architect our logging strategy by streaming GCP logs to more cost-effective sinks instead of relying exclusively on Cloud Logging.
IaC Governance: Act as a gatekeeper for infrastructure quality by reviewing Terraform Merge Requests (MRs) for correctness, efficiency, maintainability, and strict adherence to engineering standards.
Requirements:
Your Experience:
5+ years of hands-on experience in DevOps, Platform Engineering, or SRE roles, preferably within data-heavy or high-scale environments.
The DevEx Mindset: A genuine passion for your job and a drive to build tools that make other engineers happier and more productive.
Coding & Automation: Strong programming skills in Python and Go. You treat infrastructure and automation as software engineering disciplines.
Cloud & IaC Mastery: Deep expertise in GCP (Google Cloud Platform) or AWS and advanced proficiency in Terraform.
CI/CD & Containers: Extensive experience building highly optimized CI/CD pipelines and optimizing Docker/Kubernetes deployments.
Forward-Thinking: An interest in or experience with leveraging AI tools to solve traditional operational bottlenecks.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834394
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an independent Senior DevOps Engineer who thrives in a fast-paced environment. You will join our centralized DevOps team, serving as a pillar of reliability and innovation for the entire engineering organization.

This is a 50/50 "Build vs. Run" role. Splitting your time between building modern automation and improving our developer experience and operational excellence, ensuring our production environments are rock-solid, secure, and highly available while managing the live environment from server infrastructure all the way to the application.

Tasks you will take part in:
Design and implement next-generation CI/CD automation flows to streamline the path from "Code" to "Production."
Lead the evolution of our Infrastructure-as-Code (Terraform) to make our systems more modular and self-service for developers.
Collaborate with backend teams to architect scalable microservices and optimize datastore performance (RDS, MongoDB).
Embrace AI-powered development tools (Cursor, Claude Code) to automate toil and accelerate infrastructure delivery.
Full ownership of our AWS accounts and multi-cluster Kubernetes environments.
Manage and fine-tune live production environments, ensuring 99.9% uptime for our mission-critical fintech services.
Drive our Observability strategy (Datadog/Grafana) to proactively catch issues before they impact our customers.
Participate in a weekly on-call rotation, serving as the first line of defense for platform reliability.
is used for creating requisitions 'from scratch'
Requirements:
Requirements:
5+ years of experience managing high-scale production systems.
Mastery of AWS: Deep knowledge of AWS services, security best practices, and account management.
Orchestration & Containers: Expert-level experience with Kubernetes (EKS) and Docker.
Automation-First Mindset: Proficiency in Terraform and at least one high-level language (Python, Go, or Bash).
Database Ops: Solid understanding of the DevOps aspects of MySQL/RDS and MongoDB (scaling, backups, performance tuning).
SRE Discipline: Strong experience with monitoring and logging stacks like Datadog, Grafana, or Prometheus.
Team Player: Strong communication skills and a "can-do" attitude-you enjoy solving problems as part of a collaborative squad.
Experience with AI tools and a strong interest in continuously exploring and applying them in everyday work are highly valued.

Advantages:
Security Knowledge: Experience with FW, WAF, IPS/IDS, and SELinux.
Data Infrastructure: Familiarity with Data tools like Kafka, Spark, Snowflake or Data Lakes.
Multi-Cloud: Familiarity with hybrid or multi-cloud architectures.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8831065
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
5 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a Senior Infrastructure Engineer who views "Infrastructure as Software." In 2026, we dont just manage servers; we build high-performance environments that allow multi-agent systems to operate at scale.
You will be a core member of the R&D team, blending deep DevOps expertise with the coding rigor of a Backend Engineer. You arent just "configuring" AWS; you are architecting the distributed systems and data pipelines that power our autonomous security brain. Your mission is to ensure that while our agents are evolving and taking actions, our underlying platform remains immutable, observable, and infinitely scalable.
What You'll Do
Design, build, and operate our company's cloud infrastructure using AWS, Kubernetes, and Infrastructure as Code.
Build internal tools and platform services using Python and Go to improve developer productivity and system reliability.
Own infrastructure automation with Terraform, Pulumi, and modern cloud-native tooling.
Partner closely with Backend, Data Science, and Security Engineering teams to build scalable, reliable platforms.
Improve observability, monitoring, and incident response across distributed production systems.
Design and optimize infrastructure for performance, scalability, security, and cost efficiency.
Help shape engineering best practices, platform architecture, and developer experience as our company continues to grow.
Requirements:
5+ years of experience in Infrastructure, DevOps, Platform Engineering, or Backend Engineering.
Strong software engineering skills with hands-on experience building production systems in Python or Go.
Deep hands-on experience with AWS, including services such as EKS, RDS, VPC, and IAM.
Strong experience designing, operating, and scaling production Kubernetes environments.
Experience with Infrastructure as Code, CI/CD, GitOps, and modern cloud-native development practices.
A systems mindset with the ability to solve architectural challenges across infrastructure and application layers.
Comfortable using modern AI-powered developer tools and agentic workflows to improve engineering productivity.
The company Mindset: You take ownership, act with accountability, collaborate openly, and focus on delivering meaningful impact. You thrive in fast-moving environments, embrace ambiguity, and enjoy solving hard problems together.
Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.
Full professional fluency (written and verbal) in both Hebrew and English.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8829008
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Senior Front-End Infrastructure Engineer to help shape the foundation of our digital ecosystem across multiple brands and products. You will build the infrastructure, tooling, and engineering capabilities that enable teams to move faster, release with confidence, and scale effectively. Your work will directly impact developer experience, platform reliability, performance, experimentation capabilities, and the overall success of our digital brands.
This role is focused on solving complex engineering challenges with organization-wide impact. We are looking for a developer with deep technical expertise, a platform mindset, and a strong sense of ownership, someone who understands how engineering decisions drive business outcomes and enjoys building solutions that improve how teams build, test, release, and scale software across the organization. You will partner closely with Engineering, Product, QA, Data, and DevOps teams to help shape the future of our frontend ecosystem.
Roles & Responsibilities:
Build and evolve the frontend infrastructure, shared tooling, reusable components, and engineering foundations that support multiple brands and product teams
Improve frontend performance through modern web architecture patterns, including server-side rendering, caching strategies, and request middleware
Drive initiatives that improve Core Web Vitals, performance, reliability, and user experience at scale
Own and enhance frontend CI/CD pipelines, release processes, and production readiness practices to enable fast, safe, and reliable deployments
Establish monitoring, observability, and incident response practices that improve platform reliability and production visibility
Build solutions for experimentation, feature flagging, and controlled rollouts to improve reliability and support data-driven decision making
Lead cross-functional technical initiatives, partnering with Engineering, Product, QA, Data, and DevOps teams to solve complex platform challenges
Act as a technical owner for strategic frontend infrastructure initiatives, leveraging modern technologies and AI-powered development workflows to drive long-term platform improvements and organizational impact
Provide technical guidance through code reviews, engineering standards, and knowledge sharing across teams.
Requirements:
5+ years of hands-on experience in frontend development, building high-performance and scalable web applications
Strong expertise in React, Next.js, and TypeScript
Deep understanding of modern front-end architecture, infrastructure, build systems, and deployment flows
Experience building reusable infrastructure, shared libraries, internal tools, documentation, and engineering standards that improve developer productivity and are adopted across multiple teams
Strong experience with Next.js capabilities such as SSR, SSG, ISR, routing, middleware, caching, and data-fetching strategies
Strong experience with CI/CD processes for front-end applications
Hands-on experience with testing frameworks and tools such as Jest, React Testing Library, Playwright, Cypress, or similar
Strong ownership mindset with the ability to independently drive technical initiatives from ideation through adoption
Familiarity with backend concepts such as Node.js, API performance, caching, CDN behavior, and service reliability
Experience leveraging AI-powered development tools and workflows to improve engineering productivity, code quality, and delivery speed
Experience with monitoring, analytics, and observability tools such as Datadog, Google Analytics, or similar
Strong communication skills, with the ability to work closely with developers, product managers, QA, DevOps, and leadership
Ability to mentor and guide developers, fostering a culture of learning, ownership, and technical excellence
Familiarity with experimentation platforms, feature flagging, analytics, and data-driven decision making - major advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8802494
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
26/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep Navan reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.



This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.



This position is based out of our new Tel Aviv office.



What You'll Do:

Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
Ability to troubleshoot AI tools and provider/API issues, including rate limits, quota, auth, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations.
Excellent communication skills and the ability to work with stakeholders and domain experts across the company.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8797911
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 5 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior Principal DevOps Engineer, you will serve as a visionary technical leader within the Cortex Cloud DevOps group. You will define the technical strategy and architecture that ensures our massive-scale production services remain highly reliable, exceptionally secure, and performant. You will pioneer the integration of AI-driven capabilities into our daily operations, establishing elite engineering standards and fundamentally transforming the workflows of hundreds of developers through autonomous agents and intelligent procedures.
Your Career
Architectural Vision & Scalability: Design and scale massive, resilient distributed systems and global Kubernetes infrastructure, implementing robust observability and monitoring frameworks.
AI-Driven Transformation: Revolutionize the SDLC by integrating Generative AI, autonomous agents, and LLM-powered workflows into CI/CD and self-healing systems to accelerate developer velocity.
Technical Leadership & IaC: Define architectural standards, lead GitOps/IaC (FluxCD/Terraform) strategies, and mentor Senior/Staff engineers across the R&D organization.
Developer Experience & Efficiency: Build and champion AI-powered platforms that automate troubleshooting and eliminate friction. Partner directly with Engineering Directors, Principal Architects, and Product Management to align infrastructure initiatives with business goals, optimizing for scale, high availability, and multi-million-dollar cost-efficiencies.
Security & Compliance: Embed "Security by Design" principles into the platform architecture to ensure platform integrity without sacrificing delivery speed.
Requirements:
Your Experience
10+ years of progressive experience in DevOps, SRE, Platform, or Infrastructure Engineering roles, with a significant portion at the principal/ tech leadership/ staff, or architectural level.
System Design from Scratch: A proven track record of designing, building, and deploying large-scale, highly available distributed systems and cloud platforms from the ground up.
AI-Powered Automation: Proven experience designing and integrating AI-driven systems, autonomous agents, and LLM-based tools into engineering workflows to optimize development processes, procedures, and overall organizational efficiency.
Communication: Exceptional interpersonal skills, capable of articulating complex architectural and AI workflow concepts clearly to both deeply technical peers and executive leadership.
Cloud & IaC Mastery: Expert-level proficiency with GCP (or equivalent major cloud providers) and deep architectural experience with Terraform.
Advanced Container Orchestration: Deep, internal knowledge of virtualized and containerized environments, with architectural-level expertise in scaling Kubernetes, extending it via custom operators, and automating complex operational logic.
Software Engineering Approach: Advanced coding and automation skills in Python or Go. You treat infrastructure as a software engineering discipline and can build custom tooling/services when off-the-shelf solutions fall short.
Proven Leadership: Demonstrated ability to lead complex, cross-team technical initiatives from conception to delivery, including setting technical roadmaps and driving consensus among stakeholders.
OS/Systems Expertise: Mastery of Linux systems, including kernel tuning, advanced networking, and performance troubleshooting.
Nice to Have:
Deep expertise in managing and scaling stateful workloads and distributed databases (e.g., Cassandra, ScyllaDB, MemSQL, or MySQL) in containerized environments.
Experience contributing to open-source infrastructure projects, or presenting at major tech conferences.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834263
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced and highly motivated Senior DevOps Engineer to join our engineering and DevOps team. As a Senior DevOps Engineer, you will be the architect of our infrastructure, ensuring that platform is scalable, resilient, and secure. This is a hands-on role where you will bridge the gap between development and operations, automating our deployment pipelines and managing our cloud-native ecosystem. This is an incredible opportunity to shape the foundational infrastructure of a high-growth data company.

What You'll Do

Design, implement, and manage our cloud infrastructure using tools like Terraform, ensuring environment consistency and scalability.

CI/CD Automation: Take full ownership of our deployment pipelines, optimizing for speed, reliability, and developer productivity.

Cloud Orchestration: Manage and scale our Kubernetes (K8s) clusters on AWS, ensuring high availability and efficient resource utilization.

Observability & Monitoring: Implement and maintain robust monitoring, logging, and alerting systems to ensure platform health and rapid incident response.

Security & Compliance: Drive security best practices across the infrastructure, including IAM management, network security, and vulnerability scanning.

Collaborate: Work closely with software engineers to optimize application performance, containerization strategies, and database reliability.
Requirements:
We are seeking a hands-on builder who can grow into a technical leader for our infrastructure. Someone excited to own hard problems and pick up the specifics of our stack quickly.

Must-Haves

5+ years of experience in DevOps or Site Reliability Engineering (SRE), ideally within a high-growth SaaS or data-heavy environment.

Hands-on experience running production workloads on AWS (e.g., EKS, RDS, S3, IAM, VPC).

Strong experience with Kubernetes and Docker in production environments.

Proficiency with CI/CD tooling (e.g., GitHub Actions, GitLab CI, or Jenkins) and infrastructure-as-code (e.g., Terraform).

Working proficiency in Python, Bash, or Go for automation and internal tooling.

A team player with strong communication skills who can explain infrastructure concepts to cross-functional partners.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8791545
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 5 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Requirements:
Your Experience:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
Collaboration - You work effectively across engineering, security, product, and leadership to align on reliability priorities and drive shared outcomes.
Innovation balanced with pragmatism - You actively explore new approaches, particularly AI-assisted operations and agent supervision, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate reliability, risk, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834235
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 19 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do at Fiverr and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the Fiverr Technology team might be right for you. Fiverr's Engineering team is expanding its DevOps capabilities to manage our rapidly growing cloud infrastructure. We're looking for a mid-level DevOps Engineer to join this high-velocity environment, taking ownership of cloud assets, contributing to production stability, and driving key projects. You'll be instrumental in strengthening our team and supporting our continued innovation.

What am I going to do?:

* Independently deploy and maintain robust cloud infrastructure across AWS, GCP, and CloudFlare, leveraging Kubernetes, Terragrunt, and Ansible.
* Contribute to an on-call rotation, ensuring the stability of critical production services including Kafka, RabbitMQ, and various databases, with a focus on proactive incident resolution.
* Lead small to mid-sized projects from conception through completion, utilizing your expertise in CI/CD pipelines (Jenkins, ArgoCD, Argo Workflows) and scripting languages like Python, NodeJS, Go, or Kotlin.
* Partner closely with security and development teams to embed security best practices throughout the entire software development lifecycle.
* Proactively evaluate and implement new tools and technologies to enhance engineering efficiency, security posture, and operational excellence.
* Manage sensitive information securely using tools like HashiCorp Vault and ensure secure configurations for services like Kong & Nginx.

Equal opportunities:
At Fiverr, we know that talent has no single face. We welcome talent from everywhere and everyone because it makes everything we build better. Need accommodations? Just ask. And if this role excites you but you don't tick every box, apply anyway. The best people rarely fit the mold exactly.
Requirements:
* 4+ years of hands-on DevOps / Platform Engineering experience in production environments within a public cloud environment (AWS preferred)
* Strong, production-grade Kubernetes experience (design, deployment, scaling, and troubleshooting) with solid AWS experience (VPC, IAM, EC2, EKS, Load Balancers, DNS)
* Experience designing and operating highly available, scalable infrastructure systems
* Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis)
* Hands-on experience with Infrastructure as Code and configuration management (Terraform required, Terragrunt & Ansible – advantage)
* Experience with Docker and containerized workloads
* 2+ years of experience building and maintaining CI/CD pipelines (Jenkins, GitHub Actions)
* Proficiency in Python for automation and strong Linux administration skills
* Experience with monitoring and observability tools (Prometheus, Grafana)
* Development experience and familiarity with GenAI platforms (AWS Bedrock, Vertex AI, OpenAI) – advantage Working with AI At Fiverr, AI is a powerful partner in our engineering workflows. You'll leverage AI-powered tools like GitHub Copilot for code assistance, intelligent security scanning tools, and automation platforms to streamline deployments and monitoring. While AI handles repetitive tasks and provides insights, your critical thinking, architectural decisions, and complex problem-solving remain at the forefront.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776958
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a hands-on DevOps Engineer with a strong cloud-native mindset to build, maintain, and evolve our highly scalable, highly-available cloud infrastructure. This role is pivotal in driving operational excellence, security, and automation across our entire engineering organization. You will promote communication, integration, and collaboration to significantly enhance our software development productivity and reliability. You'll work closely with engineering and product teams to streamline delivery, enforce platform standards, and enable a high-velocity development environment-all while keeping reliability and security top of mind.
Responsibilities:
Design, Automate, and Manage complex cloud infrastructure on AWS using best-in-class Infrastructure as Code (IaC) practices.
Lead the operation and enhancement of our production Kubernetes environments (EKS), focusing on automation, security, observability, and seamless CI/CD integration.
Drive continuous improvement across platform tooling, developer experience, and operational processes to meet our ambitious performance and uptime goals.
Implement and enforce security-first infrastructure patterns, including strong IAM, network segmentation, and secure secrets management.
Actively contribute to high-level technical design discussions and cross-functional architectural decision-making, ensuring solutions align with long-term platform strategy.
Drive AI-Powered Automation: Design, build, and deploy infrastructure for autonomous AI agents and agentic flows, integrating AI tooling directly into developer platforms and CI/CD pipelines.
Architect AI Platform Infrastructure: Establish reliable, scalable, and secure AI/DevOps orchestration frameworks to enable continuous deployment and runtime management of AI-driven tools.
Requirements:
4+ years of experience as a DevOps Engineer, Platform Engineer, or in a similar infrastructure-focused role.
Strong hands-on expertise across the AWS Stack (e.g. EC2, EKS, RDS, VPC, IAM, S3, Lambda).
Mastery of Infrastructure as Code - Terraform or equivalent.
Deep operational knowledge of Kubernetes, including architecture, cluster management, networking, and advanced debugging in production environments.
Strong expertise in designing and managing CI/CD methodologies and platforms (e.g. Jenkins, Github Actions).
Experience with monitoring tools such as Prometheus, DataDog, Coralogix (OTEL), Grafana etc.
Hands-on experience building and operationalizing agentic flows, AI agents, and DevAI automation within modern cloud environments.
Familiarity with AI orchestration frameworks and infrastructure patterns for scaling AI-driven operational workflows (an advantage).
Proven prior experience building and maintaining highly-available, production-grade, and service-oriented systems.
Strong scripting and automation background in languages such as Python or Bash.
Exceptional communication and collaboration skills with the ability to articulate complex technical needs and influence cross-functional teams.
Strong knowledge of AWS Networking - an advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8805583
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 7 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Improve observability and incident response - raise signal quality, cut alert noise, and own the monitoring the triage agents depend on.
Eliminate toil - relentlessly identify manual, repetitive operational work and remove it through automation and agents, protecting engineering time for reliability work that only humans can do.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, capacity risks, and automation opportunities.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
דרישות:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
Collaboration - You work effectively across engineering, security, product, and leadership to align on reliability priorities#ENGLI המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8833908
סגור
שירות זה פתוח ללקוחות VIP בלבד