דרושים » ניהול ביניים » Senior DevOps Engineer

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 15 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Our platform turns paperwork-based processes into revenue-generating customer engagements for some of the largest financial institutions in Israel, the US, Europe, and APAC.

Behind that surface, our infrastructure is split the way our customers are. Enterprise customers demand isolation and get it: a dedicated single-tenant environment each, down to their own compute, data stores, identity provider, and sometimes encryption keys we cannot access. Lower-tier customers run on shared multi-tenant infrastructure. More than fifty production environments across six AWS regions, all from one GitOps pipeline.

That split is what makes this role heavy. On the shared platform, a single change is instantly global; on the dedicated fleet, it must land correctly everywhere before anyone feels it. Every new customer adds load and cost in a straight line - your job is to break that line, making the platform carry its own weight through self-service, automation, and AI agents, not more hands. Nobody sits between you and the customer.

Responsibilities

Both tenancy models, as one system. Terraform, Helm, and ArgoCD delivering to shared and dedicated environments alike. You own the pipeline and the guardrails on it: progressive rollout, blast-radius containment, and drift detection that catches a mistake before a customer does.
Efficiency and cloud cost. Dedicated infrastructure per customer makes unit economics an engineering concern, not a quarterly finance exercise. You own attribution, right-sizing, spot strategy, and autoscaling that actually scales down - and you defend efficiency as a design constraint on every new deployment.
Turning operations into product. Onboarding, version rollouts, and config changes are tickets today; they should be self-service and boring. You decide what gets automated, what goes back to the owning team, and what still needs an expert.
Scale under real load. Temporal, KEDA queue-based autoscaling, and Karpenter are live and still maturing. You own how the platform grows and shrinks against real traffic, and what that costs per customer.
Security engineering for regulated customers. KMS-based key management, SIEM pipelines, WAF policy, certificate and secret lifecycle, per-customer SSO, vulnerability remediation, and turning penetration test findings into shipped fixes. Our customers are banks, insurers, and health funds. Their auditors are effectively part of our roadmap.
AI-driven operations. Our infrastructure repositories are already built for agentic work, with an internal harness of shared skills and agents. You inherit it, extend it, and push it into places it does not reach yet: triage, rollout verification, cost anomaly investigation, log analysis, and the long tail of operational work that has never been worth a script.
Reliability of what you ship. You are on call for your own systems, paged through PagerDuty against alerts you own.
Requirements:
6+ years in DevOps, including production ownership of Kubernetes at meaningful scale. EKS preferred.
Operating a fleet, with infrastructure as code. You have run many similar-but-not-identical environments with Terraform, and felt what happens when a change has to reach all of them. You have designed modules other people build on, and you have recovered state when it went wrong.
GitOps in production. You have operated ArgoCD, or an equivalent GitOps controller, running automated sync with pruning and self-healing enabled, and you understand exactly how much power that hands to a single commit.
Deep AWS and Cloudflare as a production edge. You have run AWS at the account, network, and IAM level rather than consumed it, and you have operated Cloudflare in production: WAF policy, origin protection, DNS, and rate limiting.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8845545
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
24/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
is looking for a DevOps engineer to own our infrastructure end to end.
You would be our only one. There is no platform team to join and no one to hand you a ticket queue. There is plenty of work, and most of it is on this page. You will set the technical direction as we grow, report to the Head of Engineering, and work directly with every engineer on the team.
We run on AWS and Kubernetes and deploy through GitOps. Our use of AI in engineering is growing quickly, and the infrastructure behind it has not caught up yet. Building that is a real part of this job.
What Youll Do:
Own all of our infrastructure. Kubernetes, AWS, and the core services around them. The foundation was built properly and the team has kept it running between other priorities. You are not inheriting a mess. What it has not had is a full-time owner. Reliability, cost, security posture, and the roadmap all become yours.
Own CI/CD. Every services build and deploy pipeline, and the self-hosted runner infrastructure underneath. We ship fast and we are working on shipping faster. One of our engineers is automating more of it right now, and you will take that over and push it further. The interesting problem is adding the safety a growing company needs without turning deploys into something people dread.
Help us shape the infrastructure behind our AI development. Our engineers use coding agents every day and that use keeps growing. We do not have a clear path for the infrastructure behind it yet, and you will help us find one. One direction we are exploring is moving agents off laptops and into the cloud, with a sandbox environment per engineer so an agent keeps working after the laptop is closed. Deciding whether that is the right answer, and building what we choose, is part of this job.
Scale how we deploy to customers. Some customers need us to run inside their own environment. We already do this, but each deployment takes too much hands-on work. You will make it repeatable, faster, and secure, and give our engineers the tooling to support customers directly.
Own our observability. Metrics, logs, alerting, and dashboards. Decide what is worth waking someone for, and keep the cost of watching under control.
Own infrastructure security and support our compliance work. Hardening, IAM, secrets, and vulnerability management. You will help us keep our SOC 2 compliance, which customers require before they will work with us, and help answer their security reviews. We sell to security teams, so they ask hard questions.
Keep our data recoverable. Backups, restore testing, and a disaster recovery plan for RDS Postgres and ClickHouse that we have actually rehearsed.
Partner with engineers on solutions, not tickets. You are the person developers come to when they need infrastructure to do something new.
דרישות:
6+ years in DevOps, infrastructure, or platform engineering.
Real production Kubernetes experience, preferably EKS.
Deep AWS knowledge. VPC, IAM, and networking, at a level where you can explain the trade-offs behind your design.
Infrastructure as code. Terraform or Terragrunt.
Ownership of CI/CD pipelines, GitHub Actions preferred.
Hands-on work with AI coding agents such as Claude Code or Codex, as part of your daily workflow.
A security-first mindset. You think about blast radius and least privilege by default, not after review.
Clear communication. You will work with every engineer here, and sometimes with our customers. You can explain a technical trade-off to someone who does not share your background.
Accountability. You will be the only person who owns Tokens infrastructure, and that is what you want.
Nice to Have
GCP or Azure alongside AWS.
GitOps with ArgoCD, and writing Helm charts.
Python, or another language you use to build real internal tooling. Our services are Python.
ClickHouse, or another column store used for high volume log and event data.
RDS Postgres in production. Tuning, upgrades, and failover.
Experience building המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8832847
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
our company's platform runs inside the security stacks of enterprise customers, processing live detection and telemetry data from SIEM, EDR, and SOAR systems at real scale. Until now, infrastructure has been built and carried by the founding engineering team alongside everything else. You will be the first dedicated owner of it. That means designing our company's cloud foundation across GCP and AWS, building the deployment, scaling, and observability layers from the ground up, and setting the standards the rest of engineering builds on. It also means holding a bar that most startups can defer and our company cannot: we sell to security teams who audit their vendors seriously, so how we run our own infrastructure is part of how we win. This is a hands-on senior role with real architectural authority, and as engineering grows you become the person others learn cloud and infrastructure from.
What you'll be doing
Multi-cloud architecture: Own the design and implementation of our company's multi-cloud architecture across GCP and AWS, including the cross-cloud standards, boundaries, and tradeoffs that keep two providers from becoming twice the complexity.
Infrastructure as code: Build and maintain the entire environment in Terraform, so infrastructure is reviewable, reproducible, and scales without a person in the loop.
Deployment pipelines from scratch: Create the CI/CD foundation that lets a small team ship quickly and safely, with the testing, gating, and rollback paths that make fast releases boring rather than risky.
Kubernetes at data scale: Run and scale the Kubernetes and containerized workloads behind our company's security data processing, where volume, latency, and cost pressure all show up at once.
Security of our own stack: Own identity, access, secrets, network boundaries, and infrastructure hardening to a standard that holds up under customer security review.
Reliability and visibility: Build the logging, monitoring, and alerting layer that tells us something is wrong before a customer does, and make on-call sustainable as the system grows.
Technical direction: Partner directly with the founding engineering team on architectural decisions, and raise the infrastructure and cloud fluency of every engineer who joins after you.
Requirements:
Senior infrastructure depth: 7+ years as a DevOps or Site Reliability Engineer, with production systems you built and carried, not only inherited.
Multi-cloud experience: Hands-on work with core services across both GCP (GCE, GKE, Cloud Storage, VPC) and AWS (EC2, EKS, S3, RDS), and clear judgment on when running in two clouds is worth the cost.
Kubernetes and containers: Strong production experience with Docker and Kubernetes, including scaling, resource management, and debugging clusters under real load.
Infrastructure as code: Proven Terraform experience (or Pulumi) in complex, multi-provider environments, with a preference for codified over manual.
Pipelines and scripting: Built CI/CD from zero with GitHub Actions, GitLab CI, or Jenkins, backed by strong Python or Bash.
Security instincts: A DevSecOps mindset and real exposure to security best practices. Background in a cybersecurity or otherwise security-sensitive environment is a strong advantage.
0-to-1 ownership: Comfortable being the first and only person in a domain, making decisions without an existing playbook, and staying close to the work while setting direction for others.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837817
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Requirements:
Your Experience:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
Collaboration - You work effectively across engineering, security, product, and leadership to align on reliability priorities and drive shared outcomes.
Innovation balanced with pragmatism - You actively explore new approaches, particularly AI-assisted operations and agent supervision, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate reliability, risk, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834235
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the Technology team might be right for you. We're looking for a Senior DevOps Engineer with a strong security orientation to join our DevOps team. our platform runs at significant scale, and our DevOps team sits at the core of keeping it fast, reliable, and secure. As we grow, so does the scope of what we own - we need an engineer who can step in as a strong pillar of our production ecosystem. This is a hands-on role for someone who takes security seriously, moves fast, and knows how to get things done in a complex, high-scale environment. Our Technology Stack: AWS, GCP, CLoudFlare,Kubernetes, Terragrunt, Ansible, Jenkins, ArgoCD, Argo Workflows, Kong & Nginx, HashiCorp Vault, Kafka, RabbitMQ, Mongodb, Aurora Postgresql & Mysql, Prometheus, Grafana, VictoriaMetrics Programming languages: Python, NodeJS, Go, Kotlin


What am I going to do?:

* Full Ownership: Drive infrastructure initiatives through their entire lifecycle, taking accountability from initial design to delivery and long-term operations.
* Kubernetes Orchestration: Architect, implement, and maintain production-grade Kubernetes clusters, ensuring they remain scalable and resilient under high-scale demand.
* Cloud Architecture: Design and manage robust AWS environments, including VPC, IAM, and EKS, to support a highly available platform architecture.
* Infrastructure as Code: Utilize Terraform to build and evolve our environment, applying configuration management principles to all IaC workflows.
* CI/CD Excellence: Support and improve our deployment pipelines using Jenkins and GitHub Actions to maintain a fast development velocity.
* Observability: Implement comprehensive monitoring solutions with Prometheus and Grafana to ensure deep visibility into platform health.
* Operational Resilience: Join the DevOps on-call rotation, taking responsibility for mitigating production issues and maintaining site reliability.
* Tooling & Innovation: Continuously evaluate and adopt tools - security and otherwise - that raise the bar on engineering efficiency and security posture.
Equal opportunities:
We're not about checklists. If you don't meet 100% of the requirements for this role but still feel passionate about the position and think you have the right skills and qualifications to excel at it, we want to hear from you. We prioritize diversity. We celebrate difference and embed it into every aspect of our workplace and product, as well as our community. We are proud and committed to providing equal opportunity employment to all individuals regardless of race, color, religion, sex, sexual orientation, citizenship, national origin, disability, Veteran status, or any other characteristic protected by law. In addition, we will provide accommodation to individuals with disabilities or a special need.
Requirements:
* 6+ years of hands-on DevOps / Platform Engineering experience in large-scale production environments on a public cloud (AWS preferred).
* Proven leadership mindset - able to own projects end to end and be accountable for outcomes.
* Strong, production-grade Kubernetes experience across design, deployment, scaling, and troubleshooting.
* Solid AWS experience with VPC, IAM, EC2, EKS, Load Balancers, and DNS.
* Experience designing and operating highly available, scalable infrastructure systems.
* Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis).
* Hands-on
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8799862
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
08/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Platform and infrastructure (K8s + IaC)
Architect and operate our Kubernetes platform on AWS - scaling, networking, cost, and reliability
Own infrastructure as code (Terraform) so environments are reproducible, auditable, and fast to evolve
Build the foundations that let the team provision and scale services without friction


AI agent platform - managing and scaling agents at volume
Operate and scale the orchestration layer (Temporal) that runs dozens of agents in parallel
Tune the platform for the unusual load profile of agent workloads - bursty, long-running, data-heavy, latency-sensitive
Give engineers the primitives to deploy, version, observe, and roll back agents safely under enterprise SLAs


CI/CD and developer experience
Own build and deploy pipelines end-to-end - fast, safe, boring releases
Invest in DX as a first-class product: the team treats developer experience as leverage, and you set the bar
Reduce the time from merged to in production and from idea to running experiment


Observability and reliability
Build monitoring, alerting, and tracing that make production legible - for services and for agents
Own incident response and the reliability practices that keep enterprise customers SLAs intact
Turn incidents into systemic fixes, not repeated firefighting


Security and compliance
Own secrets management, hardening, and the day-to-day security posture of the platform
Support our compliance commitments (we are Mastercard-certified, operating in fintech - the bar is high)
Build security into the pipeline so it is the default, not a gate
Requirements:
Strong DevOps/platform/SRE experience at a company with a real engineering culture (FAANG, unicorn, or a well-established startup with high standards)
Deep Kubernetes and AWS - you have architected, scaled, and debugged production clusters, not just deployed to them
Infrastructure as code in your bones - Terraform or equivalent, with strong opinions on reproducibility and auditability
CI/CD ownership - you have built pipelines that engineers trust and rarely think about
Production observability and incident response at meaningful scale
Comfort with and curiosity about AI/LLM workloads. We are an AI-native company; you should be using AI in how you work, and excited to operate the infrastructure agents run on
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8814945
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
28/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Senior Platform Engineer to join our CTI Group. We're a startup that builds a lot, so you'll work across a wide range of systems and tools, build new platforms and services, and have real influence over where the platform goes next.
You'll work on the platform behind our CTI products: the microservices running on Kubernetes, the cloud and on-prem environments around them, and the infrastructure behind the systems. There's a lot still to build, and you'll have a real say in what comes next.
Responsibilities
Own what you build end to end, from design through implementation to running it in production.
Design and build the internal platforms, tools, and services other engineering teams rely on, and add new ones as we grow.
Make architectural decisions across infrastructure, applications, networking, security, and data.
Design, build, and maintain our cloud and on-prem environments.
Build and operate the infrastructure behind our AI/ML systems at scale.
Manage and evolve our Kubernetes environments.
Build and maintain CI/CD pipelines, deployment automation, and Infrastructure-as-Code with Terraform.
Build security into the platform, from IAM and secrets management to network policies.
Set up the monitoring and observability that show us how the platform is doing.
Design for scalability, availability, performance, and resilience in high-scale production systems.
Apply Infrastructure-as-Code practices (Terraform) end to end.
Requirements:
5+ years in DevOps or Platform Engineering roles.
Strong software engineering skills, including production-grade Python development using OOP, and proficiency in Bash scripting.
Experience designing and building platforms and infrastructure from scratch.
A strong architectural sense: you understand how infrastructure, applications, networking, security, and data fit together.
Hands-on AWS experience.
Solid Kubernetes experience, including deploying, operating, and troubleshooting clusters.
Experience with Terraform and Docker, and comfort with Linux, networking, and security fundamentals.
Experience with high-scale production systems and with the infrastructure behind AI/ML workloads.
A startup mindset: hands-on, curious, and comfortable with broad scope and a lot of room to shape your own work.
Bonus Points:
Experience in the cyber security field.
On-prem, customer private cloud, or air-gapped deployment experience.
Model serving and ML tooling such as MLFlow, and GPU scheduling on Kubernetes.
Observability stacks such as Prometheus, Grafana, or Loki.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836113
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time and Hybrid work
Required Senior DevOps Engineer
As a Senior DevOps Engineer, you'll design, operate, and evolve the platform that runs our production environment across AWS, GCP, and Azure, powering hundreds of Java microservices, ML model-serving workloads, and the internal tooling our R&D organization depends on every day.
You'll own:
A multi-cloud, multi-region Kubernetes platform managed as code via Crossplane, Helm, and ArgoCD
CI/CD pipelines (Jenkins, GitHub Actions, Argo Workflows) that deploy, test, and scale workloads across environments
Monitoring, logging, and alerting across the stack (Datadog, Coralogix, Sentry, Wiz), and the on-call rotations that keep it healthy
You'll solve:
High-scale production problems: driving continuous platform and security upgrades across the fleet, under real traffic, without impacting live tenants
The velocity-vs-safety tradeoff: designing GitOps workflows where a single PR can safely impact production across clouds and region
Advanced troubleshooting when production, staging, or dev environments break in unexpected ways
You'll impact:
Every engineer. Our platform is the substrate they deploy on, so your work directly affects their delivery speed
Cost, reliability, and security of the infrastructure that underpins the product our customers pay for
The resilience of our customer experience. The platform you build determines whether thousands of revenue teams can trust us to be there when they need it
The direction of the platform itself. You'll have the autonomy to challenge existing architecture and drive meaningful changes, not just maintain what's already there.
Requirements:
5+ years in DevOps or platform engineering, with deep AWS experience and proven ownership of high-scale production systems (reliability, performance, and cost optimization)
Strong hands-on with containers and orchestration (Docker, Kubernetes, Helm) in real production, not just demos
Fluent in IaC and GitOps (Terraform, ArgoCD) and CI/CD tooling like Jenkins, GitHub Actions, or Argo Workflows
Solid grasp of cloud security and architecture best practices: building systems that are resilient, scalable, and cost-efficient by default
Comfortable with monitoring and logging stacks (Datadog, Coralogix, Prometheus, Grafana): you design for observability, not bolt it on after
Scripting fluency in Python, bash, or Node.js to automate anything that shouldn't be done twice
Linux at a systems level: you can troubleshoot DNS, iptables, or systemd issues without hesitation
Clear communicator who collaborates across R&D, Security, and Product, documents decisions, and follows through on delivery, from RFC to rollout
Bonus: experience with big-data infrastructure, Java/JVM services, or multi-cloud (GCP/Azure).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8838170
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
we are a leader in large-scale networking solutions for AI infrastructure and service providers. The company's disaggregated networking architecture transforms the economics of large-scale infrastructures while maximizing performance, utilization, and operational efficiency. Its high-performance AI fabric maximizes GPU utilization and accelerates deployments by optimizing the AI stack end-to-end, resulting in higher tokens-per-second and lower cost-per-token. our company's solutions power production networks for global tier-1 operators like AT&T and Comcast, and scale multi-vendor AI infrastructures at foundation model labs, NeoClouds, and enterprises.
our company's Cybersecurity business unit is building an AI-powered security platform for the world's largest service provider networks. The platform combines a cloud-native microservices backend with self-hosted LLM inference, delivered into customers' own Kubernetes environments - often air-gapped, with no data leaving their premises.
Responsibilities
Own the platform's Helm charts and make them portable across any Kubernetes distribution, including OpenShift (non-root, restricted SCCs, NetworkPolicies).
Operate the self-hosted operator stack: CloudNativePG, Strimzi, ECK, MinIO, Redis, and Temporal.
Run the Nx monorepo pipelines and drive the ArgoCD GitOps rollout.
Build zero-downtime releases using blue/green and canary deployments, expand/contract migrations, and automated rollback.
Manage autoscaling with KEDA/HPA, backup/restore for stateful services, and load-validation of stamp sizing tiers.
Provide on-call support for dev/staging environments and customer stamps.
Build and maintain observability using Grafana Alloy, OpenTelemetry, and Prometheus/Grafana - dashboards and alerts across services, Kafka, databases, and GPUs.
Improve developer experience through a shared dev cluster with local-debug traffic interception (Telepresence / mirrord), targeting sub-15-minute onboarding.
Requirements:
Technical Skills
5+ years of DevOps / platform engineering experience in a product company.
Deep hands-on experience with Kubernetes and Helm, including shipping software onto clusters you don't control.
Production experience with GitOps (ArgoCD/Flux), infrastructure as code, and zero-downtime release engineering.
Experience running stateful workloads on Kubernetes via operators (Kafka, PostgreSQL, Elasticsearch).
Solid grasp of observability fundamentals: Prometheus, Grafana, OpenTelemetry.
Comfortable operating in air-gapped environments without managed cloud services.
Soft Skills
Strong cross-functional collaboration - works daily with AI/ML Ops, backend, research, and solutions teams on customer onboarding.
Comfortable owning production reliability and on-call responsibilities.
Nice to Have / Advantage
Experience with OpenShift, KEDA, and Temporal.
Hands-on AI inference infrastructure experience: deploying and operating LLM inference on Kubernetes with GPUs (vLLM, TGI, Triton, or similar).
GPU scheduling and node readiness, model packaging for offline delivery.
Experience with LLM gateways (LiteLLM) and LLM observability (Langfuse).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8828041
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. You'll define and operate our SLOs and error budgets, lead high-severity incident response, and ensure our Kubernetes and AWS infrastructure stays stable under growth. You'll also build and supervise the AI agents that handle routine alert triage and monitor tuning, focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership, incident command, and applying AI to operational work.
Your Impact:
Own reliability as an engineering discipline - define SLIs, set SLOs, and run error-budget-based decision-making so "how reliable are we" becomes a number that governs how fast we ship.
Own production incidents end-to-end - lead response, mitigation, and resolution for high-severity incidents, and drive blameless postmortems that feed real fixes back into the system.
Own the reliability and capacity of production infrastructure as we scale - forecasting headroom, validating scaling behavior under load, and keeping latency and error rates within SLO.
Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
Own, build, and supervise our SRE AI agents that triage alerts, review monitors, resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously, what needs approval, and what stays human. This is a core part of the role.
Improve observability and incident response - raise signal quality, cut alert noise, and own the monitoring the triage agents depend on.
Eliminate toil - relentlessly identify manual, repetitive operational work and remove it through automation and agents, protecting engineering time for reliability work that only humans can do.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, capacity risks, and automation opportunities.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
דרישות:
5+ years operating production cloud infrastructure, with a strong reliability focus (SRE, or DevOps/platform engineering with reliability ownership).
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Experience defining and operating SLIs, SLOs, and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
Strong observability and alerting experience in Datadog or comparable platforms, including raising signal-to-noise in production.
Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
Proven ability to own platform and reliability projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work - and in building and supervising agents, not just using them.
Ownership-driven - You take responsibility for the reliability of the systems you build and operate, from SLO definition through incident command and continuous improvement.
Reliability as engineering - You treat reliability as a software problem to be solved with code, measurement, and automation - not an ops queue to be worked by hand.
Collaboration - You work effectively across engineering, security, product, and leadership to align on reliability priorities#ENGLI המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8833908
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do at our company and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the company Technology team might be right for you.
We're looking for a Senior DevOps Engineer with a strong security orientation to join our company's DevOps team.
our company's platform runs at significant scale, and our DevOps team sits at the core of keeping it fast, reliable, and secure. As we grow, so does the scope of what we own - we need an engineer who can step in as a strong pillar of our company's production ecosystem.
This is a hands-on role for someone who takes security seriously, moves fast, and knows how to get things done in a complex, high-scale environment.
our company's Technology Stack:
AWS, GCP, CLoudFlare ,Kubernetes, Terragrunt, Ansible, Jenkins, ArgoCD, Argo Workflows, Kong & Nginx, HashiCorp Vault, Kafka, RabbitMQ, Mongodb, Aurora Postgresql & Mysql, Prometheus, Grafana, VictoriaMetrics
Programming languages: Python, NodeJS, Go, Kotlin
What am I going to do?
Full Ownership: Drive infrastructure initiatives through their entire lifecycle, taking accountability from initial design to delivery and long-term operations.
Kubernetes Orchestration: Architect, implement, and maintain production-grade Kubernetes clusters, ensuring they remain scalable and resilient under high-scale demand.
Cloud Architecture: Design and manage robust AWS environments, including VPC, IAM, and EKS, to support a highly available platform architecture.
Infrastructure as Code: Utilize Terraform to build and evolve our environment, applying configuration management principles to all IaC workflows.
CI/CD Excellence: Support and improve our deployment pipelines using Jenkins and GitHub Actions to maintain a fast development velocity.
Observability: Implement comprehensive monitoring solutions with Prometheus and Grafana to ensure deep visibility into platform health.
Operational Resilience: Join the DevOps on-call rotation, taking responsibility for mitigating production issues and maintaining site reliability.
Tooling & Innovation: Continuously evaluate and adopt tools - security and otherwise - that raise the bar on engineering efficiency and security posture at our company.
Requirements:
6+ years of hands-on DevOps / Platform Engineering experience in large-scale production environments on a public cloud (AWS preferred).
Proven leadership mindset - able to own projects end to end and be accountable for outcomes.
Strong, production-grade Kubernetes experience across design, deployment, scaling, and troubleshooting.
Solid AWS experience with VPC, IAM, EC2, EKS, Load Balancers, and DNS.
Experience designing and operating highly available, scalable infrastructure systems.
Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis).
Hands-on experience with Infrastructure as Code and configuration management (Terraform required; Terragrunt and Ansible a plus).
Experience with Docker and containerized workloads.
2+ years building and maintaining CI/CD pipelines (Jenkins, GitHub Actions)
Experience with monitoring and observability tools (Prometheus, Grafana).
Experience with GenAI platforms (AWS Bedrock, Vertex AI, OpenAI).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8806869
סגור
שירות זה פתוח ללקוחות VIP בלבד