Our platform turns paperwork-based processes into revenue-generating customer engagements for some of the largest financial institutions in Israel, the US, Europe, and APAC.
Behind that surface, our infrastructure is split the way our customers are. Enterprise customers demand isolation and get it: a dedicated single-tenant environment each, down to their own compute, data stores, identity provider, and sometimes encryption keys we cannot access. Lower-tier customers run on shared multi-tenant infrastructure. More than fifty production environments across six AWS regions, all from one GitOps pipeline.
That split is what makes this role heavy. On the shared platform, a single change is instantly global; on the dedicated fleet, it must land correctly everywhere before anyone feels it. Every new customer adds load and cost in a straight line - your job is to break that line, making the platform carry its own weight through self-service, automation, and AI agents, not more hands. Nobody sits between you and the customer.
Responsibilities
Both tenancy models, as one system. Terraform, Helm, and ArgoCD delivering to shared and dedicated environments alike. You own the pipeline and the guardrails on it: progressive rollout, blast-radius containment, and drift detection that catches a mistake before a customer does.
Efficiency and cloud cost. Dedicated infrastructure per customer makes unit economics an engineering concern, not a quarterly finance exercise. You own attribution, right-sizing, spot strategy, and autoscaling that actually scales down - and you defend efficiency as a design constraint on every new deployment.
Turning operations into product. Onboarding, version rollouts, and config changes are tickets today; they should be self-service and boring. You decide what gets automated, what goes back to the owning team, and what still needs an expert.
Scale under real load. Temporal, KEDA queue-based autoscaling, and Karpenter are live and still maturing. You own how the platform grows and shrinks against real traffic, and what that costs per customer.
Security engineering for regulated customers. KMS-based key management, SIEM pipelines, WAF policy, certificate and secret lifecycle, per-customer SSO, vulnerability remediation, and turning penetration test findings into shipped fixes. Our customers are banks, insurers, and health funds. Their auditors are effectively part of our roadmap.
AI-driven operations. Our infrastructure repositories are already built for agentic work, with an internal harness of shared skills and agents. You inherit it, extend it, and push it into places it does not reach yet: triage, rollout verification, cost anomaly investigation, log analysis, and the long tail of operational work that has never been worth a script.
Reliability of what you ship. You are on call for your own systems, paged through PagerDuty against alerts you own.
Requirements: 6+ years in DevOps, including production ownership of Kubernetes at meaningful scale. EKS preferred.
Operating a fleet, with infrastructure as code. You have run many similar-but-not-identical environments with Terraform, and felt what happens when a change has to reach all of them. You have designed modules other people build on, and you have recovered state when it went wrong.
GitOps in production. You have operated ArgoCD, or an equivalent GitOps controller, running automated sync with pruning and self-healing enabled, and you understand exactly how much power that hands to a single commit.
Deep AWS and Cloudflare as a production edge. You have run AWS at the account, network, and IAM level rather than consumed it, and you have operated Cloudflare in production: WAF policy, origin protection, DNS, and rate limiting.
This position is open to all candidates.