we are building a vertically integrated AI neocloud from the electron up. Because we own the entire stack-from clean energy generation to the GPUs running customer training jobs-we treat power as a control input, not a constraint. To support our massive scaling vector from 30,000 accelerators to 300,000 without a linear growth in headcount, we are building the Cloud Availability Platform Engineering (CAPE) organization. CAPE serves as the horizontal reliability spine beneath all of our company Cloud. As the founding Engineering Manager, Production Engineering in Tel Aviv, you will stand up our local presence from scratch, driving a cultural shift from reactive firefighting to software-defined engineering.
This is a hybrid leadership and technical role where you will hire, scale, and lead a local team of exceptional systems generalists while remaining deeply hands-on in the code and incident response. You will ensure that your team spends at least 30% of their bandwidth on strategic automation, tooling, and firmware optimization primitives to prevent the operational treadmill. If you want to bridge core software layers with physical infrastructure and write the playbook for an entire engineering site, this founding seat is for you. This is a full-time position located in Tel Aviv, Israel.
What Youll Be Working On
Team Leadership & Founding Culture: Recruit, mentor, and establish a high-performing Production Engineering footprint in Tel Aviv, setting an uncompromising cultural standard for operational discipline and systems-first engineering.
Incident & On-Call Ownership: Partner with US and Dublin teams to run a follow-the-sun global on-call rotation, while championing a strict blameless post-mortem culture that targets systemic failures over human error.
Software-Defined Operations: Drive alert-reduction initiatives to improve fleet signal-to-noise ratios, automate routine manual workflows using modern runbook automation (e.g., Temporal), and build predictive monitoring to catch SEV1/SEV2 events before customers do.
Collaborative Governance: Act as the ultimate Production Gatekeeper across cross-functional compute, storage, networking, and platform teams, holding a strict line on Production Readiness Reviews and change control.
Strategic Reliability Engineering: Protect team bandwidth to ensure engineers spend at least 30% of their time on strategic automation, tooling, and firmware optimization primitives rather than drowning in incident response.
Physical-to-Digital Automation: Instill a software-first approach to physical problems, ensuring that any physical intervention occurring twice is successfully converted into a software-defined auto-remediation.
Requirements: Years of Infrastructure Experience: Minimum of 8+ years of experience working within infrastructure, SRE, or production engineering environments.
Engineering Leadership Track Record: Minimum of 2+ years of experience directly leading first-line engineering teams within a high-growth neocloud, hyperscaler, or large-scale distributed environment.
Non-Negotiable Coding Proficiency: Strong, hands-on software engineering fundamentals in Go, Python, C++, or a comparable systems language to build automation rather than scale through headcount.
Distributed Systems Depth: Expert-level command of Linux internals, container orchestration at scale, and root-cause analysis across complex physical-to-virtual boundaries.
Operational Execution Expertise: Proven track record of running tiered on-call models, establishing clear SLIs/SLOs and error budgets, and measurably reducing paging fatigue.
Bonus Points
AI Infrastructure Experience: Prior experience working at a neocloud or AI-infrastructure company operating massive GPU clusters.
High-Performance Fabric Exposure: Hands-on exposure to high-performance networks (such as InfiniBand or RoCEv2) or hardware internals (including BMC, firmware qualification, and attestation).
Accelerator Domain Knowledge: Deep.
This position is open to all candidates.