דרושים » הנדסה » Production Operations Engineer

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Production Operations Engineer to join our fast-growing company at a breakthrough stage, as we build our dream team with the most passionate and professional people in the industry. Our team thinks differently and quickly, delivering high-quality, innovative solutions with the latest technologies - and we never forget to enjoy the ride along the way!
This is a role for someone who owns production stability end to end and knows that the best fixes often happen through other people. You won't be the one writing every manifest or resolving every incident yourself - you'll be the one who sees what production needs, sets the direction, and makes sure R&D teams follow through.
What You'll Be Doing:
Owning production stability: Act as the go-to person for the health of our production environments, proactively identifying risks before they become incidents.
Owning the incident process: Make sure incidents are properly tracked, follow-ups are driven to completion, and R&D teams write and own their postmortems - turning lessons into lasting improvements rather than letting them slip.
Driving change across R&D: Partner with engineering teams to define the reliability and operational improvements production needs, and make sure they get prioritized and implemented.
Raising the operational bar: Champion production-readiness standards, define SLOs and error budgets, and build a culture where teams own the reliability of what they ship.
Connecting the dots: Communicate clearly with stakeholders across R&D and leadership, translating between technical detail and business impact.
Requirements:
4+ years of experience in production operations, SRE, DevOps, or a related engineering role.
Strong stakeholder skills: You can influence without authority, align teams around priorities, and communicate credibly with engineers and leadership alike.
Process ownership: Experience driving operational processes to completion - you're relentless about follow-through and comfortable holding other teams accountable for their commitments.
Cloud understanding: Solid grasp of public cloud environments (GCP/AWS/Azure) - enough to reason about production architecture and drive the right decisions.
Containers & orchestration: Working knowledge of Docker, Kubernetes, and microservices architectures.
Ownership & drive: You take responsibility for outcomes, challenge the status quo constructively, and keep pushing until problems are actually solved.
Team spirit: Collaborative, curious, and driven - you thrive in an empowering, positive environment
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8822405
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 4 שעות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Production Operations Engineer to join our fast-growing company at a breakthrough stage, as we build our dream team with the most passionate and professional people in the industry. Our team thinks differently and quickly, delivering high-quality, innovative solutions with the latest technologies - and we never forget to enjoy the ride along the way!

This is a role for someone who owns production stability end to end and knows that the best fixes often happen through other people. You won't be the one writing every manifest or resolving every incident yourself - you'll be the one who sees what production needs, sets the direction, and makes sure R&D teams follow through.


What You'll Be Doing:
Owning production stability: Act as the go-to person for the health of our production environments, proactively identifying risks before they become incidents.
Owning the incident process: Make sure incidents are properly tracked, follow-ups are driven to completion, and R&D teams write and own their postmortems - turning lessons into lasting improvements rather than letting them slip.
Driving change across R&D: Partner with engineering teams to define the reliability and operational improvements production needs, and make sure they get prioritized and implemented.
Raising the operational bar: Champion production-readiness standards, define SLOs and error budgets, and build a culture where teams own the reliability of what they ship.
Connecting the dots: Communicate clearly with stakeholders across R&D and leadership, translating between technical detail and business impact.
Requirements:
4+ years of experience in production operations, SRE, DevOps, or a related engineering role.
Strong stakeholder skills: You can influence without authority, align teams around priorities, and communicate credibly with engineers and leadership alike.
Process ownership: Experience driving operational processes to completion - you're relentless about follow-through and comfortable holding other teams accountable for their commitments.
Cloud understanding: Solid grasp of public cloud environments (GCP/AWS/Azure) - enough to reason about production architecture and drive the right decisions.
Containers & orchestration: Working knowledge of Docker, Kubernetes, and microservices architectures.
Ownership & drive: You take responsibility for outcomes, challenge the status quo constructively, and keep pushing until problems are actually solved.
Team spirit: Collaborative, curious, and driven - you thrive in an empowering, positive environment
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8822400
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
26/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are building a vertically integrated AI neocloud from the electron up. Because we own the entire stack-from clean energy generation to the GPUs running customer training jobs-we treat power as a control input, not a constraint. To support our massive scaling vector from 30,000 accelerators to 300,000 without a linear growth in headcount, we are building the Cloud Availability Platform Engineering (CAPE) organization. CAPE serves as the horizontal reliability spine beneath all of our company Cloud. As the founding Engineering Manager, Production Engineering in Tel Aviv, you will stand up our local presence from scratch, driving a cultural shift from reactive firefighting to software-defined engineering.
This is a hybrid leadership and technical role where you will hire, scale, and lead a local team of exceptional systems generalists while remaining deeply hands-on in the code and incident response. You will ensure that your team spends at least 30% of their bandwidth on strategic automation, tooling, and firmware optimization primitives to prevent the operational treadmill. If you want to bridge core software layers with physical infrastructure and write the playbook for an entire engineering site, this founding seat is for you. This is a full-time position located in Tel Aviv, Israel.
What Youll Be Working On
Team Leadership & Founding Culture: Recruit, mentor, and establish a high-performing Production Engineering footprint in Tel Aviv, setting an uncompromising cultural standard for operational discipline and systems-first engineering.
Incident & On-Call Ownership: Partner with US and Dublin teams to run a follow-the-sun global on-call rotation, while championing a strict blameless post-mortem culture that targets systemic failures over human error.
Software-Defined Operations: Drive alert-reduction initiatives to improve fleet signal-to-noise ratios, automate routine manual workflows using modern runbook automation (e.g., Temporal), and build predictive monitoring to catch SEV1/SEV2 events before customers do.
Collaborative Governance: Act as the ultimate Production Gatekeeper across cross-functional compute, storage, networking, and platform teams, holding a strict line on Production Readiness Reviews and change control.
Strategic Reliability Engineering: Protect team bandwidth to ensure engineers spend at least 30% of their time on strategic automation, tooling, and firmware optimization primitives rather than drowning in incident response.
Physical-to-Digital Automation: Instill a software-first approach to physical problems, ensuring that any physical intervention occurring twice is successfully converted into a software-defined auto-remediation.
Requirements:
Years of Infrastructure Experience: Minimum of 8+ years of experience working within infrastructure, SRE, or production engineering environments.
Engineering Leadership Track Record: Minimum of 2+ years of experience directly leading first-line engineering teams within a high-growth neocloud, hyperscaler, or large-scale distributed environment.
Non-Negotiable Coding Proficiency: Strong, hands-on software engineering fundamentals in Go, Python, C++, or a comparable systems language to build automation rather than scale through headcount.
Distributed Systems Depth: Expert-level command of Linux internals, container orchestration at scale, and root-cause analysis across complex physical-to-virtual boundaries.
Operational Execution Expertise: Proven track record of running tiered on-call models, establishing clear SLIs/SLOs and error budgets, and measurably reducing paging fatigue.
Bonus Points
AI Infrastructure Experience: Prior experience working at a neocloud or AI-infrastructure company operating massive GPU clusters.
High-Performance Fabric Exposure: Hands-on exposure to high-performance networks (such as InfiniBand or RoCEv2) or hardware internals (including BMC, firmware qualification, and attestation).
Accelerator Domain Knowledge: Deep.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8797889
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
11/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do at Fiverr and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the Fiverr Technology team might be right for you. Fiverr's Engineering team is expanding its DevOps capabilities to manage our rapidly growing cloud infrastructure. We're looking for a mid-level DevOps Engineer to join this high-velocity environment, taking ownership of cloud assets, contributing to production stability, and driving key projects. You'll be instrumental in strengthening our team and supporting our continued innovation.

What am I going to do?:

* Independently deploy and maintain robust cloud infrastructure across AWS, GCP, and CloudFlare, leveraging Kubernetes, Terragrunt, and Ansible.
* Contribute to an on-call rotation, ensuring the stability of critical production services including Kafka, RabbitMQ, and various databases, with a focus on proactive incident resolution.
* Lead small to mid-sized projects from conception through completion, utilizing your expertise in CI/CD pipelines (Jenkins, ArgoCD, Argo Workflows) and scripting languages like Python, NodeJS, Go, or Kotlin.
* Partner closely with security and development teams to embed security best practices throughout the entire software development lifecycle.
* Proactively evaluate and implement new tools and technologies to enhance engineering efficiency, security posture, and operational excellence.
* Manage sensitive information securely using tools like HashiCorp Vault and ensure secure configurations for services like Kong & Nginx.

Equal opportunities:
At Fiverr, we know that talent has no single face. We welcome talent from everywhere and everyone because it makes everything we build better. Need accommodations? Just ask. And if this role excites you but you don't tick every box, apply anyway. The best people rarely fit the mold exactly.
Requirements:
* 4+ years of hands-on DevOps / Platform Engineering experience in production environments within a public cloud environment (AWS preferred)
* Strong, production-grade Kubernetes experience (design, deployment, scaling, and troubleshooting) with solid AWS experience (VPC, IAM, EC2, EKS, Load Balancers, DNS)
* Experience designing and operating highly available, scalable infrastructure systems
* Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis)
* Hands-on experience with Infrastructure as Code and configuration management (Terraform required, Terragrunt & Ansible – advantage)
* Experience with Docker and containerized workloads
* 2+ years of experience building and maintaining CI/CD pipelines (Jenkins, GitHub Actions)
* Proficiency in Python for automation and strong Linux administration skills
* Experience with monitoring and observability tools (Prometheus, Grafana)
* Development experience and familiarity with GenAI platforms (AWS Bedrock, Vertex AI, OpenAI) – advantage Working with AI At Fiverr, AI is a powerful partner in our engineering workflows. You'll leverage AI-powered tools like GitHub Copilot for code assistance, intelligent security scanning tools, and automation platforms to streamline deployments and monitoring. While AI handles repetitive tasks and provides insights, your critical thinking, architectural decisions, and complex problem-solving remain at the forefront.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776958
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
26/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep Navan reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.



This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.



This position is based out of our new Tel Aviv office.



What You'll Do:

Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
Ability to troubleshoot AI tools and provider/API issues, including rate limits, quota, auth, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations.
Excellent communication skills and the ability to work with stakeholders and domain experts across the company.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8797911
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do at our company and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the company Technology team might be right for you.
our company's Engineering team is expanding its DevOps capabilities to manage our rapidly growing cloud infrastructure. We're looking for a mid-level DevOps Engineer to join this high-velocity environment, taking ownership of cloud assets, contributing to production stability, and driving key projects. You'll be instrumental in strengthening our team and supporting our continued innovation.
What am I going to do?
Independently deploy and maintain robust cloud infrastructure across AWS, GCP, and CloudFlare, leveraging Kubernetes, Terragrunt, and Ansible.
Contribute to an on-call rotation, ensuring the stability of critical production services including Kafka, RabbitMQ, and various databases, with a focus on proactive incident resolution.
Lead small to mid-sized projects from conception through completion, utilizing your expertise in CI/CD pipelines (Jenkins, ArgoCD, Argo Workflows) and scripting languages like Python, NodeJS, Go, or Kotlin.
Partner closely with security and development teams to embed security best practices throughout the entire software development lifecycle.
Proactively evaluate and implement new tools and technologies to enhance engineering efficiency, security posture, and operational excellence.
Manage sensitive information securely using tools like HashiCorp Vault and ensure secure configurations for services like Kong & Nginx.
Requirements:
4+ years of hands-on DevOps / Platform Engineering experience in production environments within a public cloud environment (AWS preferred)
Strong, production-grade Kubernetes experience (design, deployment, scaling, and troubleshooting) with solid AWS experience (VPC, IAM, EC2, EKS, Load Balancers, DNS)
Experience designing and operating highly available, scalable infrastructure systems
Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis)
Hands-on experience with Infrastructure as Code and configuration management (Terraform required, Terragrunt & Ansible - advantage)
Experience with Docker and containerized workloads
2+ years of experience building and maintaining CI/CD pipelines (Jenkins, GitHub Actions)
Proficiency in Python for automation and strong Linux administration skills
Experience with monitoring and observability tools (Prometheus, Grafana)
Development experience and familiarity with GenAI platforms (AWS Bedrock, Vertex AI, OpenAI) - advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8806987
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
16/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are hiring a DevOps Engineer (Engineer 3) for the Identity Protection Operations team. We are looking for a senior-level engineer who is deeply fluent in modern cloud-native technologies and can leverage them to design, build, and continuously improve the systems that power our Identity Protection products. This role is open in Tel Aviv, Israel.

What You'll Do:

Design and implement new services and integrations across a hybrid environment - primarily AWS and an on-premises data center, with additional presence on GCP and OCI - leveraging shared platforms such as Kubernetes, Kafka, and modern monitoring stacks.

Build and maintain CI/CD pipelines (Jenkins) and configuration management workflows using Chef and related tooling.

Write automation, tooling, and glue logic in Bash and Python to reduce toil and accelerate delivery.

Utilize AI and automation platforms (e.g., Claude, n8n) to drive engineering efficiency and intelligent operations.

Contribute architectural insight to platform decisions - understanding system trade-offs and designing for scale and reliability.

Participate in on-call rotations and incident response for production Identity Protection services.
Requirements:
What You'll Need:

Deep knowledge and hands-on experience with Linux-based systems, including administration, performance tuning, and troubleshooting in production environments.

Deep familiarity with Kubernetes - enough to design workloads, debug failures, and make informed architectural decisions.

Strong understanding of distributed messaging systems (Kafka and/or RabbitMQ)

Deep experience with MongoDB (or comparable NoSQL): schema design, query optimization, index tuning, maintenance operations, and understanding of internals - working closely with the data layer.

Proficiency with Docker and containerized application design.

Working knowledge of observability tooling - Prometheus, Grafana, and centralized logging stacks - and the ability to instrument and interpret systems using them.

Solid scripting in Bash and Python; comfortable writing production-grade automation.

Experience with configuration management tools, Chef preferred.

Ability to reason at the architectural level: understand big-picture system design and communicate trade-offs clearly.

On-call availability on a rotational basis.

Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.


Bonus Points:

Background in identity, authentication, or security infrastructure.

Terraform or other IaC experience.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8783671
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
31/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Zipher is building the Autonomous Execution Layer for cloud data and AI workloads. Backed by $50M in funding , we dynamically orchestrate clusters, predict bottlenecks, and auto-heal infrastructure in real time — with zero human intervention . Our platform runs in production at global enterprise customers, including Fortune 500 companies , delivering mission-critical resilience and sub-second optimization We are looking for a DevOps Engineer to build and own the platform our autonomous execution engine runs on. You will own how Zipher ships, scales, and stays up — the delivery pipelines, the Kubernetes infrastructure, and the reliability of systems enterprise customers depend on around the clock.
What You’ll Do

* Architect and own GitOps-based delivery — ArgoCD, Helm, Terraform, GitHub Actions — so the engine ships to production safely, many times a day
* Build and operate the Kubernetes platform , including on-demand environments spun up per pull request and torn down automatically
* Design deployment safety into the platform: canary analysis, blue/green rollouts, automated rollback, and an observability stack (Prometheus, Grafana, OpenTelemetry) that surfaces failure first
* Partner closely with backend and data engineers to make production infrastructure secure, reproducible, and cost-aware across AWS accounts and enterprise deployments
* Drive reliability end to end: define SLOs , own incident response, and raise the operational bar for mission-critical services
What We Offer
* Own the platform behind a new category of autonomous cloud infrastructure High ownership from day one : real architectural influence, direct exposure to founders, and responsibility for mission-critical systems
* A small, technical, high-velocity team that values curiosity, speed, rigor, and engineering craftsmanship Top-of-market compensation and meaningful equity

Ready to own the platform that lets an autonomous execution engine run in production? Hit Apply.
Requirements:
What You’ll Bring 6+ years in DevOps, Platform, or Infrastructure Engineering , including ownership of production environments for a real product at scale
* Deep hands-on experience with Kubernetes, Helm, and Terraform , with GitOps (ArgoCD or equivalent) as your default way to ship
* Strong production experience on AWS — EKS, IAM, VPC networking, and managed services such as Lambda, S3, DynamoDB, or Kinesis — plus scripting in Python, Go, or Bash
* Real depth in observability and operations : metrics, tracing, log aggregation, alerting, and SLOs you defined and defended
* A high-agency, engineering-first mindset : you enjoy ambiguous, high-leverage problems and take responsibility for reliability, performance, and security
Nice to Have
* Experience building ephemeral environments with Crossplane, Terraform, or a home-grown control plane
* Experience operating data and streaming infrastructure at scale , such as Kafka, Spark, EMR, MongoDB Atlas, or Snowflake
* Familiarity with security and compliance in a fast-moving startup: SOC 2, SSO/MFA, secret scanning, IaC policy enforcement, and cloud cost optimization
* Experience in an elite IDF technology unit or another high-performance engineering environment
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8804051
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior Site Reliability Engineer at our company 911, you'll own the infrastructure that keeps our platform reliable, scalable, and secure - work that directly supports mission-critical 911 systems used by public safety agencies. You'll drive infrastructure-as-code practices across AWS, lead observability efforts through Datadog, and bring modern AI-assisted engineering approaches into how the team builds and operates.
What You'll Do
Own and evolve AWS infrastructure using Infrastructure-as-Code (Terraform / Terragrunt)
Architect and scale AWS environments
Deploy, scale, and manage containerized workloads using Kubernetes and Docker; contribute to HA/DR architecture and platform strategy
Lead deployment and release processes using Argo (reference JD also names Bitbucket, Jenkins as part of the CI/CD toolset).
Define and enforce SLOs, SLIs, and error budgets; drive toil reduction across the platform
Drive full utilization of Datadog for monitoring, dashboards, and alerting across the platform (reference JD also names Prometheus, Grafana as potential observability tooling)
Build self-service internal developer platforms that empower teams to ship faster.
Take end-to-end ownership of infrastructure projects - define success criteria, execute, and measure outcomes.
Partner cross-functionally with engineering teams (e.g., network engineering, Dev owners) on long-term technical planning.
Bring AI-assisted engineering practices (e.g., Claude, MCP integrations) into daily workflows to improve team efficiency
Document work and provide cross-training to peers.
Resolve JIRA tickets across Cloud, CI/CD, deployments, and monitoring.
Requirements:
At least 6 years of experience as a DevOps/SRE engineer in a cloud environment
Hands-on, production-level AWS experience.
Hands-on production experience with Kubernetes and containerization
Experience with Terraform/Terragrunt (or similar Infrastructure-as-Code tools) - required
Strong Bash scripting skills
Deep understanding of SRE principles: SLOs, SLIs, error budgets, toil reduction, blameless post-mortems
Strong incident management / on-call experience
Solid understanding of APIs, microservices, and distributed systems
Demonstrated experience leading a project end-to-end, from defining success criteria through delivery and measurement
Communicates effectively across teams and can drive long-term technical planning
Practical experience with AI-assisted engineering tools (e.g., Claude, Cursor) and MCP-style integrations is a strong plus
Experience building AI/ML infrastructure (model deployment, inference pipelines)-plus.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796929
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Tech Lead DevOps Engineer, you'll own the technical strategy for our cloud infrastructure end-to-end, from architecture to production. You'll build the automation frameworks, IaC tooling, and CI/CD systems that let engineering ship fast and safely at scale, while setting the bar for observability, reliability, and incident response across the org. You'll work closely with engineering leadership to shape the infrastructure roadmap and make the tradeoffs between velocity, cost, security, and reliability, acting as a technical multiplier for the broader team.

What Youll Do
Responsibilities:
Own and drive the end-to-end technical strategy for cloud infrastructure across AWS and multi-cloud environments, from conceptual architecture to production readiness.
Design and implement advanced automation frameworks, internal developer platforms, and tooling utilizing Python, Bash, and modern infrastructure-as-code technologies (e.g., Terraform, CDKTF).
Architect and continuously evolve CI/CD systems and delivery pipelines at scale, reducing friction and elevating developer productivity and deployment safety standards.
Define and enforce organization-wide observability, reliability, and performance standards, establishing best practices for monitoring, alerting, incident response, and system health.
Partner closely with engineering leadership and cross-functional teams to shape infrastructure roadmaps, enable scalable deployments, and ensure resilient production systems.
Make and own high-impact technical tradeoff decisions that judiciously balance velocity, cost efficiency, security, and reliability in a dynamic, growth-oriented environment.
Serve as a technical mentor and force multiplier, elevating the infrastructure capabilities of the broader engineering organization and championing a DevOps-first culture.
Requirements:
8+ years of hands-on DevOps, Infrastructure, or Platform Engineering experience, preferably in high-growth product environments.
Demonstrated ability to lead complex, cross-cutting technical initiatives end-to-end - from initial design and implementation through operational maturity and rollout.
Deep expertise in cloud platforms (preferably AWS) with advanced knowledge of networking, security architectures, cost optimization strategies, and scalability patterns.
Strong software engineering fundamentals and proficiency in Python, Bash, with beneficial experience in languages like TypeScript or Go for tooling and automation development.
Expert-level hands-on experience with Infrastructure as Code (e.g., Terraform, CDKTF, CloudFormation) and implementing modern GitOps workflows.
Advanced knowledge of CI/CD systems (e.g., GitHub Actions, Jenkins) and experience with production-grade container orchestration platforms (e.g., Kubernetes, EKS, or ECS).
Proven experience designing and operating robust observability stacks at scale (e.g., Prometheus, Grafana, ELK, Datadog).
Outstanding collaboration and communication skills; a technically rigorous, pragmatic leader who thrives in high-velocity environments and builds trust through deep technical knowledge and sound judgment.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796001
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced and highly motivated Senior DevOps Engineer to join our engineering and DevOps team. As a Senior DevOps Engineer, you will be the architect of our infrastructure, ensuring that platform is scalable, resilient, and secure. This is a hands-on role where you will bridge the gap between development and operations, automating our deployment pipelines and managing our cloud-native ecosystem. This is an incredible opportunity to shape the foundational infrastructure of a high-growth data company.

What You'll Do

Design, implement, and manage our cloud infrastructure using tools like Terraform, ensuring environment consistency and scalability.

CI/CD Automation: Take full ownership of our deployment pipelines, optimizing for speed, reliability, and developer productivity.

Cloud Orchestration: Manage and scale our Kubernetes (K8s) clusters on AWS, ensuring high availability and efficient resource utilization.

Observability & Monitoring: Implement and maintain robust monitoring, logging, and alerting systems to ensure platform health and rapid incident response.

Security & Compliance: Drive security best practices across the infrastructure, including IAM management, network security, and vulnerability scanning.

Collaborate: Work closely with software engineers to optimize application performance, containerization strategies, and database reliability.
Requirements:
We are seeking a hands-on builder who can grow into a technical leader for our infrastructure. Someone excited to own hard problems and pick up the specifics of our stack quickly.

Must-Haves

5+ years of experience in DevOps or Site Reliability Engineering (SRE), ideally within a high-growth SaaS or data-heavy environment.

Hands-on experience running production workloads on AWS (e.g., EKS, RDS, S3, IAM, VPC).

Strong experience with Kubernetes and Docker in production environments.

Proficiency with CI/CD tooling (e.g., GitHub Actions, GitLab CI, or Jenkins) and infrastructure-as-code (e.g., Terraform).

Working proficiency in Python, Bash, or Go for automation and internal tooling.

A team player with strong communication skills who can explain infrastructure concepts to cross-functional partners.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8791545
סגור
שירות זה פתוח ללקוחות VIP בלבד