דרושים » ניהול ביניים » Infra Engineer

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a highly skilled Infrastructure Engineer to join our team and own the scaling, management, and automation of our platforms distributed environments. If youre excited about building high-scale distributed systems and solving deep DevOps and infrastructure challenges, lets talk.
What Youll Do:
Own and scale infrastructure - Design, build, and optimize the backbone of our observability platform, ensuring seamless deployment across hundreds of distributed environments.
Solve complex scalability challenges - Tackle unique problems in multi-cluster Kubernetes environments, multi-cloud setups, and high-ingestion observability pipelines.
Manage data at scale - Build and optimize configurable data pipelines, ensuring efficient ingestion, storage, and querying of large volumes of observability data with resilience, consistency, and analytical capabilities.
Automate everything - Develop infrastructure as code, improve CI/CD processes, and automate environment provisioning for reliability and efficiency.
Enhance system reliability - Design robust monitoring, alerting, and s\elf-healing mechanisms for a high-scale production environment.
Collaborate cross-functionally - Work closely with backend engineers, product teams, and customers to design scalable, developer-friendly infrastructure.
Adopt and implement cutting-edge technologies - Continuously evaluate and introduce new tools and frameworks to improve scalability, performance, and cost efficiency.
Improve deployment efficiency - Optimize Helm charts, Kubernetes operators, and Terraform configurations to streamline environment creation and lifecycle management.
Requirements:
5+ years of experience in DevOps, SRE, or Infrastructure Engineering roles.
Strong expertise in Kubernetes, Terraform, Helm, and cloud environments (AWS, GCP, or Azure).
Experience with scalable observability stacks (e.g., ClickHouse, VictoriaMetrics, OpenTelemetry) is a huge plus.
Deep understanding of distributed systems, networking, and containerized workloads.
Proficiency in at least one programming language (Go, Python, or similar) for automation and tooling
Passion for building scalable, reliable, and efficient infrastructure.
A problem-solving mindset with the ability to tackle complex technical challenges independently.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836937
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 30 דקות
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Senior DevOps Engineer who thrives in high-scale, high-performance environments. In this role, you will design and manage mission-critical infrastructure, develop automation tools, and ensure the reliability and scalability of our platform. You will work closely with engineering teams to optimize performance, enhance observability, and prevent incidents before they happen.
If you love solving complex infrastructure challenges and building tools that empower engineers, this role is for you!
What Will You Do?
As a Senior DevOps Engineer, you will:
Design, build, and maintain a highly available, scalable, and resilient production infrastructure handling millions of requests per second.
Manage and optimize tens of AWS accounts in a multi-account cloud environment.
Developing AI-driven automation frameworks and tools that more than 400 R&D engineers use while enhancing system reliability and engineering efficiency.
Design, implement, and support the LLM and agentic workflow infrastructure, ensuring its scalability and reliability.
Manage thousands of servers and containers, ensuring seamless infrastructure operations.
Improve and maintain our logging, monitoring, and alerting stacks for enhanced observability.
Troubleshoot and mitigate production incidents, participating in on-call rotations to ensure system health and stability.
Collaborate with engineering teams to optimize performance, reduce latency, and improve deployment processes.
Requirements:
What Will You Bring To The Team?
5+ years of experience in building and maintaining production infrastructure at scale.
5+ years of experience in developing automation tools and server-side applications using Python, Go, Ruby, Java, or Node.js.
Strong cloud experience with AWS, Google Cloud, or similar platforms.
Experience with the deployment, management, and leveraging of Large Language Models (LLMs) and agentic infrastructure.
Deep knowledge of Linux systems, including troubleshooting, architecture, and system internals.
Solid understanding of web servers, load balancers, caching systems, relational databases, and networking.
A proactive, problem-solving mindset with a passion for optimizing infrastructure and AI-driven automation.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8838344
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do at our company and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the company Technology team might be right for you.
We're looking for a Senior DevOps Engineer with a strong security orientation to join our company's DevOps team.
our company's platform runs at significant scale, and our DevOps team sits at the core of keeping it fast, reliable, and secure. As we grow, so does the scope of what we own - we need an engineer who can step in as a strong pillar of our company's production ecosystem.
This is a hands-on role for someone who takes security seriously, moves fast, and knows how to get things done in a complex, high-scale environment.
our company's Technology Stack:
AWS, GCP, CLoudFlare ,Kubernetes, Terragrunt, Ansible, Jenkins, ArgoCD, Argo Workflows, Kong & Nginx, HashiCorp Vault, Kafka, RabbitMQ, Mongodb, Aurora Postgresql & Mysql, Prometheus, Grafana, VictoriaMetrics
Programming languages: Python, NodeJS, Go, Kotlin
What am I going to do?
Full Ownership: Drive infrastructure initiatives through their entire lifecycle, taking accountability from initial design to delivery and long-term operations.
Kubernetes Orchestration: Architect, implement, and maintain production-grade Kubernetes clusters, ensuring they remain scalable and resilient under high-scale demand.
Cloud Architecture: Design and manage robust AWS environments, including VPC, IAM, and EKS, to support a highly available platform architecture.
Infrastructure as Code: Utilize Terraform to build and evolve our environment, applying configuration management principles to all IaC workflows.
CI/CD Excellence: Support and improve our deployment pipelines using Jenkins and GitHub Actions to maintain a fast development velocity.
Observability: Implement comprehensive monitoring solutions with Prometheus and Grafana to ensure deep visibility into platform health.
Operational Resilience: Join the DevOps on-call rotation, taking responsibility for mitigating production issues and maintaining site reliability.
Tooling & Innovation: Continuously evaluate and adopt tools - security and otherwise - that raise the bar on engineering efficiency and security posture at our company.
Requirements:
6+ years of hands-on DevOps / Platform Engineering experience in large-scale production environments on a public cloud (AWS preferred).
Proven leadership mindset - able to own projects end to end and be accountable for outcomes.
Strong, production-grade Kubernetes experience across design, deployment, scaling, and troubleshooting.
Solid AWS experience with VPC, IAM, EC2, EKS, Load Balancers, and DNS.
Experience designing and operating highly available, scalable infrastructure systems.
Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis).
Hands-on experience with Infrastructure as Code and configuration management (Terraform required; Terragrunt and Ansible a plus).
Experience with Docker and containerized workloads.
2+ years building and maintaining CI/CD pipelines (Jenkins, GitHub Actions)
Experience with monitoring and observability tools (Prometheus, Grafana).
Experience with GenAI platforms (AWS Bedrock, Vertex AI, OpenAI).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8806869
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
7 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Senior DevOps Engineer with a strong Developer Experience (DevEx) mindset, someone who genuinely enjoys making developers more productive by building and improving the tools, infrastructure, automation, and workflows they use every day.
You will work across CI/CD, developer tooling, automation, infrastructure, observability and productivity. Your primary focus will be making development and delivery faster, simpler, more reliable, and more cost-effective.
We need someone who can independently identify developer pain points, proactively discover opportunities for improvement, and turn them into scalable engineering solutions. You will be expected to measure the impact of your work and continuously improve critical metrics such as CI feedback time, pipeline reliability, deployment time, infrastructure cost, and overall developer productivity.
This role is a perfect fit for someone who thinks far beyond just "keeping the infrastructure running" and is excited about revolutionizing the entire engineering experience.
Your Impact:
You will tackle high-impact engineering challenges across our platform. Representative areas you will own and work on include:
Optimize CI/CD Pipeline Velocity & Stability:
Improve CI Stability: Leverage AI to automatically detect recurring test failures, dynamically disable problematic tests, and open Jira tickets for follow-up.
Automate Workflows: Utilize AI to automate merge conflict resolutions and keep the merge train moving smoothly.
Accelerate CI Velocity: Architect workflows to parallelize work both across jobs and within jobs. Introduce highly effective caching strategies and local Artifactory mirrors to drastically reduce build and dependency-fetch times.
Optimize CD: Reduce container image sizes and accelerate deployment times. Partner with QA to introduce and streamline automated regression testing against tenants.
Drive Repository Security & Maintenance:
Improve Security Posture: Automate the identification and remediation of Go and Python dependencies with known CVEs.
Automate Maintenance: Design systems to automatically handle Go and Python version upgrades across the codebase.
Ensure Performance: Automate performance testing to proactively detect and alert on performance regressions before they hit production.
Cloud Infrastructure & Cost Engineering:
Reduce Infrastructure Costs: Identify and implement innovative opportunities to reduce overall infrastructure and resource consumption.
Resource Efficiency: Introduce mechanisms to make our "slim tenants" even more resource-efficient for development and testing.
Smart Observability: Re-architect our logging strategy by streaming GCP logs to more cost-effective sinks instead of relying exclusively on Cloud Logging.
IaC Governance: Act as a gatekeeper for infrastructure quality by reviewing Terraform Merge Requests (MRs) for correctness, efficiency, maintainability, and strict adherence to engineering standards.
Requirements:
Your Experience:
5+ years of hands-on experience in DevOps, Platform Engineering, or SRE roles, preferably within data-heavy or high-scale environments.
The DevEx Mindset: A genuine passion for your job and a drive to build tools that make other engineers happier and more productive.
Coding & Automation: Strong programming skills in Python and Go. You treat infrastructure and automation as software engineering disciplines.
Cloud & IaC Mastery: Deep expertise in GCP (Google Cloud Platform) or AWS and advanced proficiency in Terraform.
CI/CD & Containers: Extensive experience building highly optimized CI/CD pipelines and optimizing Docker/Kubernetes deployments.
Forward-Thinking: An interest in or experience with leveraging AI tools to solve traditional operational bottlenecks.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8831063
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an independent Senior DevOps Engineer who thrives in a fast-paced environment. You will join our centralized DevOps team, serving as a pillar of reliability and innovation for the entire engineering organization.

This is a 50/50 "Build vs. Run" role. Splitting your time between building modern automation and improving our developer experience and operational excellence, ensuring our production environments are rock-solid, secure, and highly available while managing the live environment from server infrastructure all the way to the application.

Tasks you will take part in:
Design and implement next-generation CI/CD automation flows to streamline the path from "Code" to "Production."
Lead the evolution of our Infrastructure-as-Code (Terraform) to make our systems more modular and self-service for developers.
Collaborate with backend teams to architect scalable microservices and optimize datastore performance (RDS, MongoDB).
Embrace AI-powered development tools (Cursor, Claude Code) to automate toil and accelerate infrastructure delivery.
Full ownership of our AWS accounts and multi-cluster Kubernetes environments.
Manage and fine-tune live production environments, ensuring 99.9% uptime for our mission-critical fintech services.
Drive our Observability strategy (Datadog/Grafana) to proactively catch issues before they impact our customers.
Participate in a weekly on-call rotation, serving as the first line of defense for platform reliability.
is used for creating requisitions 'from scratch'
Requirements:
Requirements:
5+ years of experience managing high-scale production systems.
Mastery of AWS: Deep knowledge of AWS services, security best practices, and account management.
Orchestration & Containers: Expert-level experience with Kubernetes (EKS) and Docker.
Automation-First Mindset: Proficiency in Terraform and at least one high-level language (Python, Go, or Bash).
Database Ops: Solid understanding of the DevOps aspects of MySQL/RDS and MongoDB (scaling, backups, performance tuning).
SRE Discipline: Strong experience with monitoring and logging stacks like Datadog, Grafana, or Prometheus.
Team Player: Strong communication skills and a "can-do" attitude-you enjoy solving problems as part of a collaborative squad.
Experience with AI tools and a strong interest in continuously exploring and applying them in everyday work are highly valued.

Advantages:
Security Knowledge: Experience with FW, WAF, IPS/IDS, and SELinux.
Data Infrastructure: Familiarity with Data tools like Kafka, Spark, Snowflake or Data Lakes.
Multi-Cloud: Familiarity with hybrid or multi-cloud architectures.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8831065
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Tech is at the center of everything we do and we're looking for people who are builders at their core. From developers to visionaries and everything in between, we want minds who aren't just interested in putting the pieces together but who can find new ways to innovate. Solid communication, creative problem solving and business understanding are all prerequisites. So, if you're tech savvy, inquisitive, and ready to take the road less traveled, the Technology team might be right for you. We're looking for a Senior DevOps Engineer with a strong security orientation to join our DevOps team. our platform runs at significant scale, and our DevOps team sits at the core of keeping it fast, reliable, and secure. As we grow, so does the scope of what we own - we need an engineer who can step in as a strong pillar of our production ecosystem. This is a hands-on role for someone who takes security seriously, moves fast, and knows how to get things done in a complex, high-scale environment. Our Technology Stack: AWS, GCP, CLoudFlare,Kubernetes, Terragrunt, Ansible, Jenkins, ArgoCD, Argo Workflows, Kong & Nginx, HashiCorp Vault, Kafka, RabbitMQ, Mongodb, Aurora Postgresql & Mysql, Prometheus, Grafana, VictoriaMetrics Programming languages: Python, NodeJS, Go, Kotlin


What am I going to do?:

* Full Ownership: Drive infrastructure initiatives through their entire lifecycle, taking accountability from initial design to delivery and long-term operations.
* Kubernetes Orchestration: Architect, implement, and maintain production-grade Kubernetes clusters, ensuring they remain scalable and resilient under high-scale demand.
* Cloud Architecture: Design and manage robust AWS environments, including VPC, IAM, and EKS, to support a highly available platform architecture.
* Infrastructure as Code: Utilize Terraform to build and evolve our environment, applying configuration management principles to all IaC workflows.
* CI/CD Excellence: Support and improve our deployment pipelines using Jenkins and GitHub Actions to maintain a fast development velocity.
* Observability: Implement comprehensive monitoring solutions with Prometheus and Grafana to ensure deep visibility into platform health.
* Operational Resilience: Join the DevOps on-call rotation, taking responsibility for mitigating production issues and maintaining site reliability.
* Tooling & Innovation: Continuously evaluate and adopt tools - security and otherwise - that raise the bar on engineering efficiency and security posture.
Equal opportunities:
We're not about checklists. If you don't meet 100% of the requirements for this role but still feel passionate about the position and think you have the right skills and qualifications to excel at it, we want to hear from you. We prioritize diversity. We celebrate difference and embed it into every aspect of our workplace and product, as well as our community. We are proud and committed to providing equal opportunity employment to all individuals regardless of race, color, religion, sex, sexual orientation, citizenship, national origin, disability, Veteran status, or any other characteristic protected by law. In addition, we will provide accommodation to individuals with disabilities or a special need.
Requirements:
* 6+ years of hands-on DevOps / Platform Engineering experience in large-scale production environments on a public cloud (AWS preferred).
* Proven leadership mindset - able to own projects end to end and be accountable for outcomes.
* Strong, production-grade Kubernetes experience across design, deployment, scaling, and troubleshooting.
* Solid AWS experience with VPC, IAM, EC2, EKS, Load Balancers, and DNS.
* Experience designing and operating highly available, scalable infrastructure systems.
* Experience with managed and distributed databases (AWS Aurora, RDS, MongoDB, Redis).
* Hands-on
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8799862
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
23/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are a well-funded, early-stage startup looking for a talented and motivated Backend Engineer specializing in infrastructure to join our founding team. The focus of this role is to build and scale the infrastructure that powers autonomous AI agents automating complex enterprise workflows. You will own the systems, pipelines, and platforms that let our AI agents run reliably, securely, and at scale in production.

Your Impact
Infrastructure & Platform

Design, build, and own the core infrastructure powering our AI agent platform, from data pipelines to production deployment systems.

Build and scale the backend systems that support high-throughput document processing and data extraction workloads.

Cloud Infrastructure and Scalability

Architect and deploy infrastructure on cloud platforms (AWS, GCP, or Azure) with a focus on scalability, reliability, and cost efficiency.

Own containerization and orchestration (Docker, Kubernetes) for all production workloads.

Build and maintain CI/CD pipelines and DevOps practices that let the team ship fast without breaking things.

Data Infrastructure

Design and manage data pipelines to process and analyze large volumes of documents and unstructured data at scale.

Build the infrastructure layer connecting AI agents to databases, vector stores, and enterprise systems (ERP, CRM).

API & Systems Integration

Build and maintain robust, well-documented APIs connecting AI agents with external systems and enterprise software.

Design for reliability: retries, observability, and graceful degradation across distributed systems.

Security and Compliance

Implement authentication and authorization mechanisms (OAuth2, JWT) to secure AI-driven systems.

Ensure compliance with data privacy standards (e.g. GDPR, HIPAA) and drive best practices for secure data handling across the infrastructure.

Monitoring and Optimization

Build observability and monitoring systems to track infrastructure health, performance, and cost.

Continuously optimize system performance for speed, reliability, and cost-efficiency at scale.

Collaboration

Work closely with AI/ML engineers, product, and the founding team to make sure infrastructure decisions support fast iteration and production-grade reliability.

Participate in code reviews, design discussions, and architecture planning to drive infrastructure strategy.
Requirements:
5+ years of experience in backend or infrastructure engineering, ideally supporting production AI/ML systems or high-throughput data pipelines.

Proven track record of building and scaling infrastructure in production environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8793026
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Required DevOps Engineer
About the role:
Our DevOps team operates the infrastructure that powers our AI and Computer Vision platform across construction sites in 15+ countries. From data pipelines and ML workloads to backend services - you'll work with a diverse, modern, Kubernetes-based stack and have real influence on how we build, deploy, and operate.
What you'll do:
Own Multi-Cloud Infrastructure: Work alongside the team to design, scale, and operate our high-scale, multi-region production infrastructure across AWS and GCP, powering construction sites globally.
Drive Kubernetes at Scale: Manage and evolve our Kubernetes platform on EKS and leveraging GitOps practices with ArgoCD and Helm to enable safe, fast, and reliable deployments.
Build Robust CI/CD: Design and maintain CI/CD pipelines that empower dozens of engineers to ship confidently - with automation, testing, and progressive delivery built in.
Tackle Diverse Infrastructure Challenges: Work hands-on with a wide variety of workloads - from heavy data processing and Computer Vision pipelines to backend services and ML inference - each with unique scaling, performance, and reliability requirements.
Ensure Reliability & Observability: Build and maintain world-class observability (metrics, logs, tracing, alerting) so that issues are caught early and resolved fast. Performance, reliability, and scalability are at the core of what you do.
Security & Cost: Partner with the team to strengthen our security posture, identity and access management, compliance, and cloud cost optimization across both clouds.
Ownership from 0 to 1: You will have real influence over our architecture and tooling. We want engineers who care about shaping what we build and how we build it, ensuring performance, security, and observability are baked in from day one.
Requirements:
A seasoned DevOps / Infrastructure engineer (5+ years) with strong hands-on experience in production cloud environments.
Proven expertise operating large-scale, distributed systems - with deep understanding of Kubernetes, networking, and cloud-native architecture.
Strong experience with multi-cloud environments (AWS and/or GCP), Infrastructure-as-Code (Terraform), and GitOps workflows (ArgoCD, Flux, or similar).
Hands-on experience with CI/CD systems (Jenkins, GitHub Actions, etc.).
Solid scripting and automation skills (Python, Bash, or Go).
Proven track record of being a collaborative team player who partners closely with developers, ML engineers, and cross-functional stakeholders across the organization.
Experience with observability stacks (Prometheus, Grafana, OpenTelemetry, Logz.io, or similar).
Experience with databases (relational and/or NoSQL) - including operational aspects like backups, migrations, and performance tuning.
AI-Native Engineering: You are an AI-native engineer who leverages LLMs and agentic tools (like Cursor, Copilot, or Claude) not just for command completion, but as a core operational partner - automating diagnostics, runbooks, and infrastructure workflows so you can focus on the critical things.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837157
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 6 שעות
חברה חסויה
Location: Tel Aviv-Yafo and Ra'anana
Job Type: Full Time
As a Senior DevOps Engineer, youll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIAs AI infrastructure.


What You'll Be Doing:
Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations.
Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints.
Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory.
Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity.
Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations.
Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows.
Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening.
Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes.
Requirements:
What We Need to See:
Bachelors degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices.
Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting.
Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible.
Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices.
Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting.
Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting.
Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment.


Ways To Stand Out From the Crowd:
Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost.
Familiarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage.
Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana.
Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers.
Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837920
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior Site Reliability Engineer at our company 911, you'll own the infrastructure that keeps our platform reliable, scalable, and secure - work that directly supports mission-critical 911 systems used by public safety agencies. You'll drive infrastructure-as-code practices across AWS, lead observability efforts through Datadog, and bring modern AI-assisted engineering approaches into how the team builds and operates.
What You'll Do
Own and evolve AWS infrastructure using Infrastructure-as-Code (Terraform / Terragrunt)
Architect and scale AWS environments
Deploy, scale, and manage containerized workloads using Kubernetes and Docker; contribute to HA/DR architecture and platform strategy
Lead deployment and release processes using Argo (reference JD also names Bitbucket, Jenkins as part of the CI/CD toolset).
Define and enforce SLOs, SLIs, and error budgets; drive toil reduction across the platform
Drive full utilization of Datadog for monitoring, dashboards, and alerting across the platform (reference JD also names Prometheus, Grafana as potential observability tooling)
Build self-service internal developer platforms that empower teams to ship faster.
Take end-to-end ownership of infrastructure projects - define success criteria, execute, and measure outcomes.
Partner cross-functionally with engineering teams (e.g., network engineering, Dev owners) on long-term technical planning.
Bring AI-assisted engineering practices (e.g., Claude, MCP integrations) into daily workflows to improve team efficiency
Document work and provide cross-training to peers.
Resolve JIRA tickets across Cloud, CI/CD, deployments, and monitoring.
Requirements:
At least 6 years of experience as a DevOps/SRE engineer in a cloud environment
Hands-on, production-level AWS experience.
Hands-on production experience with Kubernetes and containerization
Experience with Terraform/Terragrunt (or similar Infrastructure-as-Code tools) - required
Strong Bash scripting skills
Deep understanding of SRE principles: SLOs, SLIs, error budgets, toil reduction, blameless post-mortems
Strong incident management / on-call experience
Solid understanding of APIs, microservices, and distributed systems
Demonstrated experience leading a project end-to-end, from defining success criteria through delivery and measurement
Communicates effectively across teams and can drive long-term technical planning
Practical experience with AI-assisted engineering tools (e.g., Claude, Cursor) and MCP-style integrations is a strong plus
Experience building AI/ML infrastructure (model deployment, inference pipelines)-plus.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796929
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced DevOps Engineer to join our Engineering team and play a key role in building and operating our cloud-native platform. The ideal candidate will have hands-on experience managing production environments at scale, driving cloud transformation initiatives, and supporting the of enterprise systems from on-premises deployments to modern SaaS and cloud-native architectures. You will be responsible for designing, automating, and maintaining scalable infrastructure, CI/CD pipelines, and deployment processes that support both our core products and emerging AI-driven capabilities. Working closely with Engineering, QA, Product, and AI teams, you will help ensure the reliability, security, and performance of our services while driving operational excellence and continuous improvement across our technology stack.
Responsibilities:
Design, implement, and maintain CI/CD pipelines.
Manage and optimize cloud infrastructure across AWS, Azure, and/or GCP.
Develop and maintain Infrastructure as Code using Terraform.
Manage Kubernetes-based environments and GitOps deployment workflows using Argo CD and Kustomize.
Lead and support the migration of enterprise applications and infrastructure from on-premises
environments to scalable SaaS and cloud-native architectures.
Establish, maintain, and continuously improve production environments, ensuring high availability, security, scalability, and operational excellence.
Demonstrate strong production ownership, including incident management, root cause analysis, capacity planning, and performance optimization.
Collaborate with Engineering, QA, Product, and AI teams.
Support the deployment, operation, and monitoring of AI and Generative AI services.
Build and maintain monitoring, logging, and alerting systems.
Troubleshoot and resolve infrastructure, deployment, and production issues.
Requirements:
5+ years of experience as a DevOps Engineer or similar
infrastructure-focused role.
Hands-on experience with Azure, Aws, or GCP.
Experience with CI/CD tools such as Jenkins, GitHub Actions, or similar platforms.
Strong knowledge of Terraform and Infrastructure as Code practices.
Experience with Docker, Kubernetes, Argo CD, and Kustomize.
Experience designing and operating production-grade Kubernetes saas platforms or enterprise environments.
Experience with monitoring and observability tools such as Prometheus, Grafana, and ELK.
Strong troubleshooting, analytical, and communication skills.
B.Sc. in Computer Science, Computer Engineering, Information Systems, or a related field (or equivalent practical experience).
Nice to have:
Experience supporting AI, Machine Learning, or Generative AI workloads, including familiarity with MLOps concepts, AI deployment platforms, or cloud-based AI services.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796409
סגור
שירות זה פתוח ללקוחות VIP בלבד