דרושים » תוכנה » SRE Technical Lead, DevOps Group

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time and Hybrid work
This role requires working out of the Tel Aviv office three days per week.

About the Role:
As a SRE Technical Lead at our company, you will play a critical role in ensuring the reliability, scalability, and performance of the platform that powers our customers operations. Youll operate at the intersection of software engineering and production operations, taking full ownership of the systems you build and run.
This is a hands-on individual contributor role - not a managerial position. You'll be deep in the technical work, driving impact through engineering excellence rather than people management.
This role is not just about responding to incidents - its about fundamentally improving how our platform behaves under real-world conditions. You will drive reliability initiatives end-to-end: defining measurable service goals, shaping engineering priorities through error budgets, and implementing solutions that prevent issues before they occur.
Youll work closely with teams across the organization, embedding reliability and observability into every layer of the stack. At the same time, youll leverage automation, modern infrastructure practices, and emerging AI capabilities to continuously evolve how we operate and scale.

What you will do:
Develop deep product knowledge across our platform - understanding its internals, failure modes, and operational behavior well enough to own incident resolution end-to-end.
Define and track SLAs/SLOs/SLIs across critical platform services, and use error budgets to drive engineering decisions.
Own production reliability - including on-call rotations, incident response, and post-mortems - with a focus on minimizing MTTR and preventing recurrence through systemic fixes, not just firefighting.
Work hand-in-hand with engineering teams across the stack - infrastructure, application, and business layers - to embed reliability requirements everywhere.
Requirements:
What skills and experience youll bring to our company:
5+ years of experience as an SRE or Platform Developer (or similar) in a high-scale production environment, with hands-on ownership across the full stack - infrastructure and application layers.
Experience introducing or scaling AI-powered systems in real-world products (ML, LLMs, agents, or decision systems)
Strong coding skills and a software engineering mindset - you build your own tools rather than waiting for someone else to.
A true owner - you take responsibility for systems end-to-end and proactively drive improvements without waiting for direction.
Business-level reliability experience is a strong advantage.
Experience with infrastructure-as-code and modern container orchestration platforms.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8743767
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
15/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. we are building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep our company reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.
This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.
This position is based out of our new Tel Aviv office.
What You'll Do:
Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739673
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
29/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an experienced SRE Team Lead to drive the reliability, observability, and automation practices across our private cloud infrastructure and operations. In this role, you will lead a team of site reliability engineers, own the engineering roadmap for monitoring and automation, and act as a key liaison between development, operations, and platform teams. You bring at least 3-4 years of hands-on people management experience and a deep technical background in SRE or DevOps disciplines.


What will you do?

Leadership & Team Management

Lead, mentor, and grow a team of SREs, providing technical direction, career development guidance, and day-to-day management.

Own the team roadmap for reliability, observability, and automation initiatives - prioritizing work, removing blockers, and driving delivery.

Conduct regular 1:1s, performance reviews, and hiring processes to build and sustain a high-performing team.

Foster a culture of operational excellence, blameless post-mortems, and continuous improvement.

Act as an escalation point for complex incidents and reliability issues, leading post-incident reviews and ensuring follow-through on action items.


Automation & Infrastructure

Design, develop, and maintain automation tools to support infrastructure and operations teams at scale.

Manage pipelines and infrastructure workflows using Jenkins, Ansible, Python, and Bash.

Drive the adoption of infrastructure-as-code practices across the organization.

Collaborate with system engineers to improve scalability, performance, and fault tolerance of critical systems.


Monitoring & Observability

Build and extend monitoring and alerting systems using Grafana, the ELK (Elastic) stack, Zabbix, and custom scripts.

Implement and enforce observability best practices to ensure full visibility into systems, applications, and infrastructure.

Define and track SLIs, SLOs, and error budgets across key services.

Partner with development teams to embed observability earlier in the software development lifecycle.


Database & Platform Support

Support monitoring and infrastructure integration for databases including MongoDB and PostgreSQL.

Maintain documentation and champion knowledge sharing around automation, monitoring, and reliability practices.
Requirements:
Experience & Leadership:

3-4+ years of experience in a people management or team lead capacity within SRE, DevOps, or infrastructure engineering.

5-8+ years of overall experience in SRE, DevOps, or infrastructure automation roles.

Proven track record of building, coaching, and retaining high-performing engineering teams.

Experience owning an engineering roadmap and driving cross-functional reliability initiatives.


Technical Skills :

Strong scripting skills in Python and Bash; comfortable building and maintaining production-grade automation.

Hands-on experience with infrastructure automation tools, particularly Ansible.

Solid experience with monitoring and observability platforms - ELK stack, Grafana, and Zabbix.

Good understanding of CI/CD pipelines and related tooling, including Jenkins.

Familiarity with managing and monitoring MongoDB and PostgreSQL in a production environment.

Comfortable working in Linux-based environments.

Excellent problem-solving skills and strong written and verbal communication.


Ability to support the following:

Experience with cloud providers - AWS, GCP, or Azure.

Exposure to containerization technologies such as Docker and Kubernetes.

Familiarity with infrastructure provisioning using Terraform.

Experience introducing SRE practices (SLOs, error budgets, chaos engineering) at an organizational level.

Exposure and experience with migrating/ building AI tools to improve process.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8760168
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
19/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Platform Team Leader to own the foundational systems that power our R&D organization. The Platform team is a horizontal, internal-facing team - its customers are the engineers, data scientists, and analysts on our application teams, and its mission is to make them faster, safer, and more productive.

You will lead a team responsible for our production-critical data ingestion (scraping system, data normalization platform), our developer experience surface (Python mono repo, packages and services framework, CI/CD, shared images), our Airflow base layer, and our backoffice. These systems sit on the critical path for production - every other team depends on them, and uptime, incident response, and reliability are core to the role.

This role is for a hands-on technical leader who enjoys both shipping code and growing engineers. You'll spend roughly 50% of your time hands-on writing code, reviewing designs, and the rest leading the team, partnering with peer team leads, and shaping the technical direction of our platform with the VP R&D and CTO.

Responsibilities
Lead and grow a team of strong backend engineers, owning hiring, mentorship, performance, and personal growth.
Own the team's roadmap end-to-end: scoping, prioritization, delivery, and quality.
Drive the architecture and evolution of the platform layers your team owns - both the data plumbing that feeds us and the developer-experience surface other R&D teams build on.
Contribute hands-on (~50%) to design, code, code reviews, and production debugging.
Set and uphold engineering standards across the team - code quality, system design, testing, observability, and operational excellence.
Partner closely with peer team leads and the VP R&D to align on shared infrastructure, ownership boundaries, and developer experience.
Lead by example on production incidents and customer-impacting issues.
Influence broader R&D direction as a member of the R&D leadership forum.
Requirements:
Requirements
3+ years of experience as a team leader in a product-oriented R&D organization.
6+ years of hands-on Python development experience.
Strong background designing and operating distributed systems and microservices in production.
Hands-on experience with Kubernetes and Docker in production environments.
Solid working knowledge of GCP (or equivalent cloud) - Cloud Run, BigQuery, GCS, IAM.
Experience designing and operating CI/CD pipelines in a microservices / mono repo environment.
Experience with Airflow or similar orchestration frameworks.
Proven ability to balance hands-on contribution with team leadership at ~50/50.
Strong sense of ownership, end-to-end accountability, and a builder's mindset.
Excellent collaborator - comfortable working across teams and aligning on shared infrastructure ownership.

Advantages
Experience leading a platform / infrastructure / DevX team specifically (vs. a feature team).
Working knowledge of Helm charts and Terraform / Terragrunt as a consumer (you don't need to own them, but you should be able to read and reason about them).
Experience with web scraping systems at scale.
Experience with Python packaging, mono repo tooling, and shared library design.
Comfortable using GenAI tools (Cursor, Claude, etc.) as part of your engineering workflow.
B.Sc. in Computer Science or equivalent.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8744019
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a strong technical leader to lead our Runtime team - one of the most critical backend teams. The Runtime team owns the execution layer of all our automations, the core of our product. The team operates at a significant scale, processing hundreds of millions of steps and millions of security events every day in systems that simply cannot break.

You will lead a strong team of five senior engineers and take ownership of both the people and the technology. This role combines team management with deep technical leadership: developing engineers, setting the technical vision, leading design and architecture decisions, and staying close enough to the code to step in when needed.

What Will You Do?
Lead, mentor, and develop a team of five experienced backend engineers.
Own the teams technical vision, priorities, execution, and long-term direction.
Lead the design and implementation of architectural changes that keep our platform reliable at massive scale.
Drive complex, cross-team projects and influence technical decisions beyond your immediate team.
Lead production incidents and post-mortems, and continuously raise the bar for reliability and operational excellence.
Identify risks and problems before they impact production, and build the processes and infrastructure needed to prevent them.
Stay close to the technology through design discussions, code reviews, debugging, and hands-on contribution when needed.
Use data to understand how customers interact with the platform and turn emerging use cases into new product capabilities.
Help lead significant technical initiatives, including the evolution from GCP toward a Multi-Cloud architecture.
Work with modern AI tools such as Cursor and Claude Code as a core part of the development workflow, while maintaining deep technical understanding and high-quality standards.
Requirements:
What Should You Bring to the Table?
At least 2 years of experience managing an engineering team.
A strong backend engineering background with experience building complex, high-scale production systems.
Proven ability to lead technical vision, System Design, and architecture discussions.
Experience leading meaningful projects end-to-end and working across multiple teams.
Strong production ownership, including leading incidents, post-mortems, and reliability improvements.
A hands-on mindset and the ability to step into code, debugging, and code reviews when needed.
Experience mentoring and developing senior engineers.
Experience with Cloud infrastructure and services.
Familiarity with Kubernetes, Docker, CI/CD, monitoring, and observability technologies.
A strong testing mindset, including automated testing and production monitoring.
Experience with Go - an advantage.
Excellent communication skills and the ability to explain complex decisions clearly and simply.
High ownership, curiosity, and a proactive, can-do approach.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761130
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for an experienced Full Stack Team Leader to lead the engineering team behind our Customer Portal-the central platform used by every customer across all of our products. This role reports directly to the VP R&D and combines technical leadership with product ownership, cross-functional collaboration, and execution.

As the engineering leader of the Customer Portal, youll shape one of the companys most strategic products, ensuring customers enjoy a seamless experience throughout their entire journey with us. Youll lead a talented team, drive technical excellence, champion AI adoption, and collaborate closely with stakeholders across Product, Marketing, Customer Success, Support, Sales, and Design to deliver customer-facing capabilities with direct business impact.

Were looking for someone who doesnt wait for direction. You naturally identify opportunities, challenge the status quo, and drive initiatives from idea to execution. You inspire your team through technical leadership, raise the engineering bar, and continuously look for better ways to build products and develop people.

What Youll Do
Lead, mentor, and grow a team of Full Stack Engineers, fostering a culture of ownership, accountability, innovation, and continuous improvement.
Own the technical execution and delivery of the Customer Portal roadmap, serving customers across all our products.
Partner closely with Product Management and business stakeholders to translate strategic objectives into scalable technical solutions.
Collaborate daily with Product, Marketing, Customer Success, Support, Sales, and Design to deliver an outstanding customer experience.
Lead architectural decisions, ensuring the platform remains scalable, secure, maintainable, and ready for future growth.
Drive engineering excellence by establishing high standards for code quality, testing, observability, performance, reliability, and operational excellence.
Champion AI-first engineering practices by driving adoption of AI tools and workflows that improve developer productivity, code quality, and delivery velocity.
Proactively identify opportunities to improve the product, engineering processes, and team effectiveness. Bring ideas, challenge existing approaches, and lead meaningful technical and organizational improvements.
Lead planning, prioritization, estimation, and execution while balancing business priorities with technical investments.
Remove blockers, manage technical risks, and create an environment where the team can consistently deliver high-quality outcomes.
Stay hands-on when needed, contributing to architecture, design, code reviews, and implementation of complex features.
Requirements:
Requirements
3+ years of experience leading software engineering teams.
6+ years of experience as a Full Stack Software Engineer.
Strong hands-on experience with React, Node.js, REST APIs, and modern web application architecture.
Experience building and operating customer-facing SaaS platforms at scale.
Strong understanding of distributed systems, relational and NoSQL databases.
Proven ability to lead complex projects from ideation through production.
Demonstrated track record of driving technical initiatives, leading change, and continuously improving engineering practices.
Passion for leveraging AI throughout the software development lifecycle and leading teams in adopting AI-powered development workflows.
Strong sense of ownership with the ability to independently identify problems, propose solutions, and drive execution.
Excellent communication and stakeholder management skills, with experience working across Engineering, Product, Marketing, Customer Success, Support, and other business functions.
Deep understanding of software architecture, secure coding, modern development practices, CI/CD, and cloud-native applications.
Passion for building high-quality software and establishing a strong engineering quality culture.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761015
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
06/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Required ML Platform Engineering Team Lead - Sovereign AI Engineering
We're building AI that nations own and control, deployed where almost no one else can operate. ingesting and structuring complex data, and driving practical actions that can literally impact the lives of billions of people around the world. This role helps make that real.
The Dream Job
It starts with you - a technical leader driven to build both the ML platform and the engineering team behind it. You care about reliable infrastructure, great developer experience, and growing engineers through real ownership. You'll set the technical direction for our ML platform - training pipelines, model serving, feature stores, experiment tracking, and compute orchestration - shaping how models reach production across cloud and on-prem, including air-gapped deployments. A significant part of the platform supports large language models, with unique challenges across training, evaluation, and inference in mission-critical environments. You stay close enough to the codebase to debug production issues, unblock your engineers, and make sound architecture calls.
If you want to make a meaningful impact, join our mission and lead the team that builds the ML platform driving Sovereign AI products - this role is for you.
Responsibilities
Set technical direction for the ML platform - training pipelines, model serving, feature stores, experiment tracking, and compute orchestration - through RFCs, prototypes, design reviews, and build-vs-buy decisions
Lead and grow a team of ML Engineers - hire, mentor, pair on hard problems, and raise the bar through code and design reviews
Contribute to critical systems, debug production issues, and maintain deep context on the codebase to inform technical decisions
Own operational excellence for model serving - set and enforce SLAs, run capacity planning, and keep compute costs predictable
Establish ML engineering standards - reproducible experiments, automated evals, model packaging, CI/CD for models, and observability
Support the full lifecycle of our models - from training on domain-specific data to low-latency inference powering production systems
Work closely with Data Platform, AI, Data Science, and Product teams - translate business priorities into engineering work and manage cross-team dependencies
Measure and improve developer experience - deploy friction, onboarding time, CI turnaround - as seriously as model performance.
Requirements:
6+ years in software engineering, ML engineering, or platform engineering, with hands-on experience building and operating ML infrastructure at scale.
2+ years leading an engineering team - hiring, mentoring, conducting design reviews, and shipping alongside your team
Engineering craft - Strong Python, distributed systems design, testing, secure coding, API design, CI/CD discipline, and production ownership.
ML platform & serving - Model serving frameworks (e.g., Triton, TorchServe, vLLM, Ray Serve); model packaging, deployment pipelines, and inference optimization
Training infrastructure - Distributed training pipelines (e.g., frameworks like PyTorch, JAX) experiment orchestration and reproducibility
ML lifecycle tooling - Feature stores, model registries, experiment tracking (e.g., MLflow, Weights & Biases); dataset versioning and lineage
Data pipelines - Building training and inference data pipelines; familiarity with tools like Spark, Airflow/Dagster, and streaming ingestion
Comfortable with AI coding tools like Cursor, Claude Code, or Copilot
Nice to Have:
Experience operating in constrained environments - on-premise, private cloud, or air-gapped deployments
Hands-on experience with simulation environments, synthetic data generation, or reinforcement learning workflows
Platform & infra - Kubernetes, AWS, Terraform or similar IaC, CI/CD, observability, incident response
Hands-on data science or applied ML experience.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8725283
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a strong technical leader to lead one of our product-focused R&D teams at Torq.

This team owns meaningful product capabilities used across the Torq platform and works closely with Product, Design, and other engineering teams. The role combines people management, technical leadership, and strong product thinking - from understanding customer needs and shaping solutions to delivering high-quality features in production.

You will lead a team of experienced engineers, develop their skills, set clear priorities, and create an environment of ownership, collaboration, and continuous improvement. At the same time, youll stay close to the technology, lead design decisions, and step into the code when needed.

What Will You Do?
Lead, mentor, and develop a team of engineers.
Own the teams execution, priorities, and long-term direction.
Work closely with Product and Design to turn customer needs into simple, scalable solutions.
Lead features and projects from early definition and design through production.
Drive technical design and architecture decisions while balancing product needs, quality, and delivery.
Stay close to the technology through design discussions, code reviews, debugging, and hands-on contribution when needed.
Build strong collaboration with other teams and lead cross-team initiatives.
Create clear processes and ownership as the team and its scope continue to grow.
Take part in hiring and help shape the teams professional standards and culture.
Use modern AI tools such as Cursor and Claude Code as part of the development workflow while maintaining high-quality engineering standards.
Requirements:
At least 2 years of experience managing an engineering team.
A strong software engineering background in Backend or Full Stack development.
Experience leading meaningful product features and projects end-to-end.
Strong System Design and architecture skills.
A product mindset and the ability to understand customer needs and business impact.
Experience working closely with Product, Design, and other R&D teams.
The ability to develop engineers, provide clear feedback, and build a high-performing team.
A hands-on mindset and the ability to step into code and technical discussions when needed.
Experience with Cloud environments, production systems, and modern engineering practices.
Excellent communication skills and the ability to explain complex decisions clearly.
High ownership, curiosity, and a collaborative, low-ego approach.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761120
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a Senior Infrastructure Engineer who views "Infrastructure as Software." In 2026, we dont just manage servers; we build high-performance environments that allow multi-agent systems to operate at scale.
You will be a core member of the R&D team, blending deep DevOps expertise with the coding rigor of a Backend Engineer. You arent just "configuring" AWS; you are architecting the distributed systems and data pipelines that power our autonomous security brain. Your mission is to ensure that while our agents are evolving and taking actions, our underlying platform remains immutable, observable, and infinitely scalable.
What You'll Do
Design, build, and operate our company's cloud infrastructure using AWS, Kubernetes, and Infrastructure as Code.
Build internal tools and platform services using Python and Go to improve developer productivity and system reliability.
Own infrastructure automation with Terraform, Pulumi, and modern cloud-native tooling.
Partner closely with Backend, Data Science, and Security Engineering teams to build scalable, reliable platforms.
Improve observability, monitoring, and incident response across distributed production systems.
Design and optimize infrastructure for performance, scalability, security, and cost efficiency.
Help shape engineering best practices, platform architecture, and developer experience as our company continues to grow.
Requirements:
5+ years of experience in Infrastructure, DevOps, Platform Engineering, or Backend Engineering.
Strong software engineering skills with hands-on experience building production systems in Python or Go.
Deep hands-on experience with AWS, including services such as EKS, RDS, VPC, and IAM.
Strong experience designing, operating, and scaling production Kubernetes environments.
Experience with Infrastructure as Code, CI/CD, GitOps, and modern cloud-native development practices.
A systems mindset with the ability to solve architectural challenges across infrastructure and application layers.
Comfortable using modern AI-powered developer tools and agentic workflows to improve engineering productivity.
The company Mindset: You take ownership, act with accountability, collaborate openly, and focus on delivering meaningful impact. You thrive in fast-moving environments, embrace ambiguity, and enjoy solving hard problems together.
Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.
Full professional fluency (written and verbal) in both Hebrew and English.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8764502
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
3 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769987
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for an Engineering Team Lead, who will be responsible for the foundational infrastructure framework used by all our engineering teams to build, deploy, and operate AI agents safely in production. We are building the "operating system" for AI, covering agent sessions, memory management, tool orchestration, durable execution, and multi-tenant isolation. You will lead a high-impact team of 5 engineers to create the runtime and platform that defines the future of autonomous enterprise intelligence.
Youll Own:
Agentic Framework Architecture: Designing and building our internal agentic framework, leveraging and integrating industry-standard tools such as LangChain, LangSmith, ADK, and similar ecosystems.
Evaluation and Quality Systems: Building evaluation frameworks and workflows for AI agents, including offline and online evaluations, quality metrics, regression detection, and experimentation infrastructure.
Team Leadership & Mentorship: Leading a squad of 3-4 senior engineers, fostering a culture of technical excellence, and managing end-to-end delivery in a fast-paced environment. You will spend approximately 50% of your time hands-on, architecting core systems and reviewing code, and 50% leading the team, mentoring engineers, and aligning with cross-functional stakeholders.
Observability, Monitoring, and Guardrails: Providing the organization with robust observability capabilities for AI agents, including tracing, logging, monitoring, cost tracking, and safety guardrails to ensure reliable and responsible usage.
Developer Enablement Platforms: Creating APIs, SDKs, and abstractions that enable product teams to easily build, test, and operate agents while adhering to platform standards.
Cross-Language Integrations: Designing integrations and tooling across Python and Java to enable seamless adoption of the AI framework within our broader backend ecosystem.
Youll Solve:
Agent Lifecycle and Orchestration Complexity: Managing agent execution, tool usage, memory, workflows, and failure modes in production-grade systems.
AI System Reliability at Scale: Ensuring agents remain observable, debuggable, and safe as usage scales across teams and products.
Evaluation and Drift Challenges: Detecting quality regressions, model behavior changes, and unintended agent behaviors through robust evaluation and monitoring systems.
Platform Adoption Friction: Balancing flexibility with guardrails so teams can innovate quickly without compromising reliability, security, or cost controls.
Youll Impact:
Company-Wide AI Enablement: Empowering every engineering team to build agent-based solutions faster, with higher quality and confidence.
Foundational AI Infrastructure: Establishing the core frameworks, evaluations, and observability standards that all AI agents will rely on.
AI Safety and Quality Bar: Raising the bar for how AI systems are evaluated, monitored, and governed across the company.
Requirements:
8+ years of backend engineering experience, with strong system design and platform-building expertise. Tech leadership or team leading experience is an advantage.
Strong analytical and problem-solving skills, with the ability to debug and resolve complex technical issues efficiently.
Hands-on experience with agentic systems and frameworks such as LangChain, LangSmith, ADK, or equivalent agent orchestration platforms.
Strong understanding of AI evaluation methodologies, including agent evaluations, prompt evaluation, regression testing, and quality monitoring.
High proficiency in Python for building production-grade AI frameworks and services.
Familiarity with Java and experience integrating backend platforms or tooling into Java-based systems.
Experience building observability, monitoring, or platform tooling for distributed systems.
Strong analytical skills and the ability to reason about complex, evolving AI-driven systems.
Experience with cloud platforms and scalable microservices architectures.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8728090
סגור
שירות זה פתוח ללקוחות VIP בלבד