דרושים » תוכנה » Software Quality Engineering IC1

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
09/06/2026
משרה זו סומנה ע"י המעסיק כלא אקטואלית יותר
מיקום המשרה: תל אביב יפו
סוג משרה: משרה מלאה
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
04/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a Network Solution Verification Manager to join our R&D organization, leading the validation of our customer-facing networking solutions at scale using a state-of-the-art End-to-End (E2E) simulation cluster environment.

In this role, you will build and lead a team of validation engineers, owning the strategy, methodology, and execution of simulation-based validation frameworks that ensure our networking solutions meet the highest standards of quality, scale, and real-world applicability before reaching customers.


What you'll be doing:
Leading and managing a team of network validation engineers responsible for end-to-end validation of our customer-facing networking solutions at scale, using a dedicated E2E simulation cluster environment.
Defining the overall validation strategy and roadmap - establishing simulation methodologies, test coverage frameworks, and quality gates that align with product milestones and customer use cases.
Pioneering the use of agentic AI flows within the validation organization - leading the team to design, build, and operate AI-driven agents capable of autonomously performing regression analysis, identifying coverage gaps, generating new test cases, and implementing validation code. These agentic workflows will continuously learn from simulation results and product changes, dramatically accelerating the team's ability to scale test coverage and respond to emerging quality signals without manual intervention.
Overseeing the design and continuous improvement of automated regression suites for networking protocols and large-scale simulation runs, ensuring scalable, repeatable, and high-confidence validation outcomes.
Establishing a rigorous regression analysis culture - guiding the team in identifying trends, root causes, and systemic coverage gaps, and ensuring timely resolution in collaboration with engineering stakeholders.
Serving as the primary validation partner to Design, Architecture, and NCS teams - translating solution requirements into simulation scenarios, providing early-cycle quality feedback, and influencing product direction.
Analyzing customer-reported networking solution issues, driving test gap analysis, and ensuring robust regression coverage that prevents recurrence.
Staying current with emerging networking standards, simulation technologies, and industry best practices to continuously evolve the team's validation capabilities.
דרישות:
What we need to see:
B.Sc. degree or equivalent experience in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
3+ years of experience in a leadership or management role, leading software or hardware validation/test engineering teams.
7+ years of overall experience in network validation, network testing, or systems verification.
Proven track record of building and executing test automation strategies for network or distributed systems at scale, with hands-on background in Python-based automation.
Strong understanding of regression analysis methodologies - ability to drive actionable conclusions from large-scale test result datasets and translate them into engineering improvements.
Demonstrated ability to collaborate cross-functionally with architecture, design, and product teams in a fast-paced, multi-timezone environment.
Strong verbal and written communication skills, with experience presenting validation strategies and quality metrics to senior leadership.


Ways to stand out from the crowd:
Deep understanding of networking protocols and architectures (e.g., BGP, EVPN, VXLAN, RDMA/RoCE, Ethernet, IP routing, L2/L3 switching).
Experience with network simulation or emulation environments (e.g., Containerlab, GNS3, SONiC testbeds, or equivalent platforms) and the ability to guide teams in leveraging them effectively.
Experience managing validation of data center networking solutions or hyperscale network environments (spine-leaf, fat-tree, or Clos top#EN המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8768267
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a Staff Systems Software Engineer to design and build the foundational infrastructure that powers products used by billions of people worldwide. In this role, you will architect and implement large-scale distributed systems, low-level platform components, and high-performance services that underpin our core product stack. You will drive technical strategy across system reliability, performance, and scalability, partnering closely with product, infrastructure, and data engineering teams to deliver systems that operate at global scale with high availability and efficiency.
Software Engineer, Systems Responsibilities
Architect and implement large-scale distributed systems and platform services that support high-throughput, low-latency workloads across our product infrastructure
Lead the technical design of systems components including storage layers, compute pipelines, networking abstractions, and service orchestration frameworks
Identify and resolve systemic performance bottlenecks through instrumentation, profiling, and targeted optimization across the full systems stack
Define and enforce service level objectives for owned systems, building dashboards, alerting pipelines, and runbooks to reduce mean time to mitigation during incidents
Drive reliability improvements by reducing failure surface, designing resilient rollout strategies, and leading regular resiliency and overload testing exercises
Collaborate with cross-functional partners across product engineering, infrastructure, and data science to align system architecture with evolving product and business requirements
Establish and evolve coding standards, architectural patterns, and engineering best practices for systems development across the broader organization
Leverage AI-assisted development workflows to accelerate design iteration, code generation, and systems analysis, applying sound judgment on when to rely on AI versus deep systems expertise
Mentor other engineers on systems design principles, debugging methodologies, and production operations, and contribute to onboarding programs for new team members
Lead incident retrospectives, identify root causes of complex production failures, and drive implementation of systemic improvements to prevent recurrence
Requirements:
Minimum Qualifications
8+ years of experience designing and implementing large-scale distributed systems, platform infrastructure, or systems software in production environments
Experience leading major technical initiatives end-to-end, including architecture design, cross-team coordination, staged rollout, and post-launch reliability ownership
Experience debugging complex, non-reproducible systems issues including concurrency bugs, memory management failures, and distributed consistency problems
Experience defining service level objectives, building observability infrastructure, and driving reliability improvements across production systems
Experience communicating technical architecture decisions and trade-offs in writing to both engineering and non-engineering stakeholders
Preferred Qualifications
Experience with systems programming languages such as C, C++, or Rust in the context of high-performance or low-latency infrastructure
Experience building or improving developer tooling, automation frameworks, or internal platforms that measurably improve engineering efficiency across teams
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8777005
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
04/08/2026
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are seeking a Network Solution Verification Engineer to join our R&D organization, focused on validating our customer-facing networking solutions at scale using a state-of-the-art End-to-End (E2E) simulation cluster environment.

In this role, you will be a hands-on technical contributor at the intersection of networking, automation, and AI-driven validation - designing and implementing the simulation frameworks, agentic workflows, and regression pipelines that ensure our networking solutions meet the highest standards of quality and real-world applicability before reaching customers.


What you'll be doing:
Designing and implementing end-to-end validation frameworks for our customer-facing networking solutions at scale, leveraging a dedicated E2E simulation cluster environment.
Writing, maintaining, and extending automated test suites and regression pipelines for networking protocols and large-scale simulation runs, ensuring repeatable, high-confidence validation outcomes.
Performing deep regression analysis on simulation results - identifying failure trends, isolating root causes, and delivering clear, actionable findings to architecture and design teams.
Developing agentic AI flows that autonomously perform regression analysis, detect coverage gaps, generate new test cases, and implement validation code - continuously learning from simulation results and product changes to accelerate coverage without manual intervention.
Integrating validation pipelines into CI/CD workflows to enable continuous, automated regression at scale, working closely with DevOps and platform teams.
Collaborating closely with Design, Architecture, and NCS teams to understand solution requirements, translate them into simulation scenarios, and provide early-cycle quality feedback that influences product direction.
Analyzing customer-reported networking issues, mapping them to simulation coverage gaps, and building targeted test cases that prevent regression.
Continuously exploring new simulation technologies, agentic frameworks, and networking standards to evolve and improve the team's validation methodology.
Requirements:
What we need to see:
B.Sc. degree or equivalent experience in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
3+ years of hands-on experience as a software developer.
Strong proficiency in Python for test automation, tooling development, and pipeline implementation.
Hands-on experience with network simulation or emulation tools (e.g., Containerlab, GNS3, SONiC testbeds, or equivalent platforms).
Proven experience designing and building simulation agents or traffic generators that mimic real-world networking behavior at scale.
Solid experience with agentic AI frameworks and LLM-based automation (e.g., LangChain, LangGraph, AutoGen, or similar) and practical ability to apply them to validation and test generation workflows.
Strong command of regression analysis methodologies - able to triage, classify, and extract actionable conclusions from large-scale test result datasets.
Comfortable operating in a fast-paced, cross-functional, multi-timezone engineering environment with strong verbal and written communication skills.

Ways to stand out from the crowd:
Hands-on experience validating data center networking solutions or hyperscale network environments (spine-leaf, fat-tree, or Clos topologies).
Familiarity with our networking products - BlueField DPUs, ConnectX NICs, Spectrum switches, or the DOCA software stack.
Deep understanding of networking protocols and architectures (e.g., BGP, EVPN, VXLAN, RDMA/RoCE, Ethernet, IP routing, L2/L3 switching).
Hands-on experience building agentic pipelines for automated test generation, result triage, or validation code synthesis - including prompt engineering and tool-use patterns for LLM agents.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8768268
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
2 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
About us our company's mission is to protect every mobile app in the world and the people who use them. We are the leader in AI-native mobile business protection, providing cyber and fraud teams with an agentic platform that builds, monitors, and maintains security defenses in Android and IOS apps - with no SDKs, no coding, and no disruption to engineering cycles.
Our platform delivers over 400 security, anti-fraud, anti-bot, and API protection capabilities, powered by deep learning models trained on a decade of mobile defense data and trillions of live threat events. From build time to runtime, our company's AI Agents help mobile brands detect, investigate, and respond to threats faster than ever - recognized as the best AI Platform for Cyber Resilience at RSA Conference 2026 for the second consecutive year. Leading financial, healthcare, m-commerce, and B2B brands rely on our company to secure over 50,000 mobile apps and protect more than 1 billion end users globally.
our company is an Equal Opportunity Employer. We are committed to diversity, equity, and inclusion in our workplace. We do not discriminate based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other characteristic protected by law. All qualified applicants will receive consideration for employment without regard to any of these characteristics.
About the Role We are looking for a talented software engineer to join our company's Identity Group. You will take part in building advanced fraud detection capabilities, working closely with security researchers and data teams to develop cutting-edge solutions that protect hundreds of millions of mobile devices.
Responsibilities
* Design and develop core components of fraud detection systems from early prototyping through production-ready releases.
* Build scalable backend services and distributed systems that process data at scale.
* Work with large-scale datasets and collaborate with data teams to create actionable signals and features.
* Write high-performance, production-quality code in C and C ++ optimized for reliability and performance.
* Take ideas from rapid proof-of-concept through to production, tackling complex engineering challenges across the full development lifecycle.
Requirements:
* B.Sc. in Computer Science, Software Engineering, or equivalent.
* At least 2 years of experience developing complex enterprise systems.
* Experience in C / C ++ programming.
* Experience in Python, JAVA, Linux, and Git.
* Strong focus on performance, scalability, and reliability in distributed systems.
* Comfortable working with large datasets and collaborating with data science and research teams.
* Team player with strong collaboration and communication skills.
* Quick learner with ability to adapt to new tools, methods, and frameworks.
Preferred Qualifications
* Background in fraud prevention, cybersecurity, or security analytics.
* Background in mobile app research or development ( Android / IOSgreenTxtBg!).
* Experience with Machine Learning or data science.
* Experience building high-throughput data processing systems.
* Understanding of threat landscapes and attack vectors.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8778620
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769987
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
28/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Technical Lead to drive the architectural direction and engineering excellence of this group. This is a senior, deeply hands-on role for a technology leader who can own the technical roadmap, mentor a team of elite engineers, and build the infrastructure that challenges platform to its theoretical limits.
What You'll Lead:
Define and own the technical architecture of the group's distributed testing and reliability platform - designing for massive scale, real-world workload simulation, and adversarial failure injection
Lead effort involving multiple engineers, setting technical standards, running architecture reviews, driving design decisions, and mentoring engineers to grow
Build the systems that orchestrate millions of concurrent IO operations, inject chaos at the infrastructure layer (latency, packet loss, hardware failures), and expose the hardest-to-find race conditions and consistency bugs
Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines
Drive observability and reliability engineering across the group - building telemetry pipelines that track P99 latency, jitter, and system health, turning quality into a quantitative discipline
Collaborate deeply with Core R&D, Storage Kernel, and Infrastructure teams - translating architectural knowledge into targeted reliability strategies
Establish engineering practices - design docs, production-grade code reviews, testing philosophy, and cross-team technical alignment
Requirements:
Strong software engineering background with 6+ years of hands-on Python development experience is required. The ability to read, debug, and reason about C++, Rust, or Go is a significant advantage
Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system behavior under stress
Background in one or more of: storage systems, networking (TCP/IP, RDMA), cloud infrastructure, database internals, or high-performance backend systems
Experience building large-scale infrastructure platforms, internal developer platforms, or reliability engineering systems
Leadership:
Proven track record leading complex technical initiatives from architecture through delivery
Experience mentoring and growing engineers - raising the technical bar of a team, not just directing work
Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed
Comfortable operating at both the strategic and hands-on level - you write code, review designs, and shape roadmaps
Previous experience in people management roles - Advantage
Mindset:
You approach quality through the lens of Site Reliability Engineering: you care about MTTD, observability, and building self-healing systems
You have a "hacker" instinct - you don't just find bugs; you find the architectural flaws that allowed them to exist
You are an early adopter of AI tools and excited about applying LLMs and generative AI to accelerate engineering velocity
Big Advantages
Experience with storage systems, file systems, or high-performance distributed environments
Background in chaos engineering, fault injection, or simulation systems
Familiarity with observability tooling and performance engineering at scale
Experience building testing or reliability platforms as first-class engineering products
Prior experience as a Team Lead in a high-growth infrastructure company
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8757543
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779592
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
15/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. we are building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are constantly striving to make our systems reliable, scalable, and simple to operate so our services are available to travelers when they need them most. With our continued growth, we have exciting challenges ahead and we're looking for a Senior Site Reliability Engineer to join our team in Tel Aviv. This role blends classic SRE ownership with pragmatic AI SRE work: you will build and operate the platforms, automation, observability, and incident response practices that keep our company reliable, while helping teams use AI solutions, AI providers, and their APIs safely and dependably.
This is a hands-on engineering role, not a research role. You will partner with product, platform, data, security, support, and incident response teams to make production systems and AI-powered experiences more resilient. You will use software engineering, infrastructure as code, SLOs, telemetry, provider observability, and automation as your main tools, and you will apply AI where it creates measurable reliability value rather than novelty.
This position is based out of our new Tel Aviv office.
What You'll Do:
Support AI-based application solutions where reliability matters. Partner with the development teams building AI-powered travel experiences to support the development and production operation of their solution.
Work with AI solutions, providers, and APIs. Partner with teams integrating AI capabilities and providers, with attention to API reliability, authentication, quotas, rate limits, latency and provider-specific operational constraints.
Troubleshoot AI tools and provider issues. Diagnose failures across AI-powered workflows, provider APIs, configuration, permission errors, degraded responses and related areas.
Operate reliable production platforms. implement and run cloud infrastructure,and help product teams move quickly without compromising reliability.
Improve observability. Build dashboards, alerts, traces, logs, and runbooks that make service health clear, actionable, and tied to SLOs and customer impact.
Apply AI to SRE workflows. Prototype and productionize AI-assisted systems that create effective and efficient operations
Automate operational toil. Create tools, workflows, and automation that remove repetitive manual work and make operational knowledge easier to use.
Requirements:
5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
3+ years of experience operating production, 24x7 customer-facing systems.
Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation.
Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8739673
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a passionate and detail-oriented Product Owner & Engineer to join our team. In this role, you will collaborate closely with development teams to define new features, contribute to functional requirement documentation (FRD), and ensure seamless UX design and implementation.
A core part of this role is maintaining close alignment with customers. You will actively engage with customers, gather feedback, understand real-world use cases, and translate their needs into clear product requirements. You will work closely with Support and QA to ensure our testing strategy and supportability reflect actual customer environments and workflows.
Beyond core development, you will assist in critical escalations, manage complex or high-visibility installations, and develop tools that enhance the overall user experience. You will leverage telemetry data, production insights, and direct customer feedback to continuously refine and improve our products, ensuring they deliver measurable value in real-world deployments.
Key Responsibilities
Collaborate closely with Product and Engineering teams to define, refine, and prioritize features that directly address real customer needs and business impact.
Translate customer requirements and field insights into clear, structured Functional Requirement Documents (FRDs), while actively contributing to UX discussions to ensure intuitive and seamless user experiences.
Work closely with QA and Support to align testing strategies and troubleshooting workflows with real-world customer environments, ensuring reliability, operability, and supportability at scale.
Serve as a technical and product focal point during critical customer escalations and high-visibility deployments, ensuring timely resolution and long-term improvements
Develop tools and scripts to enhance user experience and operational efficiency.
Leverage telemetry, usage analytics, and direct customer feedback to drive data-informed decisions and continuously improve product performance and adoption.
Proactively identify risks, gaps, and cross-team dependencies, removing roadblocks to ensure successful delivery and measurable customer outcomes.
Requirements:
Qualifications & Skills
Proven experience in software development, technical product management, or advanced technical support within complex infrastructure environments.
Demonstrated experience with storage technologies and protocols such as NFS, S3 (object storage), SMB, and familiarity with enterprise storage architectures and distributed file systems.
Strong hands-on experience with Linux, Python, and networking concepts (TCP/IP, routing, switching, large-scale deployments).
Ability to analyze and solve complex technical challenges in scale-out Linux environments, HPC workloads, AI training infrastructures, and advanced networking architectures.
Experience collaborating across cross-functional teams - Engineering, QA, and Support - using industry-standard tools such as Jira, Slack, GitLab, Git, unit testing frameworks, and QTest.
Strong analytical skills with the ability to leverage telemetry, usage data, and customer insights to guide product decisions and prioritize effectively.
Experience working with observability and data platforms, including time-series databases (e.g., Prometheus), multi-tenant log aggregation systems, Slack and Salesforce integrations, and AI-driven automation workflows - a strong advantage.
Familiarity with scripting and automation using Python, REST APIs, OpenTelemetry (OTEL), and Bash to improve operational efficiency and supportability.
Excellent communication skills, with the ability to bridge technical depth and customer-facing clarity.
A proactive, customer-first mindset with strong ownership and accountability.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8743996
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
28/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Software Engineer - Verification and Reliability.
In this role as a SDET (Software Development Engineer in Test), you are a developer first. You will join a high-impact team of engineers who write production-grade code to build a massive-scale validation ecosystem. Your job is to act as "The Breaker"-designing the infrastructure, chaos experiments, and AI-driven tools that push our platform to its theoretical limits.
What Youll Build:
Adversarial Engineering: Design and implement Python-based distributed frameworks capable of orchestrating millions of concurrent IO operations to hunt down race conditions and memory leaks.
AI-Augmented Validation: Be at the forefront of the AI-Native transformation. You will leverage LLMs and Generative AI to automate complex scenario generation, build intelligent agents for root-cause analysis, and multiply your engineering velocity.
Simulation & Chaos: Build the "Entropy Engine." You will develop tools that inject real-world failures - latency, packet loss, and hardware crashes - to prove the resilience of our Raft and RDMA implementations.
Deep-System Observability: Move beyond "Pass/Fail." You will build telemetry pipelines to track P99 latency and jitter, providing critical architectural feedback to the Core Kernel teams.
Collaborative Architecture: You will operate with the same rigorous standards as the Core R&D team: design docs, production-grade code reviews, and high-level architectural planning.
Requirements:
Extensive Coding Experience: 5+ years of hands-on Python development experience is required. You are a Python expert who understands the language "under the hood" and are comfortable reading and debugging C++, Rust, or Go to understand how the core system works.
Systems Engineering Mindset: You have a background in distributed systems, networking (TCP/IP, RDMA), or storage protocols. You understand the complexities of consistency and metadata at scale.
AI Enthusiast: You are an early adopter of AI tools (Copilot, LLMs) and are excited about using them to automate the most tedious parts of the engineering lifecycle.
The "SRE" Lens: You approach quality through the lens of Site Reliability Engineering. You care about observability, MTTD (Mean Time to Detection), and building self-healing testing loops.
Problem Hunter: You have a "hacker" instinct. You dont just find a bug; you find the architectural flaw that allowed it to exist.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8757535
סגור
שירות זה פתוח ללקוחות VIP בלבד