דרושים » תוכנה » Manager, Performance Research and Analysis

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 4 שעות
Location: Yokne`am and Tel Aviv-Yafo
Job Type: Full Time
We are seeking a highly skilled and versatile Performance Research and Analysis Manager to join our Performance Group. This role will drive end-to-end performance strategy and execution for our next-generation data centers and solutions based on GPU systems, NIC, Switch, DPU and Networking technologies. The ideal candidate will oversee, evaluating, and optimizing end-to-end AI GPU cluster-level performance for scaling out large scale distributed training and inference jobs communication. The role will focus heavily on RDMA, Networking Protocols, Collective Communication, Congestion Control, and Load Balancing algorithms. Secondarily, you will lead our DPUs and Storage technologies for N-S use cases to support AI Inference jobs. Third, you will drive our Performance Dashboards and Observability for cluster-level performance analysis from a stream line telemetry across NICs, Switches, GPUs, and NVlink.

What you'll be doing:

Drive end-to-end performance strategy, characterization, test plans, and optimization for our next-generation AI GPU clusters, focusing on large-scale distributed training and inference workloads.

Deeply evaluate and optimize our Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.

Work on performance research and analysis of our DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.

Drive the strategy for performance observability and dashboards across our next-generation data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.

Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.
Requirements:
What we need to see:

B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.

8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.

3+ years of experience as an engineering team manager leading technical performance or R&D teams.

Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).

Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.

Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.

Ways to stand out from the crowd:

Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.

Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.

Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.

Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837819
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
2 ימים
Location: Ra'anana and Yokne`am
Job Type: Full Time
We are looking for an Engineering Manager to join our Network System Validation group. You will work on system validating advanced networking solutions across our complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. In this role you will lead the validation direction and engineering excellence of one of our technology validation teams. This is a management role for a technology leader who can own the technical roadmap and execution, mentoring a team of high-performance engineers, and push our network to its speed-of-light limits. This role combines the development of methodologies and automation tools with system validation, performance analysis, and investigation of cutting-edge AI networking technologies at scale.

What youll be doing:

Lead, mentor, and coach a team of software development and system validation engineers.

Review system and product requirements, design validation methodologies, develop comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions.

Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.

Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution.

Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection.

Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations.

Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes.

Foster a team culture centered on software quality, accountability, and technical excellence.
Requirements:
What we need to see:

B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.

8+ overall years of experience in networking, system validation, or related domains.

3+ years of experience leading software or system development team.

Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.

Strong scripting and automation experience using Python, Bash, and/or Ansible.

Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).

Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.

Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.


Ways to stand out from the crowd:

Experience with large-scale clusters or distributed systems.

Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).

Background in performance analysis, Kubernetes, or cloud environments.

Background in chaos testing, fault injection, or simulation systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8835840
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are seeking an AI Networking Architect to join the Networking Research Group. This role bridges the gap between emerging AI workloads and the data center infrastructure that powers them, working at the intersection of AI applications, distributed systems, networking hardware, and software architecture. You will join a focused team of multidisciplinary engineers driving AI workload optimization through deep application understanding, network analysis, and end-to-end systems thinking. Your insights will directly shape our products across the full stack - from applications and software libraries to hardware architecture and physical design.


What you'll be doing:
Model the performance of complex AI workloads to identify bottlenecks and recommend system-level optimizations.
Analyze new AI models, distributed training techniques, and inference workloads to understand their infrastructure requirements.
Build simulation and hardware platforms, run real AI workloads on them, and develop analytical tools to evaluate trade-offs across compute, memory, storage, and network behavior.
Translate research insights and workload behavior into actionable software, xhardware, and networking architecture requirements.
Partner with architecture, software, and product teams to influence our future networking and AI infrastructure roadmaps.
Drive architectural innovation by applying deep workload analysis to production machine learning frameworks.
Requirements:
What we need to see:
B.Sc. or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
5+ years of relevant industry or research experience. This is not an entry-level position; coursework and personal projects do not substitute for production or research experience at scale.
Hands-on experience running and analyzing AI workloads on multi-node systems - distributed training or large-scale inference - including measuring where time and resources are actually spent.
Demonstrated performance analysis work: building analytical or simulation models of real systems, validating them against measurement, and identifying bottlenecks that led to design or deployment changes.
Strong systems-level thinking across the full AI stack, from model and framework behavior down through compute, memory, storage, and network.
Track record of translating research findings and workload analysis into concrete software and hardware specifications that engineering teams acted on.
Strong programming skills in Python and C/C++, applied to performance modeling, data analysis, and prototyping.


Ways to Stand Out from the crowd:
Deep understanding of data centers, network topologies, and communication protocols.
Familiarity with GPU clusters, collective communication, storage systems, and AI networking bottlenecks.
Experience with distributed training, distributed inference, or large-scale AI serving systems, including the performance metrics and deployment strategies that govern them.
Experience in agentic programming and AI tooling.
Track record of turning academic research into concrete software, hardware, or architecture requirements, and of leading complex multidisciplinary projects with measurable production impact.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837903
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Yokne`am
Job Type: Full Time
We are looking for a Technical Lead to join our Network System Validation group and lead the validation of advanced networking solutions across complex AI cluster environments. This is a deeply hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence. You will develop validation methodologies and automation frameworks, while working hands-on on debugging, performance analysis, and cutting-edge AI networking technologies at scale. Join us to help push our networking technologies to their limits and shape how next-generation AI infrastructure is validated.

What youll be doing:
Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions
Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution
Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection
Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes
Requirements:
What we need to see:
B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.
12+ years of experience in networking, system validation, or related domains.
Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.
Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).
Strong scripting and automation experience using Python, Bash, and/or Ansible.
Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress.
Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.
Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.

Ways to stand out from the crowd:
Experience with large-scale clusters or distributed systems.
Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).
Background in performance analysis, Kubernetes, or cloud environments.
Background in chaos testing, fault injection, or simulation systems .
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836095
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: Yokne`am
Job Type: Full Time
We are looking for a Technical Lead to join our Network System Validation group and lead the validation of advanced networking solutions across complex AI cluster environments. This is a deeply hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence. You will develop validation methodologies and automation frameworks, while working hands-on on debugging, performance analysis, and cutting-edge AI networking technologies at scale. Join us to help push our networking technologies to their limits and shape how next-generation AI infrastructure is validated.

What youll be doing:

Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions

Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.

Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution

Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities

Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection

Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations

Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes
Requirements:
What we need to see:

B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.

8+ years of experience in networking, system validation, or related domains.

Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.

Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).

Strong scripting and automation experience using Python, Bash, and/or Ansible.

Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress.

Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.

Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.

Ways to stand out from the crowd:

Experience with large-scale clusters or distributed systems.

Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).

Background in performance analysis, Kubernetes, or cloud environments.

Background in chaos testing, fault injection, or simulation systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837843
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are is looking for a strong technical senior architect to join us in shaping the future. Senior Architects are innovators who can translate business needs into workable technology solutions. Their expertise is deep and broad. They are hands on, producing both detailed technical work and high-level architectural designs.

As a Senior architect in the AI Networking Research team, you will explore technological challenges on accelerate networking and building AI data centers. Research new transport functions and semantics for optimizing AI workloads, AI systems communication and accelerations and much more. You will also be leading architectural and development efforts across numerous technological fields, related to the modern AI data center, such as distributed AI and deep learning solutions, data analytics, High Performance Computing (HPC), Software Defined Networking (SDN), virtualization, storage, and more.

What youll be doing:

Enhance our GPU Networking offerings for accelerating AI workloads, such as our Dynamo, our NIXL and our UCX, tailored to the unique requirements of AI workloads.

Co-design hardware features (e.g., in GPUs, DPUs, or interconnects) that accelerate data movement and enable new capabilities for inference and model serving.

Identify and evaluate new technologies, innovations and partner relationships for alignment with our technology roadmap and business value.

Lead architecture and design of new technologies and innovations such as runtime systems, communication libraries, AI-specific technologies.

Lead proof-of-concept development to evaluate and drive such technologies.
Requirements:
What we need to see:

Hold a M.Sc. or Ph.D. in Computer Science, Electrical or Computer Engineering from a leading university (or equivalent experience).

5+ years of industry experience (or equivalent) in system architecture, AI systems architecture, scaling of AI, Parallelism of AI frameworks, or deep learning training workloads.

Experienced in algorithm design, system programming, computer architecture and operating systems.

Experienced in virtualization, networking and storage.

Deep understanding of performance profiling and optimization techniques, together with defining and using hardware features.

Strong programming and software development skills.

Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment.

Ways to stand out from the crowd:

Shown research track record.

Have experience and passion for system architecture, CPU/GPU/memory/storage/networking.

Stellar communication skills.

Knowledge in Deep Learning frameworks and AI communication libraries (NCCL, UCX, MPI and equivalents).

Deep understanding of Inference and Training workloads and optimizations, like Prefill/Decode, data parallelism, Tensor parallelism, FDSP and others.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837892
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
Our Network Architecture, Modeling and Performance Insights group is seeking a skilled and driven AI Network Modeling Engineer to conduct advanced network research and optimization while utilizing and enhancing our advanced simulation tools.

In this role, you will be a key contributor to defining the architecture and improving network performance for AI and High Performance Computing (HPC) workloads. You will be responsible for modeling advanced network architectures, topologies, configurations, logic and traffic patterns, and analyzing their effect on end to end performance. You will develop expertise in our simulation tools and contribute directly to their development road map and prioritization, develop core features in the simulator, and expose the capabilities to additional users. You will take part in complex architecture decisions, define and verify the modeling assumptions, and put them to the test vs. real HW and SW performance. If you're passionate about tackling intricate challenges and contributing to comprehensive systems and working on innovative solutions, we want to hear from you.

What you'll be doing:

Directly impact the architecture of our next generation AI and HPC offerings through modeling, simulation and analysis.

Deep dive into network behavior for training and inference use cases under advanced network topologies, configurations and traffic patterns which represent real life AI workloads, and analyze their effect on the systems performance. Develop advanced simulation solutions while contributing modular and scalable code.

Actively support the HW and SW development life cycles, map existing and candidate features into the simulation domain to enable their analysis, impact assessment and optimization.

Take full independent ownership of the performance modeling roadmap for a defined set of features. Interface directly with internal clients, communicate the analysis conclusions effectively and iterate on them. Drive projects from concept to completion.

Improve the quality of the simulation as a software product, ensuring robustness and reliability.
Requirements:
What we need to see:

BSc or above in Computer Science, Electrical Engineering, or a related field

5+ years of hands-on coding experience with strong proficiency in C++ and Python. You must be comfortable navigating, optimizing, and contributing to a complex, large-scale codebase.

End to end system perspective - capable of bridging the gap between hardware behavior, micro-architecture, and software performance (latency/throughput).

Project ownership capability with proven ability to lead technical initiatives, prioritize features, and manage project lifecycles autonomously with minimal supervision.

Ways to stand out from the crowd:

Strong background in communication networks technology. Deep understanding of Ethernet, NVLink, and/or Infiniband technologies, data center infrastructure, and network protocols.

Familiarity with discrete event simulators (e.g., OMNeT++, SystemC, or proprietary architectural simulators).

Advanced degree or experience focusing on computer architecture, algorithms, networking, or distributed systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836080
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
Our Network Architecture, Modeling and Performance Insights group is seeking a skilled and driven AI Network Modeling Architect to conduct advanced network research and optimization while utilizing and enhancing our advanced simulation tools.

In this role, you will be a key contributor to defining the architecture and improving network performance for AI and High Performance Computing (HPC) workloads. You will be responsible for modeling advanced network architectures, topologies, configurations, logic and traffic patterns, and analyzing their effect on end to end performance. You will develop expertise in our simulation tools and contribute directly to their development road map and prioritization, develop core features in the simulator, and expose the capabilities to additional users. You will take part in complex architecture decisions, define and verify the modeling assumptions, and put them to the test vs. real HW and SW performance. If you're passionate about tackling intricate challenges and contributing to comprehensive systems and working on innovative solutions, we want to hear from you.

 

What you'll be doing:

Directly impact the architecture of our next generation AI and HPC offerings through modeling, simulation and analysis.

Deep dive into network behavior for training and inference use cases under advanced network topologies, configurations and traffic patterns which represent real life AI workloads, and analyze their effect on the systems performance. Develop advanced simulation solutions while contributing modular and scalable code.

Actively support the HW and SW development life cycles, map existing and candidate features into the simulation domain to enable their analysis, impact assessment and optimization.

Take full independent ownership of the performance modeling roadmap for a defined set of features. Interface directly with internal clients, communicate the analysis conclusions effectively and iterate on them. Drive projects from concept to completion.

Improve the quality of the simulation as a software product, ensuring robustness and reliability.
Requirements:
What we need to see:

BSc or above in Computer Science, Electrical Engineering, or a related field.

5+ years of recent hands-on coding experience with strong proficiency in C++ and Python. You must be comfortable navigating, optimizing, and contributing to a complex, large-scale codebase.

End to end system perspective - capable of bridging the gap between hardware behavior, micro-architecture, and software performance (latency/throughput).

Project ownership capability with proven ability to lead technical initiatives, prioritize features, and manage project lifecycles autonomously with minimal supervision.

Ways to stand out from the crowd:

Strong background in communication networks technology. Deep understanding of Ethernet, NVLink, and/or Infiniband technologies, data center infrastructure, and network protocols.

Familiarity with discrete event simulators (e.g., OMNeT++, SystemC, or proprietary architectural simulators).

 Advanced degree or experience focusing on computer architecture, algorithms, networking, or distributed systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836077
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Join our company, an innovative startup building a fully-managed LLM-inference platform, that enables data heavy enterprises to perform any AI task at any scale without limits.
We're looking for an Experienced Performance Researcher to join our founding team. Youll be responsible for building and optimizing scalable cloud infrastructure solutions tailored for AI workloads. This role offers a unique opportunity to directly shape our infrastructure strategy, improve system reliability and performance, and contribute to establishing our company as a leader in adaptive AI compute management.
Join us to tackle the magic that make AI tick under the hood and build the backbone powering the AI revolution.
What Youll Do
- Design and build high-performance distributed inference pipelines for LLMs, focused on large-batch, non-real-time scenarios.
- Optimize GPU memory usage, kernel execution, and communication across nodes (NCCL, MPI, etc.).
- Own CUDA kernels, compiler-level tricks, and multi-GPU scheduling logic.
- Lead profiling and performance tuning for throughput, and cost- down to the kernel level.
- Collaborate with infra, product, and research teams to define SLAs, resource allocation logic, and runtime behaviors.
- Help build the core infrastructure that will run LLM workloads across hybrid GPU environments (cloud/on-prem/self-hosted).
Requirements:
- Deep experience with CUDA programming, GPU architecture, and low-level performance engineering.
- Fluency with Python and C++, and a mastery of profiling tools like Nsight, nvprof, perf, etc.
- Experience building systems for large-scale distributed training or inference (PyTorch, DeepSpeed, Ray, Horovod, etc.).
- Hands-on familiarity with cluster and container orchestration tools (Kubernetes, Slurm, Docker).
- Self-motivated and able to operate independently in a fast-moving startup environment.
- Strong analytical skills and a passion for elegant performance wins.
- A collaborative team player with strong interpersonal skills, a positive and easygoing attitude, and the potential to grow into a leadership role.
- Prior experience building inference runtimes or scheduling frameworks.
- Experience with serverless GPU models, model parallelism, tensor slicing, and batching tricks.
- Contributions to open-source HPC or ML infra projects.
- Understanding of AI/ML privacy and compliance concerns in enterprise environments.
- Track record of working on distributed systems at bleeding-edge research labs or infrastructure teams.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8807342
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are at the forefront of the AI revolution, delivering brand new accelerated compute platforms for global impact. Our Network Architecture group is seeking a talented and motivated Sr. Software Engineer to build the agentic workflows that our architects use in their daily work. The software at the center of this role is our hardware network simulation environment - you will design multi-step agent workflows over it, engineer the context that grounds them in our own specifications and source code, and optimize their runtime performance. If you are passionate about building the practical infrastructure that brings intelligent agents to life, we want to hear from you.


What you'll be doing:
Build agentic workflows - loops, graphs, and multi-step pipelines - that carry real hardware network simulation and analysis work end to end.
Engineer the context these workflows run on, turning our simulation models, specifications, design documents, and source code into context that makes agents accurate in our domain.
Work closely with network architects to understand their workflows and translate them into agent workflows they use daily.
Optimize the runtime performance of our simulation tooling on these platforms, including execution time, compute cost, and end-to-end latency.
Define evaluation and regression testing for agent workflows, so that changes to a prompt, a graph, or a context source are measurable.
Build observability across agent runs: what the agent did, where it failed, and why.
Champion guidelines for secure and reliable agent workflows, including data handling, access control, and interaction boundaries.
Serve as a key technical resource for solving sophisticated integration issues between agents and internal tooling.
Requirements:
What we need to see:
B.Sc. or above in Computer Science, Computer Engineering, or a related field, or equivalent experience.
5+ years of hands-on experience in software engineering, with demonstrated ownership of production systems from design through deployment.
Expert-level programming skills in C++, with strong Python skills alongside it.
Strong understanding of the full stack, including hardware: memory, I/O, networking, accelerators, and where real performance bottlenecks occur.
Current, practical knowledge of how to build systems around AI models: agent loops, tool interfaces, context retrieval and management, and common failure modes.
Understanding of inference serving, including request lifecycle, batching, caching, and the tradeoffs between throughput, latency, and cost.


Ways to stand out from the crowd:
Experience writing hardware simulation software - network, system, or architectural simulators, models, or testbenches.
Networking experience - protocols, fabrics, switching, or RDMA - and experience working alongside silicon, systems, or architecture teams.
Hands-on experience with inference serving engines such as vLLM, TensorRT-LLM, or Triton Inference Server, including low-level internals such as KV cache, batching and scheduling, and quantization, and related performance work such as profiling and GPU programming.
Hands-on experience building or fine-tuning LLMs or other generative models.
Agent workflows, tooling, or context pipelines adopted by other engineering teams.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837912
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: More than one
Job Type: Full Time
We are looking for a Technical Support Engineer dedicated to supporting Slurm for our customers. You will join a dedicated team of Slurm subject-matter guides, owning sophisticated support cases and helping customers run reliable, efficient, and highly scalable clusters. This role requires extensive production experience with Slurm and the ability to diagnose issues across the scheduler and the surrounding Linux, networking, storage, authentication, database, and GPU infrastructure.

What you'll be doing:

Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters.

Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.

Solve Slurm configuration and policy features, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.

Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required.

Isolate problems across Slurm and its surrounding dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.

Advise customers on Slurm configuration, upgrades, operational practices, managing system resources, and safe recovery from production incidents.

collaborate with engineering teams by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.

Develop guides, knowledge-base articles, diagnostic tools, and internal training that strengthen Slurm expertise across the support organization.
Requirements:
What we need to see:

BS.c degree in Computer Science, Engineering, or a related field, or equivalent experience.

5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments including business-critical outage incidents.

Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.

Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution.

In-depth Linux system-administration and solve experience, including systemd, cgroups, authentication, networking, and database-backed services.

Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources.

Strong analytical and research skills, showing proficiency in distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems.

Excellent written and verbal communication skills, including the ability to turn detailed technical findings into clear explanations and actionable recommendations.


Ways to stand out from the crowd:

Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs.

Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues.

Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.

Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity.

Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836131
סגור
שירות זה פתוח ללקוחות VIP בלבד