דרושים » תוכנה » ML Software Engineer, Data Plane

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774268
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
11/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications:
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications:
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776998
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We seek a versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference - from novel decoding strategies to quantization schemes - into production across our hardware lineup, from large data center servers to powerful edge devices. We work on the most advanced architectures in the field, with a focus on NVIDIA's own.

What you'll be doing:

Implement and optimize inference algorithms for LLM and omnimodal architectures, including hybrid Mamba-Transformer and mixture-of-experts models.

Profile inference pipelines using NVIDIA's profiling and simulation tools. Correlate simulation predictions against real hardware across data center and edge devices.

Write and tune GPU kernels (CUDA, Triton) for operators like fused MoE layers, SSM state updates, and quantized GEMMs.

Solve distributed inference problems: expert parallelism, communication-compute overlap, collective tuning, multi-node deployment.

Build production-grade software inside major open-source libraries - vLLM, SGLang, Dynamo, FlashInfer.

Own optimization features end-to-end, from scoping through delivery, collaborating with research, product, and engineering teams worldwide.
Requirements:
What we need to see:

B.Sc., M.Sc., or equivalent experience in Computer Science or Computer Engineering.

5+ years of hands-on software engineering experience in performance-critical systems.

Solid understanding of deep learning architectures (Transformers, SSMs, MoE, ).

Experience with systems where hardware constraints matter: GPU programming, memory hierarchy, networking, or distributed computing.

Strong software engineering fundamentals: clean design, extensibility, testability. Good judgment about when complexity is warranted.

Effective communicator who works well across teams and time zones.

Experience optimizing deep learning workloads on our GPUs using roofline models, Nsight/PyTorch profilers and end-to-end traces.


Ways to stand out from the crowd:

Contributions to open-source inference runtimes and libraries - vLLM, SGLang, FlashInfer, Dynamo or similar.

Hands-on work with LLM quantization (FP8, NVFP4, MXFP8, mixed-precision) and practical understanding of numerical precision tradeoffs.

Track record with distributed inference at scale: tensor parallelism, pipeline parallelism, expert parallelism, disaggregation, multi-node orchestration.

Deep knowledge of the latest LLM architectural trends: multi-token predictors, sparse hybrid models, attention and state-space mechanisms.

Experience with performance modeling and simulation-to-silicon correlation.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769559
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8781454
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs.



Some directions we are currently working on, and which you can be a part of:

Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5).
Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines.
Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency.
Requirements:
We expect you to have:

A profound understanding of theoretical foundations of machine learning and transformer architecture.
Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools
Understanding of GPU memory hierarchy and compute/memory tradeoffs
Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation
Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)
Strong software engineering skills (we mostly use Python)
Deep experience with modern deep learning frameworks
Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing
Strong communication and leadership abilities
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761292
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL FLOW team is looking for a Software Development Engineer to design and build automation, tooling, and monitoring systems for our next-generation ML accelerator servers. We build production software to validate, initialize, monitor, and qualify these servers - from first silicon through fleet-scale deployment. Our work spans hardware diagnostics, manufacturing test automation, CI/CD pipelines, operational dashboards, and data-driven fleet health monitoring.

Key job responsibilities
Design and develop software infrastructure - automation frameworks, deployment systems, and test orchestration platforms that run at scale across manufacturing and production environments.
Work cross-functionally with Hardware, Manufacturing, and EC2 teams to automate coordinated software delivery and qualification workflows.
Debug and root-cause hardware/software interaction failures using systematic data analysis and automation-assisted triage.
Build and own CI/CD pipelines end-to-end: from code commit through build, test, deploy, and production validation - driving fast, reliable software delivery for hardware teams.
Create data pipelines and analytics systems (ETL, aggregation, real-time reporting) that transform raw hardware test results into actionable engineering insights.
Develop monitoring dashboards, alerting systems, and data visualization tools for fleet health, yield tracking, and performance benchmarking.
Own features end-to-end: from design through implementation, testing, deployment, and operational excellence.
Requirements:
Basic Qualifications
- 3+ years of software development engineer or related occupational experience.
- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering or a related discipline or equivalent.
- Experience using Linux, demonstrating proficiency with associated tools or languages.
- Can work proactively and independently, meet deadlines, and deliver on projects and tasks.
- Knowledge of software engineering best practices across the development life cycle, including agile methodologies, coding standards, code reviews, source management, build processes, testing, and operations.

Preferred Qualifications
- Experience building monitoring dashboards and data visualization (Grafana, CloudWatch, QuickSight, or similar).
- Experience with data pipelines, ETL, or analytics (S3, Athena, Spark, or similar).
- Experience with systems programming languages (C, C++, Rust).
- Familiarity with computer architecture concepts (PCIe, memory hierarchy, power management).
- Advantage: experience with hardware bring-up, ASIC/FPGA validation, or manufacturing test development.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774254
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
04/08/2026
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are seeking an AI Networking Architect to join the Networking Research Group. This role will help bridge the gap between emerging tasks supported by advanced technologies and the data center infrastructure that powers them. In this role, you will work at the intersection of AI applications, distributed systems, networking hardware, and software architecture.

You will join a focused team of multidisciplinary engineers driving AI workload optimization through deep application understanding, network analysis, and end-to-end systems thinking. Your insights will directly shape our products across the full stack - from applications and software libraries to hardware architecture and physical design.

What Youll Be Doing:

Model the performance of complex AI workloads to identify bottlenecks and recommend system-level optimizations.

Analyze brand-new AI models, distributed training techniques, and inference workloads to understand their infrastructure requirements.

Build Platforms, simulations and HW platforms, execute AI workloads and build analytical tools to evaluate trade-offs across compute, memory, storage, and network behavior.

Translate research insights and workload behavior into actionable software, hardware, and networking architecture requirements.

Partner with architecture, software, and product teams to influence our future networking and AI infrastructure roadmaps.

Drive architectural innovation by applying deep workload analysis to real-world advanced machine learning frameworks.
Requirements:
What we need to see:

B.Sc. Or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.

3+ years of relevant industry or research experience.

Strong machine learning or data science background, with hands-on experience in LLMs, generative AI, or deep learning systems.

Strong systems-level thinking, capable of estimating end-to-end requirements across the AI stack.

Shown ability to translate research findings and product requirements into clear software and hardware specifications.

Excellent research skills, including the ability to digest academic papers, self-learn new domains, and independently test hypotheses.

Advanced programming skills for performance modeling, data analysis, and prototyping.

Excellent communication skills, demonstrating proficiency in presenting complex technical findings clearly and confidently.


Ways to Stand Out from the crowd:

Experience with distributed training, distributed inference, or large-scale AI serving systems.

Experience in Agentic programming, and AI tools

Familiarity with GPU clusters, collective communication, storage systems, or AI networking bottlenecks.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8767980
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a Senior MLOps Engineer to be a core driver in how our product empowers security teams. You will be expected to deeply understand customer needs and translate them directly into product features that deliver real value. You'll own key parts of our frontend stack, drive key architectural decisions, and turn complex security data into clear, actionable business insights.

As we scale our AIDR product and expand deeper into model-driven security intelligence, we are looking for a Senior MLOps Engineer to own the infrastructure, tooling, and operational foundations that power our NLP and LLM training, evaluation, and deployment workflows.

You will architect and operate the systems that enable us to train, fine-tune, deploy, and monitor models at scale making ML reliable, fast, cost-efficient, and production-ready.

This is a high-visibility, high-impact role where you will partner closely with DevOps, Backend, Data, and Product to establish world-class ML infrastructure from the ground up.

What Youll Do

Build & Scale ML Pipelines
Design, build, and maintain pipelines for training, fine-tuning, evaluating, and deploying NLP and LLM models across GPU and CPU environments.
Establish LLM-Focused CI/CD
Implement automated CI/CD workflows for ML models, including benchmarking, testing, performance gating, and production deployment.
Optimize Runtime & Inference
Select and optimize serving frameworks for low-latency, high-throughput inference, ensuring reliability and scalability.
Own ML Infrastructure
Manage training environments, experiment tracking, model registries, artifact versioning, and distributed training systems.
Operational Excellence
Monitor and optimize production models for performance, cost efficiency, availability, and observability.
Requirements:
5+ years in software engineering, MLOps, or ML engineering with hands-on experience deploying ML models to production.
Strong Python fundamentals and deep understanding of transformer architectures, tokenization, and NLP frameworks (PyTorch, HuggingFace).
Proven experience deploying and scaling LLMs for real-time inference-ideally on platforms like SageMaker, Vertex AI, or similar.
Expertise in GPU optimization, distributed training, and CPU-based inference optimization.
Strong cloud and Kubernetes background (EKS/GKE/AKS, Helm, Terraform, CI/CD for ML).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8762083
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are seeking a highly skilled and versatile Performance Research and Analysis Manager to join our Performance Group. This role will drive end-to-end performance strategy and execution for next-generation our data centers and solutions based on GPU systems, NIC, Switch, DPU and Networking technologies. The ideal candidate will oversee, evaluating, and optimizing end-to-end AI GPU cluster-level performance for scaling out large scale distributed training and inference jobs communication. The role will focus heavily on RDMA, Networking Protocols, Collective Communication, Congestion Control, and Load Balancing algorithms. Secondarily, you will lead our DPUs and Storage technologies for N-S use cases to support AI Inference jobs. Third, you will drive our Performance Dashboards and Observability for cluster-level performance analysis from a stream line telemetry across NICs, Switches, GPUs, and NVlink.

What you'll be doing:

Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation our AI GPU clusters, focusing on large-scale distributed training and inference workloads.

Deeply evaluate and optimize our Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.

Work on performance research and analysis of NVIDIA DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.

Drive the strategy for performance observability and dashboards across next-generation NVIDIA data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.

Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.
Requirements:
What we need to see:

B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.

8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.

3+ years of experience as an engineering team manager leading technical performance or R&D teams.

Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).

Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.

Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.


Ways to stand out from the crowd:

Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.

Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.

Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.

Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8773270
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Help build an Always-On, low-overhead GPU profiling service that runs in production, scales across cluster environments, and delivers actionable insights for ML workloads. You will be hands-on delivering our profiling solutions across system software, drivers, and CUDA to make profiling continuously available and reliable.

What youll be doing:

Develop low-overhead, high-reliability implementations in C/C++, with bounded CPU/memory budgets.

Lead end-to-end feature delivery spanning user-mode components, driver/platform layers, and performance counter/trace providers.

Establish profiling models that integrate with existing ML/AI workflows (e.g., PyTorch/XLA) to turn low-level signals into actionable insights.
Requirements:
What we need to see:

BS or MS degree or equivalent experience in Computer Engineering, Computer Science, or related degree.

5+ years of system-level C/C++ development, including concurrency, memory management, and performance engineering.

Familiarity with system software design, operating systems fundamentals, computer architectures, performance analysis, and delivering production-quality software.

Strong interpersonal, verbal, and written communication; able to influence across organizations and build trust with external collaborators.

Ways to stand out from the crowd:

Extensive experience with profiling/tracing stacks for CPU/GPU (e.g., CUPTI, Nsight, performance counters, event correlation) and debugging highly concurrent systems.

Deep hands-on knowledge of CUDA and GPU architecture, including runtime/driver APIs, CUDA streams/graphs, and kernel behavior.

Track record building continuous, always-on, or multi-client profiling systems designed for predictable overhead at scale.

Hands-on experience tuning ML training/inference loops based on deep profiling analysis, with familiarity in ML ecosystems (e.g., PyTorch, JAX) and correlating application events with GPU metrics to translate data into actionable performance insights (e.g., bottleneck triage, compute vs. memory bound).

Experience with user-mode driver development and integration within platform security and permissions models.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769934
סגור
שירות זה פתוח ללקוחות VIP בלבד