דרושים » הנדסה » Senior ML Software Engineer, Data Plane

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We are looking for an IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.

Preferred Qualifications
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8711182
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/06/2026
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are seeking an AI Networking Architect to join the Networking Research Group. This role will help bridge the gap between emerging tasks supported by advanced technologies and the data center infrastructure that powers them. In this role, you will work at the intersection of AI applications, distributed systems, networking hardware, and software architecture.

You will join a focused team of multidisciplinary engineers driving AI workload optimization through deep application understanding, network analysis, and end-to-end systems thinking. Your insights will directly shape our products across the full stack - from applications and software libraries to hardware architecture and physical design.


What Youll Be Doing:

Model the performance of complex AI workloads to identify bottlenecks and recommend system-level optimizations.

Analyze brand-new AI models, distributed training techniques, and inference workloads to understand their infrastructure requirements.

Build Platforms, simulations and HW platforms, execute AI workloads and build analytical tools to evaluate trade-offs across compute, memory, storage, and network behavior.

Translate research insights and workload behavior into actionable software, hardware, and networking architecture requirements.

Partner with architecture, software, and product teams to influence future NVIDIA networking and AI infrastructure roadmaps.

Drive architectural innovation by applying deep workload analysis to real-world advanced machine learning frameworks.
Requirements:
What we need to see:

B.Sc. Or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.

3+ years of relevant industry or research experience.

Strong machine learning or data science background, with hands-on experience in LLMs, generative AI, or deep learning systems.

Strong systems-level thinking, capable of estimating end-to-end requirements across the AI stack.

Shown ability to translate research findings and product requirements into clear software and hardware specifications.

Excellent research skills, including the ability to digest academic papers, self-learn new domains, and independently test hypotheses.

Advanced programming skills for performance modeling, data analysis, and prototyping.

Excellent communication skills, demonstrating proficiency in presenting complex technical findings clearly and confidently.


Ways to Stand Out from the crowd:

Experience with distributed training, distributed inference, or large-scale AI serving systems.

Experience in Agentic programming, and AI tools.

Familiarity with GPU clusters, collective communication, storage systems, or AI networking bottlenecks.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8702891
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
18/06/2026
Location: More than one
Job Type: Full Time
We're looking for a Senior Data Scientist to join the AI cybersecurity team in the Security and Networking Architecture group. As a Senior Data Scientist youll have the opportunity to take an active part in the research and development of our world-class networking and data center security products. This role involves creative problem solving alongside engineering teams, and is key for the continued success of AI networking security.

What youll be doing:

Developing agentic AI systems for security, combining generative models, RAG, and tool-augmented reasoning to automate threat analysis and response workflows.

Optimizing and fine-tuning models for performance, scalability, and resource utilization, considering factors such as latency, efficiency, and cost.

Developing, implementing and improving models and algorithms across media types, whether time series, images, text, audio or video.

Leveraging data pipelines to efficiently process and transform large volumes of data for training and inference purposes.

Applying alignment techniques and parameter efficient fine-tuning to improve model performance.

Measuring and benchmarking model and application performance to drive improvements.

Driving the gathering, building, and annotation of domain specific datasets for benchmarking and training.

Collaborating closely with software and hardware engineers on new features and improvements. Participate in developing and reviewing code, design documents, use case reviews, and test plan reviews.
Requirements:
What we need to see:

MS/PhD with expertise in Computer Science, Computer Engineering, Electrical Engineering or related field with a focus on Deep Learning or Machine Learning.

5+ years of experience in deep learning and machine learning in a production environment.

Excellent Python programming skills, strong software design fundamentals, and experience leveraging coding agents in development workflows.

Hands-on experience with deep learning development frameworks and libraries (e.g. TensorFlow, PyTorch).

Experience with large scale production systems and pipelines, with a track record of developing production-grade models

Experience with agentic AI systems, agent frameworks, and evaluation of agent performance and reliability.

Strong algorithm development experience, with knowledge of inference optimization techniques such as model distillation, quantization, pruning.

Background with algorithms including zero/few-shot learning, self-supervised and unsupervised learning and generative AI models for synthetic data creation.

Experience with fine-tune / training LLM models

You are proactive, take full ownership of your deliverables, have a can-do approach, and are excited to learn, explore and apply your skills and creativity to some of the most challenging and rewarding problems in the field.


What will make you stand out from the crowd:

Strong software development experience.

Familiarity with GPU based technologies like CUDA, CuDNN and TensorRT.

Experience with tools for data processing and storage.

Security and networking background, with knowledge of security protocols, network architectures, firewalls, intrusion detection systems, and other relevant security and networking concepts.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8701274
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Join our companys AI research group, a cross-functional team of ML engineers, researchers and security experts building the next generation of AI-powered security capabilities. Our mission is to leverage large language models to understand code, configuration, and human language at scale, and to turn this understanding into security AI capabilities which will drive our company AI future security solutions.
We foster a hands-on, research-driven culture where youll work with large-scale data, modern ML infrastructure, and a global product footprint that impacts over 100,000 organizations worldwide.
Key Responsibilities
Your Impact & Responsibilities
As a Senior ML Research Engineer, you will be responsible for the end-to-end lifecycle of large language models: from data definition and curation, through training and evaluation, to providing robust models that can be consumed by product and platform teams.
Own training and fine-tuning of LLMs / seq2seq models: Design and execute training pipelines for transformer-based models (encoder-decoder, decoder-only, retrievalaugmented, etc.), and fine-tune open-source LLMs on our company-specific data (security content, logs, incidents, customer interactions).
Apply advanced LLM training techniques such as instruction tuning, preference / contrastive learning, LoRA / PEFT, continual pre-training, and domain adaptation where appropriate.
Work deeply with data: define data strategies with product, research and domain experts; build and maintain data pipelines for collecting, cleaning, de-duplicating and labeling large-scale text, code and semi-structured data; and design synthetic data generation and augmentation pipelines.
Build robust evaluation and experimentation frameworks: define offline metrics for LLM quality (task-specific accuracy, calibration, hallucination rate, safety, latency and cost); implement automated evaluation suites (benchmarks, regression tests, redteaming scenarios); and track model performance over time.
Scale training and inference: use distributed training frameworks (e.g. DeepSpeed, FSDP, tensor/pipeline parallelism) to efficiently train models on multi-GPU / multi-node clusters, and optimize inference performance and cost with techniques such as quantization, distillation and caching.
Collaborate closely with security researchers and data engineers to turn domain knowledge and threat intelligence into high-value training and evaluation data, and to expose your models through well-defined interfaces to downstream product and platform teams.
Requirements:
What You Bring
5+ years of hands-on work in machine learning / deep learning, including 3+ years focused on NLP / language models.
Proven track record of training and fine-tuning transformer-based models (BERT-style, encoder-decoder, or LLMs), not just consuming hosted APIs.
Strong programming skills in Python and at least one major deep learning framework (PyTorch preferred; TensorFlow).
Solid understanding of transformer architectures, attention mechanisms, tokenization, positional encodings, and modern training techniques.
Experience building data pipelines and tools for large-scale text / log / code processing (e.g. Spark, Beam, Dask, or equivalent frameworks).
Practical experience with ML infrastructure, such as experiment tracking (Weights & Biases, MLflow or similar), job orchestration (Airflow, Argo, Kubeflow, SageMaker, etc.), and distributed training on multi-GPU systems.
Strong software engineering practices: version control, code review, testing, CI/CD, and documentation.
Ability to own research and engineering projects end-to-end: from idea, through prototype and controlled experiments, to models ready for integration by product and platform teams.
Good communication skills and the ability to work closely with non-ML stakeholders (security experts, product managers, engineers).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8722813
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, we build AI-powered vision systems that enhance safety and decision-making for some of the worlds largest vessels.
Our platform processes live video streams from multiple onboard cameras to provide real-time situational awareness, detecting and tracking marine objects, even in low visibility and highly congested environments. These systems directly support navigational decisions and help prevent collisions, reduce human error, and improve operational efficiency.
Our systems are already deployed across thousands of vessels and have processed hundreds of millions of nautical miles of real-world data, operating in unpredictable and safety-critical conditions.
This role sits at the intersection of AI and high-performance systems engineering, focused on solving real-world problems under strict constraints. You will work on systems where performance and reliability are critical and where improvements have a direct, measurable impact on real-world safety.
This is a senior, systems-focused role with end-to-end ownership over performance and reliability of production computer vision pipelines. You will define optimization strategies, identify bottlenecks across the system, and drive improvements under real-world constraints.
What youll do
Build and optimize real-time computer vision pipelines running on edge systems processing live maritime video streams (e.g, NVIDIA Jetson, Triton Inference Server)
Take models from research and turn them into production-ready, reliable components deployed on vessels
Profile and improve end-to-end system performance across: multi-camera video ingestion; preprocessing; inference; postprocessing
Identify and resolve bottlenecks across CPU, GPU, memory, and pipeline coordination
Make and justify tradeoffs between latency, accuracy, stability, and resource utilization
Design and implement robust data and inference pipelines (video -> model -> actionable output for crew)
Develop benchmarking and evaluation workflows to measure performance end-to-end and support release gating
Build and improve observability tools, including logging, monitoring, and debugging workflows for production systems
Define and maintain clear interfaces between research code and production systems
Work closely with research and backend teams to integrate new models into production systems
Continuously improve system efficiency and reliability under hardware and runtime constraints.
Requirements:
5+ years of software engineering experience, with a strong focus on systems and performance
Hands-on experience working with computer vision or deep learning systems in production
Strong programming skills in Python and/or C++
Experience working with edge or embedded systems (e.g., NVIDIA Jetson platforms)
Strong understanding of system bottlenecks, including CPU, GPU, memory, and latency constraints
Strong intuition for profiling-driven optimization and performance tuning
Experience debugging complex systems and reasoning about behavior in real-world, noisy environments
Strong advantage
Experience working with edge or embedded systems
Experience working with custom high-performance data or inference pipelines
Familiarity with multi-sensor fusion (e.g., combining vision with radar or other signals)
Experience deploying and maintaining ML models in production environments
Experience with low-level optimization and/or C++ performance tuning
Proven experience optimizing model inference (e.g., TensorRT, ONNX Runtime, quantization, pruning, or similar techniques).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8737671
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
18/06/2026
Job Type: Full Time
We're looking for a Senior AI Infrastructure Engineer to join a group that specializes in Security and Networking, and specifically ML/AI, MLOps, and agentic AI development. As a Senior AI Infrastructure Engineer, youll build and maintain the infrastructure, tools and processes necessary to support the AI lifecycle in a production environment. You will collaborate closely with data scientists, software engineers, and security architects to ensure smooth development, deployment, evaluation, and optimization of AI pipelines, models, and agents. This role requires a balance of high-level engineering rigor and a collaborative spirit; youll be a technical anchor and a supportive peer for teams across the organization.



What youll be doing:

Architecting, developing and optimizing scalable infrastructure for deploying security and networking AI models and agents in production.

Managing ML/agentic workflows to ensure performance, high availability, resource efficiency, and cost-effectiveness.

Designing and implementing pipelines and frameworks for AI training, inference, and experimentation.

Partnering with data scientists and security architects to operationalize AI agents, including packaging and integration with existing systems. This includes contributing to and reviewing code, design documents, and test plans.

Partnering with DevOps teams to integrate pipelines and workflows into CI/CD processes, ensuring reliable deployments and rollbacks.

Building proactive monitoring systems to identify issues in quality and infrastructure before they impact production.

Implementing access controls, authentication mechanisms, and encryption standards to keep our AI models and data secure.

Documenting guidelines and leading knowledge-sharing sessions to elevate the teams collective development expertise.
Requirements:
What we need to see:

BSc/MSc in CS/CE or related field (or equivalent experience).

At least 8 years of experience in ML engineering with a track record of deploying LLMs and agents to production at scale (including distributed environments).

Proficiency in Python and/or C++, with a deep understanding of ML/AI frameworks.

Hands-on experience with microservices, container orchestration, and cloud platforms for large-scale training and inference workloads.

Knowledge of ML training and inference optimization techniques.

Understanding of build infrastructure and CI/CD tools and practices (e.g. GitLab, GitHub Actions, Jenkins)

Experience with teaching and mentoring.

You are a proactive owner who takes pride in your work but remains humble and approachable. You believe that "how" we build is just as important as "what" we build.

Excellent collaboration skills, with the ability to explain complex infra concepts to non-technical stakeholders clearly and kindly.



Ways to stand out from the crowd:

Experience deploying and optimizing generative models and multi-agent systems for performance.

Deep systems knowledge (Linux internals, network protocols, or high-performance computing).

A background in security research, including knowledge of firewalls, intrusion detection, or network architectures.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8701273
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and Perplexity kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver meaningful solutions at scale, then come join us! This truly is a role at the forefront of AI/ML-you'll be working on features for the largest clusters, with the largest customers, for the largest AI models.

Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use our internal CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Requirements:
Basic Qualifications
- 5+ years of non-internship professional software development experience.
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- 3+ years as a mentor, tech lead or leading engineering teams.
- 3+years experience in SW/HW Co-Design.

Preferred Qualifications
- Bachelor's degree in computer science or equivalent.
- Experience creating automated dashboards and visualization (such as Grafana).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8711144
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
You'll design and build features that shape how hundreds of millions of customers discover products. The problems here are genuinely novel - we're defining how LLMs integrate into real-time recommendation experiences, not applying established playbooks. You'll work across multiple technical teams, ship iteratively, and see your work in the hands of customers quickly. This team values experimentation, moves fast, and gives engineers real ownership over what they build.

We are looking for a Software Development Engineer with sound technical judgment and a bias for action who takes ownership of problems end to end, communicates clearly, and cares about operational excellence, not just launching features, but making sure they hold up at scale. Someone who naturally raises the bar for the team: mentoring junior developers, advocating for engineering best practices, and thinking beyond the immediate sprint.

Key job responsibilities
- Design, build, test, and operate features for a personalized recommendation system used by multiple teams and operating at our scale.
- Deliver end-to-end solutions with focus on maintainability, scalability, performance, and reliability.
- Collaborate with Product and Science to define experiences, run experiments, and iterate based on data.
- Define and implement measurement strategies including analytics events and experiment configurations to track engagement and retention.
- Navigate ambiguity and make sound technical decisions in a problem space where established patterns don't always apply.
Requirements:
Basic Qualifications
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field.
- 5+ years of non-internship professional software development experience.
- Experience programming with at least one modern language such as Java, C++, or C# including object-oriented design.
- Experience with full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations.
- Experience contributing to the architecture and design (architecture, design patterns, reliability and scaling) of new and current systems.

Preferred Qualifications
- Master's degree or equivalent.
- Experience including, building and maintaining data flows and pipelines
- Experience with A/B testing.
- Familiarity with AI/ML integration and generative AI applications.
- Experience with end-to-end SDLC ownership, including operations and on-call, monitoring/metrics, and incident response/RCA.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8710999
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and Perplexity kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use internal our CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Requirements:
Basic Qualifications
- 5+ years of non-internship professional software development experience.
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- 3+ years as a mentor, tech lead or leading engineering teams.
- 3+years experience in SW/HW Co-Design.

Preferred Qualifications
- Bachelor's degree in computer science or equivalent.
- Experience creating automated dashboards and visualization (such as Grafana).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8748498
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
23/06/2026
Location: More than one
Job Type: Full Time
We are seeking an AI Networking Exploration Architect for our Networking Insights Group to bridge the gap between cutting-edge, hyper-scale AI workloads and the datacenter infrastructure that enables them. You will join a small, focused team of multidisciplinary engineers driving AI workload optimization through deep application understanding and end-to-end systems thinking. Your insights will directly shape our products across the full stack-from applications and software libraries to hardware architecture and physical design.

What You'll Be Doing:

Model the performance of complex AI workloads to identify bottlenecks and recommend system-level optimizations.

Translate state-of-the-art research into actionable infrastructure, software, and hardware features in partnership with architecture teams.

Rapidly master new AI domains (LLMs, generative models, multimodal systems) and distill key findings for product teams.

Incorporate your deep knowledge of AI applications into our hardware and software roadmaps.

Conduct independent research by formulating hypotheses about workload behavior and validating them through rigorous analysis.

Drive architectural innovation and network optimization by applying your domain expertise to exploratory analysis of real-world Deep Learning (DL) workloads.
Requirements:
What we need to see:

M.Sc. or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.

+5 years of experience.

Strong ML/Data Science background with hands-on experience in LLMs or generative AI.

A systems-level mindset with the ability to estimate end-to-end requirements across the entire AI stack.

Proven ability to translate research and product requirements into clear software/hardware specifications.

Exceptional research skills: you can digest academic papers, self-learn new domains, and independently test hypotheses.

Advanced Python programming skills for performance modeling and data analysis.

Excellent communication skills, with the ability to present complex findings with clarity and conviction.

A pragmatic approach: you are detail-oriented but can prioritize effectively to focus on the most critical issues.


Ways to Stand Out from the Crowd:

Deep understanding of datacenter infrastructure, network topologies, and protocols.

Expertise in distributed training methods and their impact on infrastructure.

Knowledge of AI performance metrics and the impact of different deployment strategies.

Experience extrapolating academic research into tangible hardware architecture requirements.

A track record of leading complex, multidisciplinary research projects that result in production impact.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8706931
סגור
שירות זה פתוח ללקוחות VIP בלבד