דרושים » תוכנה » ML Hardware Achitect

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 2 שעות
חברה חסויה
Location: Haifa and Tel Aviv-Yafo
Job Type: Full Time
Required ML Hardware Achitect
About the job
In this role, youll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to drive cutting-edge TPU (Tensor Processing Unit) technology that powers our most demanding AI/ML applications. Youll be part of a team that pushes boundaries, developing custom silicon solutions that power the future of our TPU. You'll contribute to the innovation behind products loved by millions worldwide, and leverage your design and verification expertise to verify complex digital designs, with a specific focus on TPU architecture and its integration within AI/ML-driven systems.
In this role, you will help shape the future of Clouds next-generation AI infrastructure, architecting high-performance Machine Learning silicon designed to power hyperscale AI inference. You will have an opportunity to drive accelerator technology that powers Generative AI models, large language models (LLMs), and emerging agentic workloads where throughput, latency, memory bandwidth, and energy efficiency are mission-critical.
You will be part of a silicon architecture team pushing the boundaries of custom computing. Leveraging your deep expertise in hardware-software co-design, machine learning algorithms, and computer architecture, you will define and optimize custom compute engines and memory hierarchies that accelerate the world's most advanced AI models across Cloud datacenters.
Responsibilities
Lead the architectural definition, modeling, and specification of next-generation, high-performance ML compute IP and acceleration blocks for Cloud AI silicon.
Own the ML IP architecture specification throughout the entire product lifecycle: concept exploration, cycle-accurate modeling, implementation, silicon bring-up, and production.
Partner closely with leading AI research and algorithm teams (e.g., Google DeepMind, Gemini research teams) and software compiler teams (XLA, PyTorch) to explore architectural trade-offs and define hardware requirements for emerging model architectures.
Drive comprehensive architecture studies, evaluating compute dataflows, numerical formats, sparsity, and specialized acceleration mechanisms such as key-value (KV) cache optimization.
Drive performance, latency, power efficiency, and silicon area projections across model topologies and workload configurations.
Requirements:
Minimum qualifications:
Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, a related field, or equivalent practical experience.
15 years of experience in computer architecture, ML accelerator design, or high-performance processor architecture.
Experience leading architectural definition and authoring architecture specifications for silicon or compute IP blocks.
Experience with performance modeling, workload profiling, and hardware-software co-design.
Preferred qualifications:
Master's degree or PhD in Electrical Engineering, Computer Engineering, or Computer Science with an emphasis on computer architecture or ML hardware systems.
5 years of experience leading the architectural definition and microarchitecture of AI/ML accelerators from concept through production.
Deep knowledge of modern deep learning workloads (Transformers, MoE, Diffusion, Generative AI inference) and their system bottlenecks (memory capacity, KV cache bandwidth, interconnect scaling).
Strong understanding of high-performance memory subsystems (custom SRAM architectures, high-bandwidth memory hierarchies, caching schemes).
Experience working with modern ML frameworks (PyTorch, JAX, TensorFlow) and ML compilers/runtimes (XLA, TVM, Triton).
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8843260
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Required Qualifications
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834186
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Machine Learning Engineering Manager, you will lead a team focused on the foundational ML & Data layers to power the ranking & recommendation systems in scope. You will drive the development of robust data & ML pipelines at scale, lead the implementation of the tools for ML scientists to test and productionize advanced ML RecSys solutions.

As a technical manager of Machine Learning Engineers and Data engineers, you should be passionate about technology, keep up to date with recent breakthroughs in the field, define and shape the teams ML and platforms roadmap, and not be afraid to get your hands dirty with code when needed.

You are expected to be the focal point for all technical aspects, make sure your team members deliver on their tasks, and work together with other stakeholders to define and shape the roadmap of our products. You will work independently and will also be responsible for making technical decisions within your team.

When it comes to management, your expertise in handling people will motivate and inspire them to reach outstanding success! You should have experience in developing people. You will mentor and coach your team while working closely with a Product Manager.



Key Job Responsibilities and Duties:

Lead and develop a high-performing team, fostering individual growth and collaboration.

Manage and mentor ML engineers and Data engineers, ensuring their professional development and effectiveness.

Develop scalable ML infrastructure and pipelines for efficient data processing and evaluations deployment.

Evaluate architecture solutions based on cost, business needs, and emerging technologies.

Collaborate closely with software engineers to ensure seamless deployment and model inference.

Monitor application health, set and track relevant metrics, and implement effective maintenance strategies.

Collaborate with stakeholders to translate business requirements into viable ML solutions.

Evaluate and integrate new ML technologies to enhance productivity and performance.
Requirements:
3+ years leading an ML engineering team of a minimum of 4 people in a fast-paced production environment.

Relevant work or academic experience (MSc + 5 years of working experience, or PhD + 3 years of working experience), involved in the application of Machine Learning to business problems.

Masters degree, PhD or equivalent experience in a quantitative field (e.g. Computer Science, Engineering Mathematics, Artificial Intelligence, Physics, etc.).

Strong knowledge in areas like e.g. Recommender Systems, Deep Learning, Information Retrieval, Causal Inference, scaling ML models, etc.

Experience designing and executing end-to-end solutions for deploying different ML models.

Experience with cloud frameworks like AWS sagemaker for training, evaluation and serving models using TensorFlow, PyTorch, or scikit-learn.

Experience with big data processing frameworks such, Pyspark, Apache Flink, Snowflake or similar frameworks.

Demonstrable experience with MySQL, Cassandra, DynamoDB or similar relational/NoSQL database systems.

Deep understanding of machine learning algorithms, statistical models, and data structures.

Experience collaborating cross functionally in the development of machine learning products (e.g. Developers, UX specialists, Product Managers, etc.).

Strong working knowledge of Python, Java, Kafka, Hadoop, SQL, and Spark or similar technologies. Working experience with version control systems.

Excellent English communication skills, both written and verbal.

Successfully driving technical, business and people related initiatives that improve productivity, performance and quality while communicating with stakeholders at all levels

Leading by example, gaining respect through actions, not your title. Developing your team and motivating them to achieve their goals. Providing feedback timely and managing your key team performance indicators.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8809571
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Required Qualifications
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications:
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834018
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Required Qualifications
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications:
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834083
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
Join our companys AI research group, a cross-functional team of ML engineers, researchers and security experts building the next generation of AI-powered security capabilities. Our mission is to leverage large language models to understand code, configuration, and human language at scale, and to turn this understanding into security AI capabilities which will drive our company AI future security solutions.
We foster a hands-on, research-driven culture where youll work with large-scale data, modern ML infrastructure, and a global product footprint that impacts over 100,000 organizations worldwide.
Key Responsibilities
Your Impact & Responsibilities
As a Senior ML Research Engineer, you will be responsible for the end-to-end lifecycle of large language models: from data definition and curation, through training and evaluation, to providing robust models that can be consumed by product and platform teams.
Own training and fine-tuning of LLMs / seq2seq models: Design and execute training pipelines for transformer-based models (encoder-decoder, decoder-only, retrievalaugmented, etc.), and fine-tune open-source LLMs on our company-specific data (security content, logs, incidents, customer interactions).
Apply advanced LLM training techniques such as instruction tuning, preference / contrastive learning, LoRA / PEFT, continual pre-training, and domain adaptation where appropriate.
Work deeply with data: define data strategies with product, research and domain experts; build and maintain data pipelines for collecting, cleaning, de-duplicating and labeling large-scale text, code and semi-structured data; and design synthetic data generation and augmentation pipelines.
Build robust evaluation and experimentation frameworks: define offline metrics for LLM quality (task-specific accuracy, calibration, hallucination rate, safety, latency and cost); implement automated evaluation suites (benchmarks, regression tests, redteaming scenarios); and track model performance over time.
Scale training and inference: use distributed training frameworks (e.g. DeepSpeed, FSDP, tensor/pipeline parallelism) to efficiently train models on multi-GPU / multi-node clusters, and optimize inference performance and cost with techniques such as quantization, distillation and caching.
Collaborate closely with security researchers and data engineers to turn domain knowledge and threat intelligence into high-value training and evaluation data, and to expose your models through well-defined interfaces to downstream product and platform teams.
Requirements:
What You Bring
5+ years of hands-on work in machine learning / deep learning, including 3+ years focused on NLP / language models.
Proven track record of training and fine-tuning transformer-based models (BERT-style, encoder-decoder, or LLMs), not just consuming hosted APIs.
Strong programming skills in Python and at least one major deep learning framework (PyTorch preferred; TensorFlow).
Solid understanding of transformer architectures, attention mechanisms, tokenization, positional encodings, and modern training techniques.
Experience building data pipelines and tools for large-scale text / log / code processing (e.g. Spark, Beam, Dask, or equivalent frameworks).
Practical experience with ML infrastructure, such as experiment tracking (Weights & Biases, MLflow or similar), job orchestration (Airflow, Argo, Kubeflow, SageMaker, etc.), and distributed training on multi-GPU systems.
Strong software engineering practices: version control, code review, testing, CI/CD, and documentation.
Ability to own research and engineering projects end-to-end: from idea, through prototype and controlled experiments, to models ready for integration by product and platform teams.
Good communication skills and the ability to work closely with non-ML stakeholders (security experts, product managers, engineers).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8840766
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 4 שעות
חברה חסויה
Location: Hod Hasharon and Haifa
Job Type: Full Time
We are looking for a Senior AI Modeling Architect to define and model the architecture of our next-generation AI processors. In this role, you will serve as a key technical authority - bringing not only strong modeling capabilities, but also the architectural vision to propose and drive end-to-end solutions to complex design challenges.
You will work at high levels of abstraction, partnering closely with HW and SW architects to co-invent optimal solutions, validate them through simulation, and influence design decisions based on experimental data.
Responsibilities:
Model CPU/AI processor functionality and performance using SystemC and pre-silicon simulation environments
Define and drive architectural solutions - not only identify problems, but come with concrete proposals and alternatives
Partner with lead HW and SW architects to co-design features and evaluate trade-offs across the stack
Analyze bottlenecks and performance on workloads reflecting future AI use cases Provide proof-of-concept implementations for new architectural features and design alternatives
Potentially lead feature definition in addition to the modeling role
Requirements:
B.Sc. or higher in Electrical Engineering, Computer Science, or related discipline
10+ years of experience in VLSI/processor architecture (exceptional candidates with less experience will be considered)
Strong hands-on experience with SystemC modeling
Solid background in AI workloads and AI hardware architecture
Experience in HW/SW co-design and architectural trade-off analysis
Ability to operate at high levels of abstraction and drive architectural decisions from simulation data
Skills:
Strong independent technical judgment - able to come with a solution, not just a model
Excellent interpersonal and collaboration skills
Creative, self-driven, and a strong team player
Advantages:
Experience in CPU/DSP/GPU processor core architecture definition
Familiarity with AI accelerator design and large-scale inference solutions
Exposure to ISA definition or new instruction design.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8842860
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Backend Engineer to help build and scale the Machine Learning Platform that powers how we use AI across the business. You'll be part of the ML Platform team, designing the infrastructure that lets our data scientists move faster, ship smarter, and operate with confidence in production.
We believe three things matter for every role: drive to push through challenges, efficiency that keeps standards high while moving fast, and adaptability that lets you pivot with data and AI insights. These aren't buzzwords, they're how we actually work. Our AI-first approach isn't just a tagline either. We're building the future of insurance with AI at the center, and we need people who are genuinely excited to learn and grow alongside these tools.
In this role you'll:
Design and build the foundational ML platform and AI agents to accelerate data science model delivery across all business units
Architect cloud-native microservices running on Kubernetes, using infrastructure-as-code to automate model deployment and management
Own the end-to-end ML lifecycle, covering training, testing, deployment, and real-time monitoring
Evaluate and choose the right tools and technologies based on workload demands and performance requirements
Collaborate with engineering, data science, and product teams to keep ML projects aligned with business goals
Identify and fix reliability, scalability, and performance gaps before they become problems.
Requirements:
What you'll need
3+ years of software engineering experience, with a strong record of delivering high-scale, production-grade systems
Strong proficiency in Python
Hands-on experience with relational and NoSQL databases, and at least one major cloud platform (AWS, Azure, or GCP)
Experience with training, testing, deploying, and monitoring real-time or near real-time ML models in production
Fluent with AI-powered development tools like Cursor and Claude Code, and genuinely curious about what's next in GenAI, LLMs, and AI agents
Familiarity with AI concepts like RAG, embeddings, mixture-of-experts, prompt crafting, and LLM context engineering - an advantage
Sharp problem-solving instincts and the ability to move fast without cutting corners
Bachelor's or Master's degree in Computer Science, Engineering, Statistics, or a related field
Ready to work in an office environment most days of the week.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837916
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
27/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are building the first Neuron performance engineering team in Tel Aviv. As a Machine Learning Performance Engineer, you'll help shape the direction of this team from the ground up - profiling and optimizing workloads across the full ML software stack, writing high-performance kernels, and improving the Neuron SDK that external developers depend on. You'll work at the boundary between software and hardware, collaborating directly with compiler, runtime, and chip design engineers to close performance gaps customers care about.

The team is new and small, which means broad scope, direct ownership, and real influence over the technical direction we take. If you enjoy digging into performance bottlenecks and turning analysis into measurable wins, this role is for you.

Key job responsibilities
Design and implement high-performance compute kernels for ML operations, leveraging the Neuron architecture and programming models.
Profile ML workloads end-to-end to identify bottlenecks - memory, compute, or communication - and drive optimizations through to a measured improvement.
Enhance the programming model and tooling that kernel and model developers rely on, improving usability and debugging workflows.
Identify and drive optimization opportunities across the Neuron software stack (compiler, runtime, frameworks).
Document software designs, operational runbooks, and performance findings so the broader team can build on your work.

A day in the life
You might start your morning reviewing profiling data from a customer's large diffusion model training job, tracing a utilization gap back to a specific kernel. After a design discussion with compiler engineers about a new operator fusion strategy, you spend the afternoon writing and benchmarking a kernel prototype. Later, you review a teammate's pull request for a runtime optimization and share your findings in a short write-up for the broader Neuron organization. Your work directly translates into faster model execution and lower cost for AWS customers running ML workloads at scale.
Requirements:
Basic Qualifications
- 3+ years of non-internship professional software development experience.
- Knowledge of Python and/or C++ programming.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Experience with PyTorch, TensorFlow, and/or JAX.

Preferred Qualifications
- Master's degree in Computer Science, Engineering, Mathematics, or a related field.
- Experience optimizing performance for LLM, Vision, or other deep-learning models.
- Experience with kernel writing or parallel programming (CUDA, Triton, CUTLASS, Pallas, Mojo, SIMD, MPI).
- Experience with compiler optimization or hardware-software co-design.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8834492
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
10/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs.



Some directions we are currently working on, and which you can be a part of:

Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5).
Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines.
Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency.
Requirements:
A profound understanding of theoretical foundations of machine learning and transformer architecture.
Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools
Understanding of GPU memory hierarchy and compute/memory tradeoffs
Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation
Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)
Strong software engineering skills (we mostly use Python)
Deep experience with modern deep learning frameworks
Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing
Strong communication and leadership abilities
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8817627
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior Machine Learning Engineer, you will work closely with top notch engineers and data scientists to design, develop, evaluate and deploy Gen AI-powered solutions for scalable, customer-facing applications. Your work will focus on building and applying state-of-the-art agentic capabilities to drive business impact and improve efficiency.



Key Job Responsibilities and Duties:

Design, develop, and deploy high-quality, performant, and efficient Generative AI-powered solutions and agentic systems into production environments.

Evaluate and define optimal architectural solutions by considering emerging technologies, business needs, and technical requirements for latency, throughput, and scale.

Own services end-to-end, including implementing robust monitoring and maintenance strategies to ensure application and ML health, quality, and performance.

Write and maintain clean, scalable, and well-tested production code, ensuring reproducibility and seamless integration via CI/CD pipelines.

Pioneer and promote best practices and the adoption of cutting-edge technology in GenAI application development.

Collaborate effectively with Product Managers, Data Scientists, and Analysts to understand business requirements and translate them into technical ML and agentic solutions.

Provide technical guidance and mentorship to other engineers, contributing to the team's overall technical development.
Requirements:
Bachelors or masters degree in Computer Science, Engineering, Statistics, or a related field.

Minimum of 6 years of experience as a Machine Learning Engineer or a similar role, with a consistent record of successfully delivering ML solutions to production.

Experience of working on products that impact a large customer base.

Demonstrable experience and capabilities with Generative AI applications, including Large Language Models (LLMs), Agentic Systems, and MCP in production environments. Experience deploying large-scale language models (e.g., GPT, BERT, or similar architectures) - an advantage.

Deep understanding of core machine learning algorithms, statistical models, evaluation methods, and data structures.

Experience in designing, building, and deploying models using cloud frameworks (e.g., AWS Sagemaker) and standard ML libraries (e.g., TensorFlow, PyTorch, or scikit-learn).

Strong programming proficiency in languages such as Python and Java.

Strong coding practices, including writing and reviewing production-quality, maintainable, and well-tested code, with the ability to effectively leverage modern AI coding assistants while maintaining high standards for correctness, readability, and system design.

Experience with big data processing frameworks (e.g., Pyspark, Apache Flink, Snowflake) and demonstrable experience with relational/NoSQL database systems (e.g., MySQL, Cassandra, DynamoDB).

Excellent English communication and presentation skills, both written and verbal.

Proficiency in data manipulation, analysis, and visualization using tools like NumPy, pandas, and matplotlib - an advantage.

Experience with experimental design, A/B testing, and evaluation metrics for ML models - an advantage.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8809567
סגור
שירות זה פתוח ללקוחות VIP בלבד