דרושים » תוכנה » Senior Software Engineer, LLM Inference

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 22 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
We seek a versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference - from novel decoding strategies to quantization schemes - into production across our hardware lineup, from large data center servers to powerful edge devices. We work on the most advanced architectures in the field, with a focus on NVIDIA's own.

What you'll be doing:

Implement and optimize inference algorithms for LLM and omnimodal architectures, including hybrid Mamba-Transformer and mixture-of-experts models.

Profile inference pipelines using NVIDIA's profiling and simulation tools. Correlate simulation predictions against real hardware across data center and edge devices.

Write and tune GPU kernels (CUDA, Triton) for operators like fused MoE layers, SSM state updates, and quantized GEMMs.

Solve distributed inference problems: expert parallelism, communication-compute overlap, collective tuning, multi-node deployment.

Build production-grade software inside major open-source libraries - vLLM, SGLang, Dynamo, FlashInfer.

Own optimization features end-to-end, from scoping through delivery, collaborating with research, product, and engineering teams worldwide.
Requirements:
What we need to see:

B.Sc., M.Sc., or equivalent experience in Computer Science or Computer Engineering.

5+ years of hands-on software engineering experience in performance-critical systems.

Solid understanding of deep learning architectures (Transformers, SSMs, MoE, ).

Experience with systems where hardware constraints matter: GPU programming, memory hierarchy, networking, or distributed computing.

Strong software engineering fundamentals: clean design, extensibility, testability. Good judgment about when complexity is warranted.

Effective communicator who works well across teams and time zones.

Experience optimizing deep learning workloads on our GPUs using roofline models, Nsight/PyTorch profilers and end-to-end traces.


Ways to stand out from the crowd:

Contributions to open-source inference runtimes and libraries - vLLM, SGLang, FlashInfer, Dynamo or similar.

Hands-on work with LLM quantization (FP8, NVFP4, MXFP8, mixed-precision) and practical understanding of numerical precision tradeoffs.

Track record with distributed inference at scale: tensor parallelism, pipeline parallelism, expert parallelism, disaggregation, multi-node orchestration.

Deep knowledge of the latest LLM architectural trends: multi-token predictors, sparse hybrid models, attention and state-space mechanisms.

Experience with performance modeling and simulation-to-silicon correlation.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769559
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Join our companys AI research group, a cross-functional team of ML engineers, researchers and security experts building the next generation of AI-powered security capabilities. Our mission is to leverage large language models to understand code, configuration, and human language at scale, and to turn this understanding into security AI capabilities which will drive our company AI future security solutions.
We foster a hands-on, research-driven culture where youll work with large-scale data, modern ML infrastructure, and a global product footprint that impacts over 100,000 organizations worldwide.
Key Responsibilities
Your Impact & Responsibilities
As a Senior ML Research Engineer, you will be responsible for the end-to-end lifecycle of large language models: from data definition and curation, through training and evaluation, to providing robust models that can be consumed by product and platform teams.
Own training and fine-tuning of LLMs / seq2seq models: Design and execute training pipelines for transformer-based models (encoder-decoder, decoder-only, retrievalaugmented, etc.), and fine-tune open-source LLMs on our company-specific data (security content, logs, incidents, customer interactions).
Apply advanced LLM training techniques such as instruction tuning, preference / contrastive learning, LoRA / PEFT, continual pre-training, and domain adaptation where appropriate.
Work deeply with data: define data strategies with product, research and domain experts; build and maintain data pipelines for collecting, cleaning, de-duplicating and labeling large-scale text, code and semi-structured data; and design synthetic data generation and augmentation pipelines.
Build robust evaluation and experimentation frameworks: define offline metrics for LLM quality (task-specific accuracy, calibration, hallucination rate, safety, latency and cost); implement automated evaluation suites (benchmarks, regression tests, redteaming scenarios); and track model performance over time.
Scale training and inference: use distributed training frameworks (e.g. DeepSpeed, FSDP, tensor/pipeline parallelism) to efficiently train models on multi-GPU / multi-node clusters, and optimize inference performance and cost with techniques such as quantization, distillation and caching.
Collaborate closely with security researchers and data engineers to turn domain knowledge and threat intelligence into high-value training and evaluation data, and to expose your models through well-defined interfaces to downstream product and platform teams.
Requirements:
What You Bring
5+ years of hands-on work in machine learning / deep learning, including 3+ years focused on NLP / language models.
Proven track record of training and fine-tuning transformer-based models (BERT-style, encoder-decoder, or LLMs), not just consuming hosted APIs.
Strong programming skills in Python and at least one major deep learning framework (PyTorch preferred; TensorFlow).
Solid understanding of transformer architectures, attention mechanisms, tokenization, positional encodings, and modern training techniques.
Experience building data pipelines and tools for large-scale text / log / code processing (e.g. Spark, Beam, Dask, or equivalent frameworks).
Practical experience with ML infrastructure, such as experiment tracking (Weights & Biases, MLflow or similar), job orchestration (Airflow, Argo, Kubeflow, SageMaker, etc.), and distributed training on multi-GPU systems.
Strong software engineering practices: version control, code review, testing, CI/CD, and documentation.
Ability to own research and engineering projects end-to-end: from idea, through prototype and controlled experiments, to models ready for integration by product and platform teams.
Good communication skills and the ability to work closely with non-ML stakeholders (security experts, product managers, engineers).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8722813
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
7 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs.



Some directions we are currently working on, and which you can be a part of:

Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5).
Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines.
Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency.
Requirements:
We expect you to have:

A profound understanding of theoretical foundations of machine learning and transformer architecture.
Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools
Understanding of GPU memory hierarchy and compute/memory tradeoffs
Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation
Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)
Strong software engineering skills (we mostly use Python)
Deep experience with modern deep learning frameworks
Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing
Strong communication and leadership abilities
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761292
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are seeking an AI Networking Architect to join the Networking Research Group. This role will help bridge the gap between emerging tasks supported by advanced technologies and the data center infrastructure that powers them. In this role, you will work at the intersection of AI applications, distributed systems, networking hardware, and software architecture.

You will join a focused team of multidisciplinary engineers driving AI workload optimization through deep application understanding, network analysis, and end-to-end systems thinking. Your insights will directly shape our products across the full stack - from applications and software libraries to hardware architecture and physical design.

What Youll Be Doing:

Model the performance of complex AI workloads to identify bottlenecks and recommend system-level optimizations.

Analyze brand-new AI models, distributed training techniques, and inference workloads to understand their infrastructure requirements.

Build Platforms, simulations and HW platforms, execute AI workloads and build analytical tools to evaluate trade-offs across compute, memory, storage, and network behavior.

Translate research insights and workload behavior into actionable software, hardware, and networking architecture requirements.

Partner with architecture, software, and product teams to influence our future networking and AI infrastructure roadmaps.

Drive architectural innovation by applying deep workload analysis to real-world advanced machine learning frameworks.
Requirements:
What we need to see:

B.Sc. Or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.

3+ years of relevant industry or research experience.

Strong machine learning or data science background, with hands-on experience in LLMs, generative AI, or deep learning systems.

Strong systems-level thinking, capable of estimating end-to-end requirements across the AI stack.

Shown ability to translate research findings and product requirements into clear software and hardware specifications.

Excellent research skills, including the ability to digest academic papers, self-learn new domains, and independently test hypotheses.

Advanced programming skills for performance modeling, data analysis, and prototyping.

Excellent communication skills, demonstrating proficiency in presenting complex technical findings clearly and confidently.


Ways to Stand Out from the crowd:

Experience with distributed training, distributed inference, or large-scale AI serving systems.

Experience in Agentic programming, and AI tools

Familiarity with GPU clusters, collective communication, storage systems, or AI networking bottlenecks.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8767980
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
At our company, we build AI-powered vision systems that enhance safety and decision-making for some of the worlds largest vessels.
Our platform processes live video streams from multiple onboard cameras to provide real-time situational awareness, detecting and tracking marine objects, even in low visibility and highly congested environments. These systems directly support navigational decisions and help prevent collisions, reduce human error, and improve operational efficiency.
Our systems are already deployed across thousands of vessels and have processed hundreds of millions of nautical miles of real-world data, operating in unpredictable and safety-critical conditions.
This role sits at the intersection of AI and high-performance systems engineering, focused on solving real-world problems under strict constraints. You will work on systems where performance and reliability are critical and where improvements have a direct, measurable impact on real-world safety.
This is a senior, systems-focused role with end-to-end ownership over performance and reliability of production computer vision pipelines. You will define optimization strategies, identify bottlenecks across the system, and drive improvements under real-world constraints.
What youll do
Build and optimize real-time computer vision pipelines running on edge systems processing live maritime video streams (e.g, NVIDIA Jetson, Triton Inference Server)
Take models from research and turn them into production-ready, reliable components deployed on vessels
Profile and improve end-to-end system performance across: multi-camera video ingestion; preprocessing; inference; postprocessing
Identify and resolve bottlenecks across CPU, GPU, memory, and pipeline coordination
Make and justify tradeoffs between latency, accuracy, stability, and resource utilization
Design and implement robust data and inference pipelines (video -> model -> actionable output for crew)
Develop benchmarking and evaluation workflows to measure performance end-to-end and support release gating
Build and improve observability tools, including logging, monitoring, and debugging workflows for production systems
Define and maintain clear interfaces between research code and production systems
Work closely with research and backend teams to integrate new models into production systems
Continuously improve system efficiency and reliability under hardware and runtime constraints.
Requirements:
5+ years of software engineering experience, with a strong focus on systems and performance
Hands-on experience working with computer vision or deep learning systems in production
Strong programming skills in Python and/or C++
Experience working with edge or embedded systems (e.g., NVIDIA Jetson platforms)
Strong understanding of system bottlenecks, including CPU, GPU, memory, and latency constraints
Strong intuition for profiling-driven optimization and performance tuning
Experience debugging complex systems and reasoning about behavior in real-world, noisy environments
Strong advantage
Experience working with edge or embedded systems
Experience working with custom high-performance data or inference pipelines
Familiarity with multi-sensor fusion (e.g., combining vision with radar or other signals)
Experience deploying and maintaining ML models in production environments
Experience with low-level optimization and/or C++ performance tuning
Proven experience optimizing model inference (e.g., TensorRT, ONNX Runtime, quantization, pruning, or similar techniques).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8737671
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 40 דקות
Location: More than one
Job Type: Full Time
We are seeking a highly skilled and modern software engineer to develop and prototype brand new advancements in distributed training and inference using our Spectrum-X AI fabric. This role offers a rare chance to pioneer AI and networking technology, contributing to ground-breaking projects that will define the landscape of large-scale AI systems. Improve AI app-networking connection by refining communication, crafting congestion control, coding NIC firmware, and expanding switch SDK features for enhanced AI factory efficiency. Your work impacts large AI system development, scaling, and speed.

What youll be doing:

Prototype end-to-end solutions to improve distributed training and disaggregated inference performance.

Analyze and optimize communication flows across application, transport, and network layers.

Develop system software spanning communication libraries, drivers, and firmware integrations.

Collaborate with hardware, firmware, and SDK teams to co-design network features.

Validate and integrate prototypes into our AI infrastructure and products.
Requirements:
What we need to see:

BSc/MSc/PhD in Computer Science or Electrical Engineering.

5+ years of relevant experience and/or knowledge.

Deep understanding of networking and communication internals - NCCL, RDMA/RoCE, congestion control.

Hands-on experience with HW/SW/FW integration and low-level programming (C/C++, kernel, drivers).

Some background in distributed training systems (such as PyTorch DDP, Megatron-LM, DeepSpeed).


Ways to stand out from the crowd:

Demonstrated innovation and leadership turning prototypes into impactful product features.

Experience with programmable data planes (P4, eBPF, DOCA SDK, or switch SDKs).

Familiarity with NIC firmware scheduling, in-network compute, or congestion management.

Contributions to open-source projects, academic papers, or performance benchmarking tools.

Strong background in AI factory architectures, distributed inference, or network telemetry.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8771228
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a hands-on Applied AI Scientist to join our core R&D team and drive the development of next-generation AI systems for autonomous driving. This role sits at the intersection of applied research and deployment. This role sits at the intersection of applied research and deployment. You will work directly on our multi-layered autonomy architecture, with a primary focus on real-time predictive models for driving decisions. A deep technical role for someone who thrives on turning cutting-edge research into real, working systems under hard constraints.

Responsibilities:
Own the research-to-deployment cycle for driving models - from literature review and prototyping through to production integration.
Design, implement, and iterate on real-time predictive models, including vision-language-action (VLA) models.
Collaborate on reasoning systems, contributing to VLA models that handle planning across varied horizons.
Bridge cloud-scale training with edge deployment - work on model compression, quantization, speculative decoding, and efficient inference for embedded automotive platforms.
Evaluate and integrate state-of-the-art techniques from the broader AI research community into our autonomy stack.
Collaborate closely with internal R&D teams to unblock technical challenges, accelerate delivery, and raise the overall technical bar.
Requirements:
Requirements:
Ph.D. in Computer Science, Electrical Engineering, Machine Learning, Robotics, or a related field (an MSc with an exceptional background will also be considered).
Strong publication or deployment track record in one or more of: deep learning, computer vision, generative AI, reinforcement learning, or motion prediction.
Demonstrated ability to go from paper to working implementation - not just theory, but shipped systems.
Strong coding skills in Python; experience with C++ is a plus.
Familiarity with modern ML infrastructure: PyTorch, ONNX, Triton, Dynamo, distributed training, model optimization.
Solid mathematical foundations in probability, optimization, and statistics.

Attributes:
Experience with CUDA or low-level GPU optimization.
Hands-on work with model quantization, distillation, or efficient inference on edge devices.
Background in real-time, safety-critical, or embodied AI systems (robotics, autonomous vehicles, drones, etc.).
Experience with foundation models (Language, Vision, Tabular, VLAs) and their on-device deployment.
Familiarity with driving datasets, simulation environments, or sensor fusion pipelines.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8757491
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Senior AI Engineer to design and build production-grade, LLM-powered systems. You'll work at the intersection of software engineering and applied AI - shipping agents, RAG pipelines, and tool-using systems that solve real problems at scale. This is a hands-on, high-ownership role for someone who thrives at the frontier of what's possible with modern LLMs and isn't afraid to write the glue, the infrastructure, and the prompts that make it all work.
This is a **cross-functional, company-wide role**. You won't be embedded in a single product team - instead, you'll partner with every department to identify high-leverage opportunities and build AI-powered tools and workflows that boost productivity and efficiency across the entire organization.
This is a great opportunity to be part of one of the fastest-growing infrastructure companies in history, an organization that is in the center of the hurricane being created by the revolution in artificial intelligence.
"our company's data management vision is the future of the market."- Forbes
we are the data platform company for the AI era. We are building the enterprise software infrastructure to capture, catalog, refine, enrich, and protect massive datasets and make them available for real-time data analysis and AI training and inference. Designed from the ground up to make AI simple to deploy and manage, our company takes the cost and complexity out of deploying enterprise and AI infrastructure across data center, edge, and cloud.
Our success has been built through intense innovation, a customer-first mentality and a team of fearless workers who leverage their skills & experiences to make real market impact. This is an opportunity to be a key contributor at a pivotal time in our companys growth and at a pivotal point in computing history.
What You'll Do:
- Design, build, and operate LLM-powered applications, agents, and workflows end-to-end - from prototype to production.
- Architect retrieval, context engineering, and tool-use strategies that make models reliable, accurate, and cost-efficient.
- Integrate LLMs with internal services, third-party APIs, and data stores to automate complex business and engineering workflows.
- Build, evaluate, and continuously improve evaluation harnesses for non-deterministic systems.
- Collaborate closely with product, research, and platform teams to translate ambiguous problems into shipped capabilities.
- Stay ahead of the rapidly evolving LLM ecosystem (models, frameworks, agentic patterns) and bring the best ideas into our stack.
Requirements:
Engineering Foundations:
- Strong Python skills- you write clean, idiomatic, well-tested code and understand the language deeply.
- Hands-on experience using coding agents(Cursor, Claude Code, GitHub Copilot, or similar) to build complex software systems. You know how to delegate effectively to AI assistants and review their output critically.
- Experience with multiple database paradigms- both SQL (PostgreSQL, MySQL) and NoSQL (MongoDB, Redis, DynamoDB, or similar). You can choose the right tool for the job.
- Experience designing and integrating with third-party APIs- REST and gRPC. Comfortable building robust clients, handling auth, retries, rate limits, and schema evolution.
- Production experience with Docker and Kubernetes- containerizing services, writing manifests, and debugging deployments.
- Strong Linux fundamentals- confident in bash and the terminal; you can navigate, script, and troubleshoot a server without reaching for a GUI.
- Experience building cloud-native tools on AWS, GCP, or Azure (compute, storage, queues, serverless, IAM).
AI / LLM Expertise:
- Solid understanding of what an LLM is and how it works- tokenization, attention, context windows, sampling, and the practical implications of each for system design.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8744445
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a Senior MLOps Engineer to be a core driver in how our product empowers security teams. You will be expected to deeply understand customer needs and translate them directly into product features that deliver real value. You'll own key parts of our frontend stack, drive key architectural decisions, and turn complex security data into clear, actionable business insights.

As we scale our AIDR product and expand deeper into model-driven security intelligence, we are looking for a Senior MLOps Engineer to own the infrastructure, tooling, and operational foundations that power our NLP and LLM training, evaluation, and deployment workflows.

You will architect and operate the systems that enable us to train, fine-tune, deploy, and monitor models at scale making ML reliable, fast, cost-efficient, and production-ready.

This is a high-visibility, high-impact role where you will partner closely with DevOps, Backend, Data, and Product to establish world-class ML infrastructure from the ground up.

What Youll Do

Build & Scale ML Pipelines
Design, build, and maintain pipelines for training, fine-tuning, evaluating, and deploying NLP and LLM models across GPU and CPU environments.
Establish LLM-Focused CI/CD
Implement automated CI/CD workflows for ML models, including benchmarking, testing, performance gating, and production deployment.
Optimize Runtime & Inference
Select and optimize serving frameworks for low-latency, high-throughput inference, ensuring reliability and scalability.
Own ML Infrastructure
Manage training environments, experiment tracking, model registries, artifact versioning, and distributed training systems.
Operational Excellence
Monitor and optimize production models for performance, cost efficiency, availability, and observability.
Requirements:
5+ years in software engineering, MLOps, or ML engineering with hands-on experience deploying ML models to production.
Strong Python fundamentals and deep understanding of transformer architectures, tokenization, and NLP frameworks (PyTorch, HuggingFace).
Proven experience deploying and scaling LLMs for real-time inference-ideally on platforms like SageMaker, Vertex AI, or similar.
Expertise in GPU optimization, distributed training, and CPU-based inference optimization.
Strong cloud and Kubernetes background (EKS/GKE/AKS, Helm, Terraform, CI/CD for ML).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8762083
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
חברה חסויה
Job Type: Full Time
We are seeking a Director of Software Architecture to lead and accelerate the evolution of next-generation AI data center and networking technologies. This role is designed for a hands-on, execution-driven leader who pushes boundaries, translates vision into reality, and delivers high-impact solutions at scale. You will operate at the intersection of innovation, architecture, and delivery-shaping the future of AI networking and infrastructure.

What you will be doing:
Drive identification, evaluation, and rapid adoption of emerging technologies, ensuring strong alignment with strategic roadmap and measurable business outcomes.
Lead the design and delivery of advanced networking applications, leveraging data plane programming and modern networking protocols to solve complex, large-scale challenges.
Architect solutions across AI data center environments, integrating GPU-based systems, hardware acceleration, and high-performance networking.
Act as a thought leader and industry influencer-engaging directly with customers, publishing technical content, and representing us at key conferences and forums.
Define and execute a bold architectural vision for our networking in close collaboration with cross-functional software and hardware leaders.
Requirements:
What we need to see:
M.Sc. or PhD. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
12+ years of deep experience in software architecture, systems design, and applied research.
8+ years of proven leadership, building and driving high-performing engineering teams in fast-paced environments.
Strong expertise in AI inference technologies, frameworks, and large-scale distributed systems.
Hands-on experience with networking protocols (e.g., TCP/IP, RDMA, RoCE, InfiniBand) and data center networking architectures.
Deep familiarity with AI DC architectures, hardware acceleration technologies, and SDKs (e.g., DOCA, CUDA or similar).
Experience with AI data center design, including compute, networking, and AI storage systems.
Exceptional communication and influence skills, with a demonstrated ability to align stakeholders and drive decisions across complex organizations.

Ways to stand out from the crowd:
A track record of aggressively prototyping, validating, and scaling new ideas into production.
Strong foundations in system software, including operating systems and low-level architecture.
Experience with hyperscale cloud and AI data center environments.
Expertise in AI storage systems, high-performance computing (HPC), and end-to-end accelerated infrastructure.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8767985
סגור
שירות זה פתוח ללקוחות VIP בלבד