דרושים » תוכנה » ML Engineer

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
06/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time and Hybrid work
we are looking for a ML Engineer.
Own the model plane end to end - vLLM inference, fine-tuning (GRPO/RL and LoRA), and the evaluation-and-promote pipeline that takes an owned model from checkpoint to production serving. We train and serve our own models on GPUs in our own cluster.
Requirements:
2+ years of ML engineering or applied ML research
Deep familiarity with HuggingFace transformers
Experience running inference servers (vLLM, TGI, or Triton)
Python and CUDA fundamentals
Understanding of quantization (AWQ, GPTQ, GGUF)
Bonus: GRPO/RLHF/DPO training, GPU workloads on Kubernetes
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8811620
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for our **first Forward Deployed Engineer**. You'll sit directly with our customers' engineering teams and get their workloads onto Impala - from the first scoping conversation to a production endpoint carrying real traffic, at a cost and latency profile they couldn't hit anywhere else.
This is an engineering role. You will write code every day, in customer repos and in ours. It also carries pieces of solutions architecture, product management, and pre-sales, and you should want that mix rather than tolerate it. Ambiguous business goals come in; observable, benchmarked, production services go out.
As the founding FDE you also define the function: what a POC looks like, what we promise and measure, which patterns get productized, and how the field feeds the roadmap. The next FDEs will work from what you build here.
What You'll Do
- **Own customer outcomes end to end:** problem framing, evaluation design, migration, deployment, benchmarking, monitoring. You're the technical owner from first call through expansion.
- **Win the POC:** turn a vague objective into a tight spec and a working proof of concept fast, with explicit quality, latency, throughput, and cost-per-token targets - and hit them.
- **Tune serverless inference for real workloads:** model and engine selection, batching strategy, KV cache behavior, speculative decoding, quantization, parallelism, cold-start and autoscaling behavior under bursty traffic. Diagnose regressions down to the inference engine.
- **Migrate workloads onto Impala:** move customers off OpenAI-compatible APIs, self-managed vLLM, SageMaker, or their own GPU fleets - and prove out the quality and cost delta with numbers.
- **Guide model strategy:** advise on open-weight model selection, distillation, and fine-tuning for specific tasks; help customers get from a general-purpose frontier model to a smaller, faster, cheaper one that holds quality.
- **Close the product loop:** bring the field back into the roadmap - write the PRDs, land the PRs, and turn one-off customer work into platform features.
- **Be the technical anchor in the room:** support sales on complex evaluations, run technical onboarding, earn trust with staff engineers and CTOs.
Requirements:
- **4+ years** building and shipping production software
- **Inference in production:** hands-on experience serving LLMs with **vLLM, SGLang, TensorRT-LLM** or equivalent, and real intuition for what makes inference fast or expensive.
- **Optimization fundamentals:** working knowledge of batching, KV cache, quantization, speculative decoding, tensor and pipeline parallelism - and the tradeoffs between them.
- **Model judgment:** fluency with the open-weight model landscape and good instincts on model selection for a given task, hardware profile, and latency budget.
- **Production cloud comfort:** containers, Kubernetes, observability, CI/CD. You don't need to be an SRE, but nothing here should be a black box.
- **Range in the room:** you can hold a technical conversation with a skeptical staff engineer and a commercial one with their VP, in the same meeting.
- **Default ownership:** ambiguity, unfamiliar codebases, and a customer waiting on you are the normal conditions of this job, not exceptions.
- Prior forward-deployed, solutions architecture, or applied ML engineering experience at an infrastructure company.
- Post-training experience: LoRA, SFT, DPO, RLHF, GRPO, distillation.
- GPU-level performance work CUDA, Triton, memory bandwidth and throughput profiling.
- Experience selling into or building for regulated, data-sensitive enterprises.
- Open-source contributions to inference or serving projects.
- You've been the first or second person in a function before.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8807348
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8781454
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774268
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
11/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications:
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications:
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776998
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
26/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
What You'll Do
- Design and ship the ML backbone of Gini AI Workers - routing, tool selection, reasoning, memory, evaluation.
- Build evaluation and feedback loops - offline evals, online A/B, regression harnesses, human-in-the-loop labeling pipelines.
- Optimize cost and latency across the agent stack: prompt engineering, model routing (frontier ↔ small ↔ fine-tuned), caching, speculative decoding, distillation.
- Fine-tune and/or RAG-tune models for vertical enterprise tasks (invoice extraction, PO matching, ticket triage, forecasting).
- Own the ML infra - training pipelines, experiment tracking, model registry, deployment, monitoring, drift detection.
- Partner with backend + product to turn research into shipped features on a weekly cadence.
Requirements:
- 4+ years of ML engineering in production (not just research or notebooks).
- Hands-on LLM experience in 2025-2026: agentic systems, tool-use, function-calling, RAG, structured output, eval design.
- Strong Python. Comfortable with PyTorch/JAX and one serving stack (vLLM, TGI, TensorRT-LLM, SageMaker, or similar).
- You've built an eval pipeline that actually caught a regression in prod.
- You read the papers and know which ones to ignore.
Nice to Have
- Experience with MCP, LangGraph, DSPy, or custom agent frameworks.
- Fine-tuning (LoRA/QLoRA, DPO/ORPO, RLAIF) on open-weight models (Llama, Qwen, Mistral, DeepSeek).
- Vector DBs (pgvector, Pinecone, Weaviate, Qdrant), reranking, hybrid retrieval.
- Prior work on multi-agent systems or enterprise copilots.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8798773
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a strong Backend Software Engineer to bridge the gap between our Machine Learning research team and our enterprise production systems. You will act as the technical backbone for our ML Scientists - by advising, designing and implementing the production facing features. If you are a backend expert who wants to solve complex system architecture challenges and dive into the world of ML platforms & Agentic LLM pipelines, this is the role for you - An exciting role collaborating with ML science team, data/infra team and DevOps to drive real customer impact.



As a ML Engineer, you will:



Lead ML delivery: transforming research output (code, models, ideas) into robust, scalable, low-latency microservices in production

Help architect e2e solutions to real customer pains ranging from ingestion, integration, ETLs, DB design up to low-latency services

Design, build, and maintain automated workflows for ML models, including auto-trains, benchmarking, testing, performance gating, and production deployment.

Tackle complex backend challenges: optimizing API response times, managing database connectivity and concurrency at scale, balancing accuracys drive for complex questions with the business needs of fast responsiveness by making hard technical trade-offs between customer gains and business costs.

Design and optimize data pipelines and ETL processes, connecting our Snowflake data warehouse to our training environments.

Work within our existing ML infrastructure (Kubeflow, MLflow, KServe) to ensure smooth model lifecycles and performance monitoring.

Collaborate closely with ML Scientists, guiding them on software engineering best practices without slowing down their research.

Monitor and optimize production models for performance, cost efficiency, availability, and observability.
Requirements:
6+ years of backend software engineering experience designing, building, and maintaining large-scale, high-throughput production systems

Strong coding skills, Ability to write clean, maintainable code, OOP familiarity, package design, microservices etc.
Note: Work is in python, but strong engineers with deep Java/C# backgrounds who have some Python experience and are willing to transition fully are highly encouraged to apply.

Solid Database design & SQL skills, Deep understanding of SQL, experience working with relational and/or bigdata (columnar) databases, ORMs, and efficient query design.

API & Performant Design Proven experience - building robust systems, you understand how to handle concurrency, ETL tradeoffs, building fault-tolerant best effort data flows
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8818291
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
17/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Join our companys AI research group, a cross-functional team of ML engineers, researchers and security experts building the next generation of AI-powered security capabilities. Our mission is to leverage large language models to understand code, configuration, and human language at scale, and to turn this understanding into security AI capabilities which will drive our company AI future security solutions.
We foster a hands-on, research-driven culture where youll work with large-scale data, modern ML infrastructure, and a global product footprint that impacts over 100,000 organizations worldwide.
Key Responsibilities
Your Impact & Responsibilities
As a Senior ML Research Engineer, you will be responsible for the end-to-end lifecycle of large language models: from data definition and curation, through training and evaluation, to providing robust models that can be consumed by product and platform teams.
Own training and fine-tuning of LLMs / seq2seq models: Design and execute training pipelines for transformer-based models (encoder-decoder, decoder-only, retrievalaugmented, etc.), and fine-tune open-source LLMs on our company-specific data (security content, logs, incidents, customer interactions).
Apply advanced LLM training techniques such as instruction tuning, preference / contrastive learning, LoRA / PEFT, continual pre-training, and domain adaptation where appropriate.
Work deeply with data: define data strategies with product, research and domain experts; build and maintain data pipelines for collecting, cleaning, de-duplicating and labeling large-scale text, code and semi-structured data; and design synthetic data generation and augmentation pipelines.
Build robust evaluation and experimentation frameworks: define offline metrics for LLM quality (task-specific accuracy, calibration, hallucination rate, safety, latency and cost); implement automated evaluation suites (benchmarks, regression tests, redteaming scenarios); and track model performance over time.
Scale training and inference: use distributed training frameworks (e.g. DeepSpeed, FSDP, tensor/pipeline parallelism) to efficiently train models on multi-GPU / multi-node clusters, and optimize inference performance and cost with techniques such as quantization, distillation and caching.
Collaborate closely with security researchers and data engineers to turn domain knowledge and threat intelligence into high-value training and evaluation data, and to expose your models through well-defined interfaces to downstream product and platform teams.
Requirements:
What You Bring
5+ years of hands-on work in machine learning / deep learning, including 3+ years focused on NLP / language models.
Proven track record of training and fine-tuning transformer-based models (BERT-style, encoder-decoder, or LLMs), not just consuming hosted APIs.
Strong programming skills in Python and at least one major deep learning framework (PyTorch preferred; TensorFlow).
Solid understanding of transformer architectures, attention mechanisms, tokenization, positional encodings, and modern training techniques.
Experience building data pipelines and tools for large-scale text / log / code processing (e.g. Spark, Beam, Dask, or equivalent frameworks).
Practical experience with ML infrastructure, such as experiment tracking (Weights & Biases, MLflow or similar), job orchestration (Airflow, Argo, Kubeflow, SageMaker, etc.), and distributed training on multi-GPU systems.
Strong software engineering practices: version control, code review, testing, CI/CD, and documentation.
Ability to own research and engineering projects end-to-end: from idea, through prototype and controlled experiments, to models ready for integration by product and platform teams.
Good communication skills and the ability to work closely with non-ML stakeholders (security experts, product managers, engineers).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8785689
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for an experienced and passionate ML Engineering Team Lead to lead our ML Engineering team and shape the next generation of our AI infrastructure. This is a hands-on leadership role where you'll combine technical leadership, software architecture, and people management to build scalable, production-ready AI systems running on edge devices.
About The Role:
Lead, mentor, recruit, and grow a team of software engineers, fostering a culture of ownership, collaboration, and continuous improvement.
Own the team's technical roadmap, architecture, execution, and project prioritization, aligning delivery with business goals.
Design, build, and maintain scalable software and ML infrastructure across cloud and edge environments.
Partner with AI Researchers to productionize Computer Vision and Deep Learning models into reliable, high-performance systems.
Design and optimize inference pipelines with a focus on scalability, latency, and reliability.
Drive engineering excellence through architecture reviews, code reviews, development best practices, and modern AI-assisted engineering workflows.
Requirements:
6+ years of software development experience, including 3+ years leading software engineering or ML engineering teams.
Strong hands-on experience with Python and C++ or Rust.
Experience building, deploying, and maintaining production-grade Machine Learning systems.
Strong understanding of software architecture, scalable system design, and performance optimization.
Experience collaborating with AI, Machine Learning, or Computer Vision teams.
Excellent leadership, communication, and organizational skills, with a strong ownership mindset.
Experience using modern AI-assisted development tools (such as Cursor, Claude Code, or Codex) while maintaining high engineering quality.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8822310
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a hands-on AI Platform Team Lead to build and lead the team behind this platform: a high-throughput, low-latency engine that runs GPU-based models, from MMBERT-style models to LLMs, together with CPU-based heuristics and security logic.
This is a core infrastructure role for someone who wants to own the runtime layer of AI security at scale: performance, reliability, orchestration, GPU efficiency, and production-grade execution in the traffic path.
The team will also own the model lifecycle required to take AI security algorithms from research to large-scale production, working closely with research and algorithm teams.


Responsibilities
Build and lead Catos AI Platform team: hiring, mentoring, architecture, technical direction, and execution.
Own the AI security runtime platform for high-throughput, low-latency inline security decisions across Catos global cloud and PoPs.
Design the orchestration layer for running GPU models, CPU heuristics, and security logic as one production engine.
Own production readiness: observability, SLOs, autoscaling, reliability, rollout, rollback, and operational health.
Own the model lifecycle platform: registry, versioning, deployment, monitoring, and safe production rollout.
Work closely with research and algorithm teams to productionize AI security models and algorithms at scale.
Define the long-term platform strategy for AI runtime and model serving at Cato.
Requirements:
3+ years of leadership experience as a team lead, tech lead, or engineering manager.
3+ years of hands-on experience in AI inference, production ML infrastructure, model serving, or AI runtime platforms.
Strong experience with production inference technologies such as Triton, vLLM, CUDA, Kubernetes, Docker, PyTorch, ONNX, TensorRT, or similar.
3+ years of experience with Go, or strong experience with a similar high-performance backend language such as C++, Rust, or Java.
Experience with performance optimization, scalability, observability, and SLO-driven production ownership.
Strong system design skills, especially around distributed systems, performance, reliability, and production infrastructure.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8822396
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are on a mission to create public transportation systems that provide far greater access to jobs, healthcare, and education. Our platform serves as the technology backbone for modern transit networks, transforming antiquated and siloed public transportation systems into smart, data-driven, and efficient digital networks. With hundreds of agency partners around the world, we are recognized as the leading transportation technology and service provider globally.
As a Staff ML Engineer at our company, you will play a central role in shaping how millions of riders and drivers move through our cities every day. Our team sits at the intersection of machine learning, optimization, and real-world operations - turning complex, multi-dimensional challenges into the real-time intelligence that powers our company's mass-scale automated dispatch system. This is a rare opportunity to work on problems that are genuinely hard, at a scale that is genuinely rare, where the solutions you build have a direct and visible impact on the efficiency and reliability of transit networks around the world.
About the Role:
Own the development of ML models and optimization algorithms that drive our company's real-time dispatch system - making smart, scalable decisions across thousands of simultaneous rides, drivers, and operational constraints.
Design and implement online algorithms for real-time decision-making, balancing system utilization with a consistently high quality of service for riders - where every millisecond and every percentage point of efficiency matters.
Model and mathematically represent competing demands on our company's system, translating messy real-world operational complexity into elegant, tractable formulations that can be solved at scale.
Use sophisticated statistical methods to analyse demand patterns, traffic dynamics, and fleet performance - generating insights that directly inform algorithm development and operational strategy.
Collaborate closely with engineering, product, and operations teams to bring complex algorithmic work to life in production - owning the full journey from research and prototyping through to real-world deployment and iteration.
Requirements:
Advanced degree (M.Sc. or PhD) in Computer Science, Mathematics or a closely related field, with a strong background in Machine Learning.
8+ years of industry experience shipping machine learning models at production scale - you've taken hard problems from whiteboard to deployment and know what it takes to make research work in the real world.
Deep, hands-on expertise in Machine Learning, with significant experience in reinforcement learning for complex, dynamic or constrained systems.
Excellent coding skills in Python or similar, with the ability to turn rigorous ideas into working, maintainable solutions that perform under real operational load.
Strong applied research mindset: able to translate ambiguous business/operational challenges into tractable ML formulations and measurable impact.
Naturally curious, fast-learning, and collaborative - you bring strong communication skills and genuine intellectual generosity to the teams you work with.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796348
סגור
שירות זה פתוח ללקוחות VIP בלבד