דרושים » הנדסה » Researcher

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Join our company, an innovative startup building a fully-managed LLM-inference platform, that enables data heavy enterprises to perform any AI task at any scale without limits.
We're looking for an Experienced Performance Researcher to join our founding team. Youll be responsible for building and optimizing scalable cloud infrastructure solutions tailored for AI workloads. This role offers a unique opportunity to directly shape our infrastructure strategy, improve system reliability and performance, and contribute to establishing our company as a leader in adaptive AI compute management.
Join us to tackle the magic that make AI tick under the hood and build the backbone powering the AI revolution.
What Youll Do
- Design and build high-performance distributed inference pipelines for LLMs, focused on large-batch, non-real-time scenarios.
- Optimize GPU memory usage, kernel execution, and communication across nodes (NCCL, MPI, etc.).
- Own CUDA kernels, compiler-level tricks, and multi-GPU scheduling logic.
- Lead profiling and performance tuning for throughput, and cost- down to the kernel level.
- Collaborate with infra, product, and research teams to define SLAs, resource allocation logic, and runtime behaviors.
- Help build the core infrastructure that will run LLM workloads across hybrid GPU environments (cloud/on-prem/self-hosted).
Requirements:
- Deep experience with CUDA programming, GPU architecture, and low-level performance engineering.
- Fluency with Python and C++, and a mastery of profiling tools like Nsight, nvprof, perf, etc.
- Experience building systems for large-scale distributed training or inference (PyTorch, DeepSpeed, Ray, Horovod, etc.).
- Hands-on familiarity with cluster and container orchestration tools (Kubernetes, Slurm, Docker).
- Self-motivated and able to operate independently in a fast-moving startup environment.
- Strong analytical skills and a passion for elegant performance wins.
- A collaborative team player with strong interpersonal skills, a positive and easygoing attitude, and the potential to grow into a leadership role.
- Prior experience building inference runtimes or scheduling frameworks.
- Experience with serverless GPU models, model parallelism, tensor slicing, and batching tricks.
- Contributions to open-source HPC or ML infra projects.
- Understanding of AI/ML privacy and compliance concerns in enterprise environments.
- Track record of working on distributed systems at bleeding-edge research labs or infrastructure teams.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8807342
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for our **first Forward Deployed Engineer**. You'll sit directly with our customers' engineering teams and get their workloads onto Impala - from the first scoping conversation to a production endpoint carrying real traffic, at a cost and latency profile they couldn't hit anywhere else.
This is an engineering role. You will write code every day, in customer repos and in ours. It also carries pieces of solutions architecture, product management, and pre-sales, and you should want that mix rather than tolerate it. Ambiguous business goals come in; observable, benchmarked, production services go out.
As the founding FDE you also define the function: what a POC looks like, what we promise and measure, which patterns get productized, and how the field feeds the roadmap. The next FDEs will work from what you build here.
What You'll Do
- **Own customer outcomes end to end:** problem framing, evaluation design, migration, deployment, benchmarking, monitoring. You're the technical owner from first call through expansion.
- **Win the POC:** turn a vague objective into a tight spec and a working proof of concept fast, with explicit quality, latency, throughput, and cost-per-token targets - and hit them.
- **Tune serverless inference for real workloads:** model and engine selection, batching strategy, KV cache behavior, speculative decoding, quantization, parallelism, cold-start and autoscaling behavior under bursty traffic. Diagnose regressions down to the inference engine.
- **Migrate workloads onto Impala:** move customers off OpenAI-compatible APIs, self-managed vLLM, SageMaker, or their own GPU fleets - and prove out the quality and cost delta with numbers.
- **Guide model strategy:** advise on open-weight model selection, distillation, and fine-tuning for specific tasks; help customers get from a general-purpose frontier model to a smaller, faster, cheaper one that holds quality.
- **Close the product loop:** bring the field back into the roadmap - write the PRDs, land the PRs, and turn one-off customer work into platform features.
- **Be the technical anchor in the room:** support sales on complex evaluations, run technical onboarding, earn trust with staff engineers and CTOs.
Requirements:
- **4+ years** building and shipping production software
- **Inference in production:** hands-on experience serving LLMs with **vLLM, SGLang, TensorRT-LLM** or equivalent, and real intuition for what makes inference fast or expensive.
- **Optimization fundamentals:** working knowledge of batching, KV cache, quantization, speculative decoding, tensor and pipeline parallelism - and the tradeoffs between them.
- **Model judgment:** fluency with the open-weight model landscape and good instincts on model selection for a given task, hardware profile, and latency budget.
- **Production cloud comfort:** containers, Kubernetes, observability, CI/CD. You don't need to be an SRE, but nothing here should be a black box.
- **Range in the room:** you can hold a technical conversation with a skeptical staff engineer and a commercial one with their VP, in the same meeting.
- **Default ownership:** ambiguity, unfamiliar codebases, and a customer waiting on you are the normal conditions of this job, not exceptions.
- Prior forward-deployed, solutions architecture, or applied ML engineering experience at an infrastructure company.
- Post-training experience: LoRA, SFT, DPO, RLHF, GRPO, distillation.
- GPU-level performance work CUDA, Triton, memory bandwidth and throughput profiling.
- Experience selling into or building for regulated, data-sensitive enterprises.
- Open-source contributions to inference or serving projects.
- You've been the first or second person in a function before.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8807348
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
23/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are a well-funded, early-stage startup looking for a talented and motivated Backend Engineer specializing in infrastructure to join our founding team. The focus of this role is to build and scale the infrastructure that powers autonomous AI agents automating complex enterprise workflows. You will own the systems, pipelines, and platforms that let our AI agents run reliably, securely, and at scale in production.

Your Impact
Infrastructure & Platform

Design, build, and own the core infrastructure powering our AI agent platform, from data pipelines to production deployment systems.

Build and scale the backend systems that support high-throughput document processing and data extraction workloads.

Cloud Infrastructure and Scalability

Architect and deploy infrastructure on cloud platforms (AWS, GCP, or Azure) with a focus on scalability, reliability, and cost efficiency.

Own containerization and orchestration (Docker, Kubernetes) for all production workloads.

Build and maintain CI/CD pipelines and DevOps practices that let the team ship fast without breaking things.

Data Infrastructure

Design and manage data pipelines to process and analyze large volumes of documents and unstructured data at scale.

Build the infrastructure layer connecting AI agents to databases, vector stores, and enterprise systems (ERP, CRM).

API & Systems Integration

Build and maintain robust, well-documented APIs connecting AI agents with external systems and enterprise software.

Design for reliability: retries, observability, and graceful degradation across distributed systems.

Security and Compliance

Implement authentication and authorization mechanisms (OAuth2, JWT) to secure AI-driven systems.

Ensure compliance with data privacy standards (e.g. GDPR, HIPAA) and drive best practices for secure data handling across the infrastructure.

Monitoring and Optimization

Build observability and monitoring systems to track infrastructure health, performance, and cost.

Continuously optimize system performance for speed, reliability, and cost-efficiency at scale.

Collaboration

Work closely with AI/ML engineers, product, and the founding team to make sure infrastructure decisions support fast iteration and production-grade reliability.

Participate in code reviews, design discussions, and architecture planning to drive infrastructure strategy.
Requirements:
5+ years of experience in backend or infrastructure engineering, ideally supporting production AI/ML systems or high-throughput data pipelines.

Proven track record of building and scaling infrastructure in production environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8793026
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
11/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications:
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications:
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776998
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
This role sits at the intersection of AI and high-performance systems engineering, focused on solving real-world problems under strict constraints. You will work on systems where performance and reliability are critical and where improvements have a direct, measurable impact on real-world safety.

This is a senior, systems-focused role with end-to-end ownership over performance and reliability of production computer vision pipelines. You will define optimization strategies, identify bottlenecks across the system, and drive improvements under real-world constraints.

What youll do
Build and optimize real-time computer vision pipelines running on edge systems processing live maritime video streams (e.g, NVIDIA Jetson, Triton Inference Server).
Take models from research and turn them into production-ready, reliable components deployed on vessels.
Profile and improve end-to-end system performance across: multi-camera video ingestion; preprocessing; inference; postprocessing
Identify and resolve bottlenecks across CPU, GPU, memory, and pipeline coordination.
Make and justify tradeoffs between latency, accuracy, stability, and resource utilization.
Design and implement robust data and inference pipelines (video -> model -> actionable output for crew).
Develop benchmarking and evaluation workflows to measure performance end-to-end and support release gating.
Build and improve observability tools, including logging, monitoring, and debugging workflows for production systems.
Define and maintain clear interfaces between research code and production systems.
Work closely with research and backend teams to integrate new models into production systems.
Continuously improve system efficiency and reliability under hardware and runtime constraints.
Requirements:
Requirements:
5+ years of software engineering experience, with a strong focus on systems and performance.
Hands-on experience working with computer vision or deep learning systems in production.
Strong programming skills in Python and/or C++.
Experience working with edge or embedded systems (e.g., NVIDIA Jetson platforms).
Strong understanding of system bottlenecks, including CPU, GPU, memory, and latency constraints.
Strong intuition for profiling-driven optimization and performance tuning.
Experience debugging complex systems and reasoning about behavior in real-world, noisy environments.

Strong advantage:
Experience working with edge or embedded systems.
Experience working with custom high-performance data or inference pipelines.
Familiarity with multi-sensor fusion (e.g., combining vision with radar or other signals).
Experience deploying and maintaining ML models in production environments.
Experience with low-level optimization and/or C++ performance tuning.
Proven experience optimizing model inference (e.g., TensorRT, ONNX Runtime, quantization, pruning, or similar techniques).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796312
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8781454
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
10/09/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs.



Some directions we are currently working on, and which you can be a part of:

Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5).
Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines.
Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency.
Requirements:
A profound understanding of theoretical foundations of machine learning and transformer architecture.
Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools
Understanding of GPU memory hierarchy and compute/memory tradeoffs
Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation
Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)
Strong software engineering skills (we mostly use Python)
Deep experience with modern deep learning frameworks
Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing
Strong communication and leadership abilities
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8817627
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
23/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We're forming a new AI Group and looking for a Senior AI Engineer to help shape it from an early stage - a greenfield, long-term effort to evolve how decisions are made across the platform using AI-driven systems. You won't just integrate APIs or build demos; you'll build the AI brain that works alongside (and increasingly drives) our core automation engine, with real production impact from day one and room to grow into technical leadership as the group scales.

Agentic AI Architecture: Design and build autonomous AI agents that analyze infrastructure in real time and make intelligent decisions. Work with modern agentic frameworks (LangGraph, PydanticAI) and conversational AI to create multi-agent systems - including troubleshooting, optimization, FinOps, and how-to agents. Leverage core LLM capabilities (tool-use, memory, retrieval) to operate safely in production.
Platform Integration & Intelligent Decision Systems: Develop MCPs to expose capabilities to AI agents that reason over infrastructure environments, metrics, configurations, and cost signals. Build integrations with tools like Slack, Jira, and AI-powered IDEs (Cursor, Windsurf) to deliver context-aware insights, from "why is this pod not scheduling?" to "how can we reduce costs by 30% safely?"
AI Model Development & MLOps: Build and deploy machine learning models that learn from infrastructure patterns - detecting the right resource policies for workloads, predicting optimal scaling triggers, and recommending GPU configurations. Own the complete ML pipeline from training to production, ensuring models are reliable, monitored, and continuously improving.
R&D AI Tools Development & Adoption: Build and embed internal AI tools to accelerate engineering, development, research, and support.
AI Tools for Business Impact: Develop AI-powered tools that help Sales and Support teams demonstrate value instantly - agents that analyze customer infrastructure, generate cost optimization reports automatically, and turn technical data into clear business recommendations.
End-to-End Ownership: Own AI systems from concept to production, ensuring they're fast (sub-2-second responses), reliable, safe, and cost-effective. Build evaluation frameworks to measure quality, implement security controls, and balance performance tradeoffs in production.
Technical Leadership: Define AI architecture and best practices as a founding member of the AI team. Make key technical decisions - choosing frameworks, designing multi-agent systems, establishing data governance - and shape how evolves from AI-enhanced internal tools to customer-facing AI products.
Requirements:
Core Engineering: Significant software engineering experience (typically 4+ years) with strong Python skills and solid backend engineering fundamentals.
Production Experience: Experience building and operating production systems in cloud environments.
Real-World GenAI Experience: Practical experience bringing LLM-based systems into production, including handling latency, cost control, and failure modes. Familiarity with additional agentic frameworks (e.g., LangChain, MetaGPT) and evaluation frameworks.
Builder Mentality: Strong ownership and the ability to operate independently while collaborating closely across teams, with the motivation to grow into technical leadership as the group expands.
(Advantage) Data & RAG: Experience enabling LLMs to consume structured or operational data (configurations, logs, metrics) and experience with retrieval systems (RAG) or vector databases.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8792284
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a strong Backend Software Engineer to bridge the gap between our Machine Learning research team and our enterprise production systems. You will act as the technical backbone for our ML Scientists - by advising, designing and implementing the production facing features. If you are a backend expert who wants to solve complex system architecture challenges and dive into the world of ML platforms & Agentic LLM pipelines, this is the role for you - An exciting role collaborating with ML science team, data/infra team and DevOps to drive real customer impact.



As a ML Engineer, you will:



Lead ML delivery: transforming research output (code, models, ideas) into robust, scalable, low-latency microservices in production

Help architect e2e solutions to real customer pains ranging from ingestion, integration, ETLs, DB design up to low-latency services

Design, build, and maintain automated workflows for ML models, including auto-trains, benchmarking, testing, performance gating, and production deployment.

Tackle complex backend challenges: optimizing API response times, managing database connectivity and concurrency at scale, balancing accuracys drive for complex questions with the business needs of fast responsiveness by making hard technical trade-offs between customer gains and business costs.

Design and optimize data pipelines and ETL processes, connecting our Snowflake data warehouse to our training environments.

Work within our existing ML infrastructure (Kubeflow, MLflow, KServe) to ensure smooth model lifecycles and performance monitoring.

Collaborate closely with ML Scientists, guiding them on software engineering best practices without slowing down their research.

Monitor and optimize production models for performance, cost efficiency, availability, and observability.
Requirements:
6+ years of backend software engineering experience designing, building, and maintaining large-scale, high-throughput production systems

Strong coding skills, Ability to write clean, maintainable code, OOP familiarity, package design, microservices etc.
Note: Work is in python, but strong engineers with deep Java/C# backgrounds who have some Python experience and are willing to transition fully are highly encouraged to apply.

Solid Database design & SQL skills, Deep understanding of SQL, experience working with relational and/or bigdata (columnar) databases, ORMs, and efficient query design.

API & Performant Design Proven experience - building robust systems, you understand how to handle concurrency, ETL tradeoffs, building fault-tolerant best effort data flows
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8818291
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for an Engineering Team Lead, who will be responsible for the foundational infrastructure framework used by all our company engineering teams to build, deploy, and operate AI agents safely in production. We are building the "operating system" for AI at our company, covering agent sessions, memory management, tool orchestration, durable execution, and multi-tenant isolation. You will lead a high-impact team of 5 engineers to create the runtime and platform that defines the future of autonomous enterprise intelligence.
Youll Own:
Agentic Framework Architecture: Designing and building our companys internal agentic framework, leveraging and integrating industry-standard tools such as LangChain, LangSmith, ADK, and similar ecosystems.
Evaluation and Quality Systems: Building evaluation frameworks and workflows for AI agents, including offline and online evaluations, quality metrics, regression detection, and experimentation infrastructure.
Team Leadership & Mentorship: Leading a squad of 3-4 senior engineers, fostering a culture of technical excellence, and managing end-to-end delivery in a fast-paced environment. You will spend approximately 50% of your time hands-on, architecting core systems and reviewing code, and 50% leading the team, mentoring engineers, and aligning with cross-functional stakeholders.
Observability, Monitoring, and Guardrails: Providing the organization with robust observability capabilities for AI agents, including tracing, logging, monitoring, cost tracking, and safety guardrails to ensure reliable and responsible usage.
Developer Enablement Platforms: Creating APIs, SDKs, and abstractions that enable product teams to easily build, test, and operate agents while adhering to platform standards.
Cross-Language Integrations: Designing integrations and tooling across Python and Java to enable seamless adoption of the AI framework within our companys broader backend ecosystem.
Youll Solve:
Agent Lifecycle and Orchestration Complexity: Managing agent execution, tool usage, memory, workflows, and failure modes in production-grade systems.
AI System Reliability at Scale: Ensuring agents remain observable, debuggable, and safe as usage scales across teams and products.
Evaluation and Drift Challenges: Detecting quality regressions, model behavior changes, and unintended agent behaviors through robust evaluation and monitoring systems.
Platform Adoption Friction: Balancing flexibility with guardrails so teams can innovate quickly without compromising reliability, security, or cost controls.
Youll Impact:
Company-Wide AI Enablement: Empowering every engineering team at our company to build agent-based solutions faster, with higher quality and confidence.
Foundational AI Infrastructure: Establishing the core frameworks, evaluations, and observability standards that all AI agents at our company will rely on.
AI Safety and Quality Bar: Raising the bar for how AI systems are evaluated, monitored, and governed across the company.
Requirements:
8+ years of backend engineering experience, with strong system design and platform-building expertise. Tech leadership or team leading experience is an advantage.
Strong analytical and problem-solving skills, with the ability to debug and resolve complex technical issues efficiently.
Hands-on experience with agentic systems and frameworks such as LangChain, LangSmith, ADK, or equivalent agent orchestration platforms.
Strong understanding of AI evaluation methodologies, including agent evaluations, prompt evaluation, regression testing, and quality monitoring.
High proficiency in Python for building production-grade AI frameworks and services.
Familiarity with Java and experience integrating backend platforms or tooling into Java-based systems.
Experience building observability, monitoring, or platform tooling for distributed systems.
Strong analytical skills and the ability to reason about complex, evolving AI-driven systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8788410
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
08/09/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Senior Data Engineer to build and scale the data foundation behind platform and products. You will own complex data end-to-end - from ingestion and transformation through modeling, quality, observability, and production delivery.
This is a hands-on senior IC role for a strong builder who can solve difficult data problems independently, set a high technical bar, and collaborate closely with engineering, DS, and product. You will turn large, fragmented datasets into reliable, reusable capabilities that power every product.
Responsibilities
Build and own scalable data pipelines- Design, implement, and operate robust pipelines for high-volume structured and unstructured data, with validation, monitoring, lineage, and recovery built in.
Scale the platform for growth- A key near-term initiative is re-architecting the system to support a significantly larger customer base. You will own performance and cost-efficiency across pipelines and services, keeping reliability and operating costs under control as the platform scales.
Build across the stack- This is not a pipelines-only role. You will also write backend services and some frontend, including the internal backoffice the team runs on. We hire builders, not narrow specialists.
Own the core data tables- Own schema design and evolution, data contracts, and the modeling standards the team follows - naming, shared dimensions, normalization, documentation. Be accountable when a table is wrong, late, or drifting.
Level up the teams data work- Pair with and advise software engineers and data scientists on Spark, SQL, and modeling, and help turn notebook-grade code into production-grade pipelines.
Partner cross-functionally- Translate product, client, compliance, and business requirements into clear technical designs and dependable production systems.
Requirements:
Spark at scale- You have tuned real Spark jobs for performance and cost - skew, shuffle, partitioning, memory, spill - run pipelines over TB-scale or billions of rows in production, and can reason about the physical execution plan, not just write DataFrame code.
5+ years of professional experience building and owning production systems.
Strong Python and SQL, with maintainable, tested production code.
Strong software engineering fundamentals across the stack. You can own backend services and pick up frontend when the work needs it - not a pipelines-only specialist.
AI-first way of working- You build with AI in your day-to-day development, using it to move faster and raise the quality of what you ship.
Deep experience designing and operating ETL/ELT pipelines, data models, and distributed data-processing systems.
Comfortable advising and pairing with other engineers and data scientists on data work.
Strong AWS experience: S3, Glue, EMR, Athena, and related compute and orchestration services.
Experience with modern data lakehouse or warehouse architectures. Apache Iceberg is a strong advantage.
Experience with workflow orchestration (Airflow or similar), CI/CD, Docker, Git, and infrastructure as code such as AWS CDK and CloudFormation.
Strong understanding of data quality, schema evolution, lineage, observability, privacy, security, and access controls. Experience with regulated or sensitive data, such as healthcare / PHI, is an advantage.
High comfort in a fast-moving environment with incomplete requirements, high ownership, and a strong sense of urgency.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8814966
סגור
שירות זה פתוח ללקוחות VIP בלבד