דרושים » הנדסה » Software Development Engineer (AWS ML), Machine Learning Israel (MLIL) - FLOW sub-team (Fleet Lifecy

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL FLOW team is looking for a Software Development Engineer to design and build automation, tooling, and monitoring systems for our next-generation ML accelerator servers. We build production software to validate, initialize, monitor, and qualify these servers - from first silicon through fleet-scale deployment. Our work spans hardware diagnostics, manufacturing test automation, CI/CD pipelines, operational dashboards, and data-driven fleet health monitoring.

Key job responsibilities
Design and develop software infrastructure - automation frameworks, deployment systems, and test orchestration platforms that run at scale across manufacturing and production environments.
Work cross-functionally with Hardware, Manufacturing, and EC2 teams to automate coordinated software delivery and qualification workflows.
Debug and root-cause hardware/software interaction failures using systematic data analysis and automation-assisted triage.
Build and own CI/CD pipelines end-to-end: from code commit through build, test, deploy, and production validation - driving fast, reliable software delivery for hardware teams.
Create data pipelines and analytics systems (ETL, aggregation, real-time reporting) that transform raw hardware test results into actionable engineering insights.
Develop monitoring dashboards, alerting systems, and data visualization tools for fleet health, yield tracking, and performance benchmarking.
Own features end-to-end: from design through implementation, testing, deployment, and operational excellence.
Requirements:
Basic Qualifications
- 3+ years of software development engineer or related occupational experience.
- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering or a related discipline or equivalent.
- Experience using Linux, demonstrating proficiency with associated tools or languages.
- Can work proactively and independently, meet deadlines, and deliver on projects and tasks.
- Knowledge of software engineering best practices across the development life cycle, including agile methodologies, coding standards, code reviews, source management, build processes, testing, and operations.

Preferred Qualifications
- Experience building monitoring dashboards and data visualization (Grafana, CloudWatch, QuickSight, or similar).
- Experience with data pipelines, ETL, or analytics (S3, Athena, Spark, or similar).
- Experience with systems programming languages (C, C++, Rust).
- Familiarity with computer architecture concepts (PCIe, memory hierarchy, power management).
- Advantage: experience with hardware bring-up, ASIC/FPGA validation, or manufacturing test development.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774254
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
11/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications:
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications:
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776998
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774268
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a Staff Systems Software Engineer to design and build the foundational infrastructure that powers products used by billions of people worldwide. In this role, you will architect and implement large-scale distributed systems, low-level platform components, and high-performance services that underpin our core product stack. You will drive technical strategy across system reliability, performance, and scalability, partnering closely with product, infrastructure, and data engineering teams to deliver systems that operate at global scale with high availability and efficiency.
Software Engineer, Systems Responsibilities
Architect and implement large-scale distributed systems and platform services that support high-throughput, low-latency workloads across our product infrastructure
Lead the technical design of systems components including storage layers, compute pipelines, networking abstractions, and service orchestration frameworks
Identify and resolve systemic performance bottlenecks through instrumentation, profiling, and targeted optimization across the full systems stack
Define and enforce service level objectives for owned systems, building dashboards, alerting pipelines, and runbooks to reduce mean time to mitigation during incidents
Drive reliability improvements by reducing failure surface, designing resilient rollout strategies, and leading regular resiliency and overload testing exercises
Collaborate with cross-functional partners across product engineering, infrastructure, and data science to align system architecture with evolving product and business requirements
Establish and evolve coding standards, architectural patterns, and engineering best practices for systems development across the broader organization
Leverage AI-assisted development workflows to accelerate design iteration, code generation, and systems analysis, applying sound judgment on when to rely on AI versus deep systems expertise
Mentor other engineers on systems design principles, debugging methodologies, and production operations, and contribute to onboarding programs for new team members
Lead incident retrospectives, identify root causes of complex production failures, and drive implementation of systemic improvements to prevent recurrence
Requirements:
Minimum Qualifications
8+ years of experience designing and implementing large-scale distributed systems, platform infrastructure, or systems software in production environments
Experience leading major technical initiatives end-to-end, including architecture design, cross-team coordination, staged rollout, and post-launch reliability ownership
Experience debugging complex, non-reproducible systems issues including concurrency bugs, memory management failures, and distributed consistency problems
Experience defining service level objectives, building observability infrastructure, and driving reliability improvements across production systems
Experience communicating technical architecture decisions and trade-offs in writing to both engineering and non-engineering stakeholders
Preferred Qualifications
Experience with systems programming languages such as C, C++, or Rust in the context of high-performance or low-latency infrastructure
Experience building or improving developer tooling, automation frameworks, or internal platforms that measurably improve engineering efficiency across teams
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8777005
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8781454
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 21 שעות
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced Senior Delivery Consultant - Modernization with deep expertise in Artificial Intelligence to join our Professional Services (ProServe). This role combines strategic architectural vision with hands-on technical leadership to deliver innovative AI solutions that drive customer success and business transformation across diverse industries and use cases.

Key job responsibilities
* Architecture & Design: Design and architect end-to-end AI-powered application solutions aligned with customer business objectives and technical requirements.
* Define application architecture patterns, standards, and best practices for AI/ML integration on us.
* Create technical roadmaps for customer AI application development and modernization initiatives
* Evaluate and recommend AWS AI/ML services and technologies including our Bedrock, SageMaker, and generative AI solutions
* Design data pipelines and ETL processes to support AI model training and inference using AWS services
* Customer Engagement & Consulting:
Lead customer engagements from discovery through implementation, serving as trusted technical advisor
* Conduct AI readiness assessments and develop adoption strategies tailored to customer maturity levels
* Facilitate architecture workshops and design sessions with customer stakeholders
* Deliver Well-Architected reviews focused on AI/ML workloads
* Build strong relationships with customer technical teams and executive leadership
* Guide customers in constructing AI processes aligned with AWS best practices
* Technical Leadership: Lead cross-functional teams in implementing AI solutions from concept to production
* Provide technical guidance on AI model integration, deployment strategies, and optimization on AWS
* Conduct architecture reviews ensuring solutions meet scalability, performance, security, and cost-efficiency requirements
* Mentor customer teams and junior ProServe consultants on AI best practices and AWS technologies
* Collaborate with data scientists, ML engineers, and software developers to translate AI models into production applications
* AI Solution Development: Design architectures for generative AI applications including RAG (Retrieval-Augmented Generation) systems, chatbots, and intelligent agents using Amazon Bedrock
* Architect real-time and batch AI inference pipelines with appropriate monitoring and observability
* Implement MLOps practices using SageMaker for model versioning, deployment automation, and continuous improvement
* Design solutions for responsible AI including bias detection, explainability, and governance frameworks
* Optimize AI application performance, cost, and resource utilization across AWS services
Knowledge Sharing & Thought Leadership
* Develop reusable assets, reference architectures, and best practice documentation
* Contribute to AWS ProServe knowledge base and customer-facing content
דרישות:
Basic Qualifications
- 10+ years of software development experience.
- 5+ years of machine learning, statistical modeling, data mining, and analytics techniques experience.
- Knowledge of AWS services including compute, storage, networking, security, databases, machine learning, and serverless technologies.
- Knowledge of programming languages such as C/C++, Python, Java or Perl.
- Master's degree in computer science, engineering, mathematics or equivalent, or experience in defining and creating benchmarks for assessing GenAI model performance.
- Understanding of various AI domains: NLP, computer vision, recommendation systems, predictive analytics.
- Willingness to travel to customer sites as needed.

Preferred Qualifications
- Certified Machine Learning Specialty or AI Practitioner or Generative AI - Associate.
- Contributions to open-source AI projects or published research.
- Experience with responsible AI frameworks, governance practices, and compliance requirements.
- Prior experience in ProServe, consulting, or systems integrati המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8802053
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior SRE Engineer, you will be a key player in ensuring the reliability, scalability, and performance of our critical IT infrastructure. You will leverage SRE principles and an automation-first mindset to build and maintain resilient hybrid cloud environments. This role is ideal for a candidate who thrives in a fast-paced, innovative setting and is passionate about solving complex challenges with cutting-edge technology.
Key Responsibilities
Provision, configure, and support resilient hybrid cloud deployment architectures using an Infrastructure-as-Code framework.
Proactively collaborate with development teams to ensure new applications are production-ready, scalable, and reliable from inception.
Develop and maintain tools and frameworks to automate operational tasks, including deployment, monitoring, and recovery.
Conduct thorough root cause analysis of production issues and implement preventative measures to improve system resilience, demonstrating strong problem-solving skills.
Manage CI/CD platforms, Linux infrastructure, and contribute to capacity planning and operational runbooks.
Design and implement proactive service monitoring, alerting, and trend analysis to maintain service availability and performance SLAs.
Participate in an on-call rotation to support critical applications and services, responding to and resolving incidents efficiently.
Contribute to comprehensive documentation related to infrastructure design, deployment, and operational procedures.
Requirements:
Your Expereience:
6+ years of Devops engineering experience on mission-critical, enterprise-level systems in a hybrid (both cloud and on-prem) environment.
3+ years of hands-on experience with cloud environments, preferably Google Cloud Platform (GCP).
Expertise in configuration management and Infrastructure-as-Code using frameworks such as Terraform and Ansible.
Strong programming/scripting knowledge in languages like Python, Bash, or Go for infrastructure automation.
Demonstrated experience with CI/CD pipelines (e.g., GitHub, Jenkins, Artifactory) and a strong foundation in Linux/Unix administration.
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
Preferred Qualifications
Experience with containerization and orchestration technologies, particularly Kubernetes.
Hands-on experience with monitoring and observability tools such as Datadog, Grafana, or Prometheus.
Understanding of networking principles including firewalls, load balancers, and complex network designs.
A curious and positive mindset with a passion for applied learning and challenging existing processes for continuous improvement.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779502
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
09/08/2026
Job Type: Full Time
We are looking for an outstanding Senior Networking Software Architect to join the NIC/DPU Software and Firmware Architecture group. In this role, you will help define the next generation of our datacenter and AI networking platforms, with focus on DPU management, QoS, performance, telemetry, and software architecture across stacks. You will work closely with hardware designers, firmware/kernel driver teams, system engineers, validation, product management and customers. The role spans early architecture definition, pre-silicon design, bring-up, and production readiness for large-scale AI and cloud datacenter deployments.

What Youll Be Doing:

Own software and system architecture for next-generation DPU management, QoS, performance, telemetry, and observability features.

Define end-to-end control and management flows across DOCA, host drivers, embedded firmware, BMC, management controllers and external management systems.

Specify telemetry and observability requirements, including counters, logs, traces, events, health monitoring, debug data, and streaming telemetry.

Define management interfaces and APIs for configuration, provisioning, lifecycle operations, diagnostics, and field serviceability.

Write clear architecture specifications, interface definitions, flow diagrams, and design documents for software, firmware, and system teams.

Partner with R&D teams to translate high-level architecture into implementable designs and guide features through development, validation, silicon bring-up, and production.

Analyze system performance bottlenecks, interoperability issues, telemetry gaps, and customer-reported issues, then feed learnings into future architecture.

Collaborate with system and cluster architects to ensure NIC/DPU features fit end-to-end AI datacenter and cloud networking designs.
Requirements:
What We Need To See:

B.Sc. or M.Sc. in Computer Engineering, Computer Science, Electrical Engineering, or equivalent experience.

9+ years of experience in networking, system software, embedded software, firmware, or datacenter infrastructure.

Proven experience in software architecture, or technical leadership roles.

Deep understanding of networking concepts and protocols such as Ethernet, TCP/IP, RDMA/RoCE, congestion control, QoS, virtualization overlays, and traffic management.

Strong background with DPUs, SmartNICs, or other high-performance networking devices.

Experience with system management, provisioning, monitoring, telemetry, diagnostics, or lifecycle-management flows.

Familiarity with management protocols and frameworks such as Redfish, PLDM, MCTP, IPMI, gNMI, SNMP, Netconf, REST, or gRPC-based APIs.

Ability to lead cross-functional architecture discussions across software, firmware, hardware, validation, product, and customer-facing teams.

Excellent written and verbal communication skills, including the ability to create clear architecture documents and present trade-offs.


Ways To Stand Out From The Crowd:

Experience defining software architecture for DPU products, including management, telemetry, QoS, performance, security, virtualization, or offload features.

Hands-on background with Linux networking, device drivers, firmware, embedded Linux, BMC software, DOCA, DPDK, OVS or Kubernetes networking.

Experience with performance counters, profiling tools, eBPF, Prometheus, Grafana, dashboards, heat maps, or large-scale telemetry systems.

Experience in defining and developing GAI-based analysis tools to extract insights from telemetry data and streams.

Background in RAS, diagnosability, serviceability, field failure analysis, production debug, or customer escalation handling.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8773378
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a detail-oriented and collaborative Senior ML Engineer to help support and maintain our machine learning capabilities. This role is ideal for someone who enjoys working closely with production systems, ensuring reliability, scalability, and explainability of models while enabling research teams to deliver impact faster.



Responsibilities



Collaborate with cross-functional teams to ensure ML systems remain robust, explainable, and aligned with business needs.
Monitor and report on ML model performance, reliability, and explainability metrics.
Participate in model retraining procedures, implement automation and optimization of MLOps pipelines.
Extend and scale monitoring pipelines, including support for new features in development.
Investigate, troubleshoot, and resolve issues in production ML workflows (tiered support from initial triage to root-cause analysis with model owners).
Develop and maintain repositories for feature engineering, inference monitoring pipelines, and artifact monitoring tools.
Perform exploratory data analysis (EDA) on historical datasets to identify quality issues and maintain data health.
Implement and oversee production based adjusters across customer deployments.
Evaluate and track critical ML artifacts such as explainability files, coverage metrics, and alignment of features.
Support development and maintenance of internal tools (e.g., interfaces, registries, and feature monitoring frameworks).
Build and maintain static and temporal features, including seasonality, event-based, and price-related features.
Requirements:
5+ years of hands-on experience in data science, ML operations, or applied ML support.
Proficiency in Python and standard data/ML libraries (Pandas/Polars, NumPy, Scikit-learn, SQL; experience with PyTorch or TensorFlow is a plus).
Strong data visualization and exploratory data analysis skills for monitoring and debugging pipelines.
Experience with time-series data and feature engineering.
Familiarity with explainability tools and model monitoring best practices.
Strong problem-solving skills with the ability to troubleshoot across data, code, and model workflows.
Excellent communication skills to summarize findings for both technical and non-technical audiences.
Experience with cloud-based ML platforms - preferably GCP
Familiarity with containerization (Docker), K8s, CI/CD workflows, or ML observability tools.
Familiarity with orchestration tools such as Airflow, Kedro or Dagster is a plus.
Prior exposure to demand forecasting, pricing, or revenue management.
Bachelor's or Master's in Computer Science, Machine Learning, Statistics, Engineering or a relevant field.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8762148
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
7 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We are a well-funded, early-stage startup looking for a talented and motivated Backend Engineer specializing in infrastructure to join our founding team. The focus of this role is to build and scale the infrastructure that powers autonomous AI agents automating complex enterprise workflows. You will own the systems, pipelines, and platforms that let our AI agents run reliably, securely, and at scale in production.

Your Impact
Infrastructure & Platform

Design, build, and own the core infrastructure powering our AI agent platform, from data pipelines to production deployment systems.

Build and scale the backend systems that support high-throughput document processing and data extraction workloads.

Cloud Infrastructure and Scalability

Architect and deploy infrastructure on cloud platforms (AWS, GCP, or Azure) with a focus on scalability, reliability, and cost efficiency.

Own containerization and orchestration (Docker, Kubernetes) for all production workloads.

Build and maintain CI/CD pipelines and DevOps practices that let the team ship fast without breaking things.

Data Infrastructure

Design and manage data pipelines to process and analyze large volumes of documents and unstructured data at scale.

Build the infrastructure layer connecting AI agents to databases, vector stores, and enterprise systems (ERP, CRM).

API & Systems Integration

Build and maintain robust, well-documented APIs connecting AI agents with external systems and enterprise software.

Design for reliability: retries, observability, and graceful degradation across distributed systems.

Security and Compliance

Implement authentication and authorization mechanisms (OAuth2, JWT) to secure AI-driven systems.

Ensure compliance with data privacy standards (e.g. GDPR, HIPAA) and drive best practices for secure data handling across the infrastructure.

Monitoring and Optimization

Build observability and monitoring systems to track infrastructure health, performance, and cost.

Continuously optimize system performance for speed, reliability, and cost-efficiency at scale.

Collaboration

Work closely with AI/ML engineers, product, and the founding team to make sure infrastructure decisions support fast iteration and production-grade reliability.

Participate in code reviews, design discussions, and architecture planning to drive infrastructure strategy.
Requirements:
5+ years of experience in backend or infrastructure engineering, ideally supporting production AI/ML systems or high-throughput data pipelines.

Proven track record of building and scaling infrastructure in production environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8793026
סגור
שירות זה פתוח ללקוחות VIP בלבד