דרושים » הנדסה » Software Team Lead- AI Datacenter Networking

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Software Team Lead.
Job Summary
Develop and optimize high-performance software for AI datacenter networking, implementing NOS enhancements, smart NIC integration, and specialized routing features for AI workloads.
Key Responsibilities:
Implement network operating system features for AI-optimized switching scenarios
Integrate smart NIC solutions for end-to-end AI-optimized networking
Create monitoring and telemetry collection mechanisms for datapath analysis
Implement load balancing algorithms optimized for AI workload patterns
Debug and optimize datapath performance bottlenecks
Collaborate with QA teams on feature testing and validation
Requirements:
BSc degree in Computer Science or Engineering
2+ years of team leads.
5+ years of software development experience in networking or systems programming
Understanding of networking protocols and packet processing optimization
Preferred Qualifications
Experience with SONiC development, SAI implementation, or switch software
Rust
Knowledge of Linux kernel networking
Experience with smart NIC programming interfaces and SDK development
Familiarity with DPDK, RDMA, and high-performance networking libraries
Experience with AI/ML framework networking integration
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8763837
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time and Hybrid work
we are looking for a Software Engineer - AI Datacenter Networking.
Job Summary:
Develop and optimize high-performance software for AI datacenter networking, implementing NOS enhancements, smart NIC integration, and specialized routing features for AI workloads.
Key Responsibilities:
Implement network operating system features for AI-optimized switching scenarios
Integrate smart NIC solutions for end-to-end AI-optimized networking
Create monitoring and telemetry collection mechanisms for datapath analysis
Implement load balancing algorithms optimized for AI workload patterns
Debug and optimize datapath performance bottlenecks
Collaborate with QA teams on feature testing and validation
Requirements:
BSc degree in Computer Science or Engineering
5+ years of software development experience in networking or systems programming
Understanding of networking protocols and packet processing optimization
Preferred Qualifications:
Experience with SONiC development, SAI implementation, or switch software
Rust
Knowledge of Linux kernel networking
Experience with smart NIC programming interfaces and SDK development
Familiarity with DPDK, RDMA, and high-performance networking libraries
Experience with AI/ML framework networking integration
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8763871
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and Perplexity kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver meaningful solutions at scale, then come join us! This truly is a role at the forefront of AI/ML-you'll be working on features for the largest clusters, with the largest customers, for the largest AI models.

Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use our internal CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Requirements:
Basic Qualifications
- 5+ years of non-internship professional software development experience.
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- 3+ years as a mentor, tech lead or leading engineering teams.
- 3+years experience in SW/HW Co-Design.

Preferred Qualifications
- Bachelor's degree in computer science or equivalent.
- Experience creating automated dashboards and visualization (such as Grafana).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8711144
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and Perplexity kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use internal our CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Requirements:
Basic Qualifications
- 5+ years of non-internship professional software development experience.
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- 3+ years as a mentor, tech lead or leading engineering teams.
- 3+years experience in SW/HW Co-Design.

Preferred Qualifications
- Bachelor's degree in computer science or equivalent.
- Experience creating automated dashboards and visualization (such as Grafana).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8748498
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
6 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Technical Lead to drive the architectural direction and engineering excellence of this group. This is a senior, deeply hands-on role for a technology leader who can own the technical roadmap, mentor a team of elite engineers, and build the infrastructure that challenges platform to its theoretical limits.
What You'll Lead:
Define and own the technical architecture of the group's distributed testing and reliability platform - designing for massive scale, real-world workload simulation, and adversarial failure injection
Lead effort involving multiple engineers, setting technical standards, running architecture reviews, driving design decisions, and mentoring engineers to grow
Build the systems that orchestrate millions of concurrent IO operations, inject chaos at the infrastructure layer (latency, packet loss, hardware failures), and expose the hardest-to-find race conditions and consistency bugs
Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines
Drive observability and reliability engineering across the group - building telemetry pipelines that track P99 latency, jitter, and system health, turning quality into a quantitative discipline
Collaborate deeply with Core R&D, Storage Kernel, and Infrastructure teams - translating architectural knowledge into targeted reliability strategies
Establish engineering practices - design docs, production-grade code reviews, testing philosophy, and cross-team technical alignment
Requirements:
Strong software engineering background with 6+ years of hands-on Python development experience is required. The ability to read, debug, and reason about C++, Rust, or Go is a significant advantage
Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system behavior under stress
Background in one or more of: storage systems, networking (TCP/IP, RDMA), cloud infrastructure, database internals, or high-performance backend systems
Experience building large-scale infrastructure platforms, internal developer platforms, or reliability engineering systems
Leadership:
Proven track record leading complex technical initiatives from architecture through delivery
Experience mentoring and growing engineers - raising the technical bar of a team, not just directing work
Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed
Comfortable operating at both the strategic and hands-on level - you write code, review designs, and shape roadmaps
Previous experience in people management roles - Advantage
Mindset:
You approach quality through the lens of Site Reliability Engineering: you care about MTTD, observability, and building self-healing systems
You have a "hacker" instinct - you don't just find bugs; you find the architectural flaws that allowed them to exist
You are an early adopter of AI tools and excited about applying LLMs and generative AI to accelerate engineering velocity
Big Advantages
Experience with storage systems, file systems, or high-performance distributed environments
Background in chaos engineering, fault injection, or simulation systems
Familiarity with observability tooling and performance engineering at scale
Experience building testing or reliability platforms as first-class engineering products
Prior experience as a Team Lead in a high-growth infrastructure company
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8757543
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
01/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
we are looking for a Software Team Leader to shape the next revolution in electronics!
Responsibilities
Lead a team of software engineers delivering core software systems and backend services in Java and/or Python.
Provide both technical and managerial leadership through architecture guidance, design reviews, code reviews, coaching, mentoring and performance management.
Own the full agile software development lifecycle, from requirements analysis, design and implementation through testing, release and production support.
Drive the design and delivery of scalable, reliable software systems, including REST APIs, database-backed services and integrations with messaging or streaming platforms.
Plan, prioritize and coordinate team execution, including estimates, milestones, resource needs, delivery risks and dependencies.
Enforce high engineering standards for code quality, maintainability, documentation, testing and release readiness.
Monitor production environments, lead incident response when needed and ensure fast resolution of production issues.
Collaborate closely with product, QA, DevOps and other engineering teams to ensure high-quality, on-time delivery.
Grow and develop the team through hiring, onboarding, mentoring and continuous improvement of engineering practices.
Promote practical AI innovation by staying current with AI development tools and helping the team identify opportunities to improve productivity, product capabilities and competitive differentiation.
Requirements:
Bachelors degree in Computer Science or a related field.
7+ years of hands-on software development experience in Java and/or Python.
5+ years of experience leading software engineering teams.
Strong experience designing and building core, complex software systems.
Strong backend development experience, including REST APIs, service-oriented architecture and production-grade software systems.
Strong understanding of Spring and Spring Boot.
Experience working with relational databases and strong SQL/RDBMS knowledge.
Solid understanding of Java/Python development fundamentals, including streams, I/O, collections and functional programming concepts.
Experience with high code standards, including clean design, formatting, naming, documentation and maintainability.
Strong familiarity with open-source frameworks and standard software development workflows.
Proven leadership, communication, problem-solving and teamwork skills.
Preferred qualifications
Experience designing or integrating AI-driven software solutions.
Experience with JPA/Hibernate.
Experience with streaming or messaging platforms such as Kafka or RabbitMQ.
Experience with Docker and Kubernetes.
Experience with cloud platforms such as AWS, GCP or Azure.
Experience operating and troubleshooting production environments.
Experience improving engineering practices, development velocity, release quality or team productivity through modern tooling, including AI-assisted development tools.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8719359
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
You'll design and build features that shape how hundreds of millions of customers discover products. The problems here are genuinely novel - we're defining how LLMs integrate into real-time recommendation experiences, not applying established playbooks. You'll work across multiple technical teams, ship iteratively, and see your work in the hands of customers quickly. This team values experimentation, moves fast, and gives engineers real ownership over what they build.

We are looking for a Software Development Engineer with sound technical judgment and a bias for action who takes ownership of problems end to end, communicates clearly, and cares about operational excellence, not just launching features, but making sure they hold up at scale. Someone who naturally raises the bar for the team: mentoring junior developers, advocating for engineering best practices, and thinking beyond the immediate sprint.

Key job responsibilities
- Design, build, test, and operate features for a personalized recommendation system used by multiple teams and operating at our scale.
- Deliver end-to-end solutions with focus on maintainability, scalability, performance, and reliability.
- Collaborate with Product and Science to define experiences, run experiments, and iterate based on data.
- Define and implement measurement strategies including analytics events and experiment configurations to track engagement and retention.
- Navigate ambiguity and make sound technical decisions in a problem space where established patterns don't always apply.
Requirements:
Basic Qualifications
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field.
- 5+ years of non-internship professional software development experience.
- Experience programming with at least one modern language such as Java, C++, or C# including object-oriented design.
- Experience with full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations.
- Experience contributing to the architecture and design (architecture, design patterns, reliability and scaling) of new and current systems.

Preferred Qualifications
- Master's degree or equivalent.
- Experience including, building and maintaining data flows and pipelines
- Experience with A/B testing.
- Familiarity with AI/ML integration and generative AI applications.
- Experience with end-to-end SDLC ownership, including operations and on-call, monitoring/metrics, and incident response/RCA.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8710999
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time and Hybrid work
we are looking for a Software Engineer Infra.
Responsibilities:
- Design and develop the core infrastructure components that power our products across a large-scale distributed system.
- Own and evolve key infrastructure domains such as high-availability, data-management, security, and telemetry.
- Produce excellent software design, write clean and efficient code, and debug complex issues across the stack.
- Mentor and guide junior engineers, sharing knowledge and raising the technical bar of the team.
- Turn ambitious, out-of-the-box ideas into real, production-grade value for our customers.
Requirements:
- BSc in Computer Science or a related degree, or equivalent practical experience.
- 5+ years of experience working as a Software Engineer.
- Strong proficiency in C++ / C.
- Experience with Rust and Python.
- Fast and self-driven learner of new technologies and programming languages.
Soft Skills
- Genuine passion for software development and craftsmanship.
- Ability to mentor and guide junior engineers.
- Creative, out-of-the-box thinker who can also execute and deliver.
Nice to Have / Advantage
- Experience with networking concepts (IPv4/6, ARP, DAD, routing, neighbors, etc.) - significant advantage.
- Experience with Linux networking, including Netlink and Linux network interfaces - significant advantage.
- Experience developing distributed systems.
- Experience with Linux kernel development or kernel internals.
- Experience working in a Linux environment.
- Experience developing software infrastructures.
- Experience with computer networking software.
- Experience with high-scale systems and performance optimizations.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8763851
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a talented Software Developer to join our Cloud Firewall team.
In this role, you will be part of a skilled software development team responsible for designing and developing advanced network security solutions for cloud environments such as AWS, Microsoft Azure, and Google Cloud Platform.
You will work with multiple programming languages and technologies, contribute to core product capabilities, and take part in building scalable, high-quality security solutions for modern cloud infrastructure.
Key Responsibilities
Participate in the full software development lifecycle, from initial concept and design through development, testing, and release.
Design, develop, and maintain features for cloud-based network security solutions.
Write well-designed, testable, efficient, and maintainable code.
Actively use AI-assisted development tools to improve productivity, code quality, testing, troubleshooting, documentation, research, prototyping, and engineering workflows, while ensuring secure and responsible usage.
Ensure software designs and implementations comply with product requirements and technical specifications.
Troubleshoot, debug, and resolve complex technical issues.
Support continuous improvement by researching alternative technologies and presenting recommendations for architectural review.
Collaborate closely with team members, architects, QA, and other stakeholders.
Requirements:
B.Sc. in Computer Science or a related field (must)
1+ year of software development experience
Comfortable working in a Linux environment
Deep understanding of networking concepts, including the TCP/IP stack, OSI model, and routing
Proficiency in Python and Bash scripting
Good familiarity with C
Strong coding, design, debugging, and troubleshooting skills
Hands-on experience using AI-assisted development tools as part of daily engineering work
Strong communication skills and ability to work effectively as part of a team
Advantages
Knowledge of Linux OS internals, such as the kernel networking stack, device drivers, and core subsystems
Experience working on security gateways, such as firewalls, VPN gateways, or deep packet inspection components
Background in cyber security or network security
Strong analytical and problem-solving skills
Fast learner with the ability to quickly understand complex systems
Ability to work in a dynamic, multi-tasking environment.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8723143
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
25/06/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We are looking for an IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.

Preferred Qualifications
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8711182
סגור
שירות זה פתוח ללקוחות VIP בלבד