דרושים » תוכנה » Senior Software Engineer - ML Network Stack, ML Network Stack

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
25/06/2026
משרה זו סומנה ע"י המעסיק כלא אקטואלית יותר
מיקום המשרה: תל אביב יפו
סוג משרה: משרה מלאה
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and Perplexity kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use internal our CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Requirements:
Basic Qualifications
- 5+ years of non-internship professional software development experience.
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- 3+ years as a mentor, tech lead or leading engineering teams.
- 3+years experience in SW/HW Co-Design.

Preferred Qualifications
- Bachelor's degree in computer science or equivalent.
- Experience creating automated dashboards and visualization (such as Grafana).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8748498
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
3 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications:
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications:
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776998
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
5 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL FLOW team is looking for a Software Development Engineer to design and build automation, tooling, and monitoring systems for our next-generation ML accelerator servers. We build production software to validate, initialize, monitor, and qualify these servers - from first silicon through fleet-scale deployment. Our work spans hardware diagnostics, manufacturing test automation, CI/CD pipelines, operational dashboards, and data-driven fleet health monitoring.

Key job responsibilities
Design and develop software infrastructure - automation frameworks, deployment systems, and test orchestration platforms that run at scale across manufacturing and production environments.
Work cross-functionally with Hardware, Manufacturing, and EC2 teams to automate coordinated software delivery and qualification workflows.
Debug and root-cause hardware/software interaction failures using systematic data analysis and automation-assisted triage.
Build and own CI/CD pipelines end-to-end: from code commit through build, test, deploy, and production validation - driving fast, reliable software delivery for hardware teams.
Create data pipelines and analytics systems (ETL, aggregation, real-time reporting) that transform raw hardware test results into actionable engineering insights.
Develop monitoring dashboards, alerting systems, and data visualization tools for fleet health, yield tracking, and performance benchmarking.
Own features end-to-end: from design through implementation, testing, deployment, and operational excellence.
Requirements:
Basic Qualifications
- 3+ years of software development engineer or related occupational experience.
- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering or a related discipline or equivalent.
- Experience using Linux, demonstrating proficiency with associated tools or languages.
- Can work proactively and independently, meet deadlines, and deliver on projects and tasks.
- Knowledge of software engineering best practices across the development life cycle, including agile methodologies, coding standards, code reviews, source management, build processes, testing, and operations.

Preferred Qualifications
- Experience building monitoring dashboards and data visualization (Grafana, CloudWatch, QuickSight, or similar).
- Experience with data pipelines, ETL, or analytics (S3, Athena, Spark, or similar).
- Experience with systems programming languages (C, C++, Rust).
- Familiarity with computer architecture concepts (PCIe, memory hierarchy, power management).
- Advantage: experience with hardware bring-up, ASIC/FPGA validation, or manufacturing test development.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774254
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
5 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774268
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Backend Software Engineering, your job responsibilities will include:
Build and ship high-quality, production-grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high-quality code.
Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.
Contribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably.
Critically evaluate code (human or AI-generated) for correctness, quality, security, and performance.
Build new and exciting components in an ever-growing and evolving market technology to provide scale and efficiency.
Develop high-quality, production-ready code that can be used by millions of users of our cloud platform.
Make design decisions on the basis of performance, scalability, and future expansion.
Work in a Hybrid Engineering model and contribute to all phases of SDLC including design, implementation, code reviews, automation, and testing of the features.
Build efficient components/algorithms on a microservice multi-tenant SaaS cloud environment
Code review, mentoring junior engineers, and providing technical guidance to the team (depending on the seniority level).
Requirements:
5+ years of development experience as a software engineer.
Deep knowledge of object-oriented programming and other scripting languages: Java, Python, Scala C#, Go, Node.JS and C++.
Strong SQL skills and experience with relational and non-relational databases e.g. (Postgress/Trino/redshift/Mongo).
Experience with developing SAAS products over public cloud infrastructure - AWS/Azure/GCP.
Proven experience designing and developing distributed systems at scale.
Proficiency in queues, locks, scheduling, event-driven architecture, and workload distribution, along with a deep understanding of relational database and non-relational databases.
A demonstrated, genuine AI-first approach to engineering - using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty.
Experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows.
Advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-ready.
A deeper understanding of software development best practices and demonstrate leadership skills.
Degree or equivalent relevant experience required. Experience will be evaluated based on the core competencies for the role (e.g. extracurricular leadership roles, military experience, volunteer roles, work experience, etc.)
Desired Skills:
Technical expertise in Generative AI, particularly with RAG systems and Agentic workflows that use large language models.
Experience with Big-Data/ML and S3
Hands-on experience with Streaming technologies like Kafka
Experience with Elastic Search
Experience with Terraform, Kubernetes, Docker
Experience working in a high-paced and rapidly growing multinational organization.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8737818
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Software Engineer to join our core engineering team and help build the infrastructure behind one of the fastest-growing AI APIs. You'll work on systems that handle massive scale across a distributed microservices architecture running on AWS. You'll ship fast, take ownership of critical systems, and solve hard infrastructure problems as we grow.

This is a great role for an engineer who loves building ambitious systems from scratch, and wants to tackle the kind of scale and complexity challenges typically reserved for much larger companies.

What Youll Do

Design and build high-performance distributed systems

Design and implement backend infrastructure and API endpoints

Build and optimize real-time data pipelines that process billions of events per day across distributed queues and stream processors

Improve performance, monitoring, and reliability across the stack

Own core systems and contribute to key architectural decisions

Help shape a strong engineering culture focused on velocity and quality
Requirements:
What You Bring

5 years of professional software engineering experience

Strong backend development skills

Proven experience designing and operating large-scale, distributed systems, with a solid understanding of API design, reliability, and performance at scale

Hands-on expertise with AWS infrastructure and cloud-native services, bringing practical knowledge of deploying and managing services in real-world environments

Comfortable in a fast-paced startup environment with lots of ownership

Strong sense of ownership and accountability over outcomes

Curiosity about LLMs, retrieval and the future of AI systems, with a drive to stay at the forefront of new technology

Based in Tel-Aviv or open to relocating
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8761222
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
03/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We're looking for a Senior Data Engineer to help build next-generation data platform - the lakehouse foundation that will power data processing across the entire product. This is not a "write pipelines on top of someone else's platform" role, and it's not a pure infrastructure role either. It's both, deliberately.

You'll own the platform end to end: the infrastructure it runs on (Spark on Kubernetes, Apache Iceberg, AWS Glue, Airflow), the frameworks and tooling that let dozens of other engineers build on it without reinventing the wheel, and the design of the data pipelines themselves. Everything you build becomes leverage for the teams around you - your abstractions, base images, CI/CD flows, and operational patterns are what make the platform usable at scale.

You'll also own one of the hardest ongoing trade-offs in a high-scale data platform: balancing cost and performance. Compute sizing, storage layout, partitioning and compaction strategy, job scheduling - every decision has a price tag and a latency profile, and you'll be the one making those calls with data.

This role is ideal for an engineer who is equally comfortable debugging a Spark executor OOM on Kubernetes at 10am, designing a clean Python framework API at noon, and modeling the cost impact of a table layout change in the afternoon.



What You'll Do

Platform & Infrastructure

- Design, deploy, and operate our Spark-on-Kubernetes compute platform, including autoscaling, resource tuning, and multi-tenancy considerations.

- Own the lakehouse storage layer built on Apache Iceberg and AWS Glue catalog - table design, partitioning, compaction, schema evolution, and retention.

- Build and operate orchestration on Airflow: DAG standards, deployment flows, environment promotion, and reliability.

- Own production operations of the platform: monitoring, alerting, incident response, and continuous hardening.

Frameworks & Developer Enablement

- Build the code frameworks, libraries, and templates that other engineers use to write pipelines - so that spinning up a new production-grade Spark job is measured in hours, not weeks.

- Define and enforce standards for pipeline structure, testing, observability, and deployment across teams.

- Own CI/CD for data workloads: image builds, artifact promotion, and GitOps-based delivery.

- Act as a technical partner to product and research teams building on the platform - your customers are other engineers.

Data Pipelines & Architecture

- Design and build scalable batch and streaming pipelines processing complex, high-volume datasets from diverse sources.

- Lead large-scale backfills and migration initiatives, ensuring data consistency and integrity across evolving storage and compute platforms.

- Design event-driven data flows over large-scale queue systems (Kafka) for reliable, efficient data movement.

Cost & Performance

- Continuously balance cost against performance: right-size compute, tune queries and jobs, optimize storage layout and file sizes, and choose the correct engine for each workload.

- Build cost visibility and attribution into the platform so trade-offs are made with data, not guesswork.
דרישות:
- 5+ years of experience in software engineering, with meaningful time spent building and operating large-scale data platforms.

- Strong hands-on experience with distributed processing engines (Spark strongly preferred), including performance tuning and debugging in production.

- Practical experience deploying and operating workloads in Kubernetes-based environments - you're not afraid of infra work; you enjoy it.

- Experience building shared frameworks, libraries, or internal tooling used by other engineers, with the product mindset that comes with it (clean APIs, docs, versioning, backward compatibility).

- Strong proficiency in SQL and data modeling: complex analytical queries, query tuning, partitioning strategies.

- Solid software engineerin המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8766023
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
02/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Are you a results-driven backend or full-stack engineer passionate about building scalable, cloud-native microservices? Do you thrive in an agile, startup-like environment while having the backing of a market leader? At our company, you wont just be a coder; you'll be an architect of innovation, shaping the future of identity security.
Our engineering team is the core of our success. We build high-quality, professional-grade products on a mature, event-driven microservices architecture hosted in AWS. We're now seeking a Senior Backend Software Engineer to be a foundational member of a new team building a cutting-edge, cloud-based SaaS identity analytics product from the ground up.
Your Mission: The First 12 Months

First 3 Months: You will be fully integrated into our agile team, actively contributing to the design and development of core microservices for our new SaaS product. You will have a comprehensive understanding of the product architecture and roadmap, and you'll be writing high-quality, test-covered code in Go.

First 6 Months: You will take ownership of significant features, from design and estimation to implementation and deployment. You'll be expected to contribute to our continuous delivery pipeline and improve code quality by producing unit and end-to-end tests, aiming to increase code coverage by 15-20%.

First 12 Months: You will be a key contributor to the product's evolution, mentoring junior engineers and collaborating with product management to define and implement new features. You will have made a measurable impact on the product's performance and scalability, and you'll be a go-to expert for critical components of the system.

What You'll Do

Design, develop, and deploy backend microservices in Go. Success will be measured by the delivery of well-tested, scalable, and maintainable code that meets product requirements and is deployed to production in a timely manner.

Integrate AI-assisted tooling into day-to-day DevOps and engineering workflows. Success will be measured by improvements in productivity, scalability, and operational efficiency. You will achieve this by leveraging AI tools to generate initial configuration drafts, validate infrastructure code, and recommend workflow improvements, while utilizing AI-driven automation to reduce repetitive manual tasks and accelerate engineering execution without sacrificing high-quality standards.

Collaborate on designs, code reviews, and testing. Success will be measured by your active participation in team meetings, providing constructive feedback on peer code reviews, and contributing to a collaborative and positive team environment.

Produce unit and end-to-end tests. Success will be measured by your consistent contributions to our test suites, with a focus on increasing code coverage and improving the overall quality and reliability of the product.

Create and refine design and engineering best practices. Success will be measured by tangible improvements to our development processes and documentation, leading to increased team efficiency and code quality.
Requirements:
5+ years of professional software development experience.
3+ years of hands-on experience with Go.
A Bachelor of Science in Computer Science, a related field, or equivalent practical knowledge.
Advanced experience with object-oriented analysis and design.
Strong communication skills, with the ability to articulate complex technical concepts to both technical and non-technical audiences.

Preferred Qualifications:
Proven experience with AWS.
A deep understanding of Continuous Delivery principles and practices.
Prior experience working on Big Data or Machine Learning products.
Familiarity with instrumenting code for production performance metrics.
While this is a backend-focused role, a solid understanding of modern JavaScript frameworks (React, Angular, Backbone) and ES6+ is a plus.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8764428
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
04/08/2026
Location: More than one
Job Type: Full Time
We are seeking a highly motivated High-Performance System Architect to join our team of experts and help shape the future of high-performance and ML / AI computing. Our next-generation NVL systems will be at the forefront of connecting and powering the world's most advanced compute clusters, which would be used to train the most advanced AI models such as GPT and DeepSeek. As a high-performance system architect, you will have the opportunity to work on some of the most cutting-edge technology and help to drive the innovation of our next generation networks that will be used by top researchers and engineers around the world.

What youll be doing:

Define the NVL system architecture end-to-end, by internal requirements and customers requirements through all product life cycles (post/pre silicon, on deployments).

Research various of solutions to enable the next large-scale-high-performance computing clusters. The position spans over various layers from algorithms, software, firmware, and HW.

Collaborate with cross-functional teams, including other architecture teams, logic design, system software, firmware, and research teams, to ensure the successful execution of the project.
Requirements:
What we need to see:

B.Sc, M.Sc, or Ph.D degree in Computer Science, Computer Engineer, or Electrical Engineer.

At least 5 years of industry or research experience in computer networks.

Excellent understanding of large-scale networks behavior and the effect of distributed computing workloads effect on the network.

Experience in developing models for simulations, analyzing simulation results and development of optimization algorithms.

Possess strong managerial, problem solving and critical thinking skills.

Ability to work and operate in a highly dynamic environment.

Partner with multiple groups in the organization.


Ways to stand out from the crowd:

Good knowledge in network protocols - such as InfiniBand, IP, TCP and RoCE and network topologies.

Good knowledge in Python, C++.

Familiarity with HPC environments, routing algorithms, Omnet++ and NS3 simulation environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8768260
סגור
שירות זה פתוח ללקוחות VIP בלבד