דרושים » הנדסה » HPC Operations Engineer

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 2 שעות
חברה חסויה
Location: Yokne`am and Tel Aviv-Yafo
Job Type: Full Time
We are now looking for a HPC Operations Engineer to join our mission and continue improving our HPC infrastructure. A meaningful part of NVIDIAs strength is our unique and advanced development tools and environments that enable our incredible pace of innovation. We are looking for architects to help us evolve the way our private compute cloud is architected and optimized.

What youll be doing:
Troubleshoot incoming support requests in a large-scale HPC environment.
Contribute improvements to existing deployment automation, configuration management, observability, and operational monitoring and day to day operation through automation.
Ensure compute servers are running an accurate Operating System and configuration.
Address Complex Issues: Perform comprehensive issue resolution from bare metal to application level, ensuring system reliability and efficiency.
Collaborate with specialist teams to drive issues to closure.
Collaborate with domain experts to improve how our chip development process utilizes our infrastructure.
Directly contribute to the overall quality and improve time to market for our next generation chips.
Requirements:
What we need to see:
BS in Computer Science or similar degree or equivalent experience.
2+ years of experience Proficient in coordinating Centos/RHEL Linux distributions.
Understating of container technologies like Docker.
Proficiency in Python and UNIX scripting languages such as bash.
Excellent problem-solving skills, with the ability to analyze sophisticated systems, identify bottlenecks, and implement scalable solutions.
Excellent communication and teamwork skills, with the ability to work optimally with teams with varied strengths and individuals.
Proven understanding of cluster configuration managements tools such as Ansible.

Ways to stand out from the crowd:
Understanding of key Linux technologies such as NFS, automounter, LDAP, DNS, and TCP/IP networking in Red Hat Linux distribution flavors.
Familiarity with job scheduler administration (e.g. IBM Spectrum LSF or SLURM) and experience building/ operating large scale compute infrastructure.
Knowledge of the FlexLM license management system.
Proficiency in Perl for maintaining legacy automation scripts.
Familiarity with High-Speed Networking (InfiniBand, RDMA, RoCE etc.) and fast, distributed storage systems (Lustre, GPFS, etc.)
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837857
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: More than one
Job Type: Full Time
We are looking for a Technical Support Engineer dedicated to supporting Slurm for our customers. You will join a dedicated team of Slurm subject-matter guides, owning sophisticated support cases and helping customers run reliable, efficient, and highly scalable clusters. This role requires extensive production experience with Slurm and the ability to diagnose issues across the scheduler and the surrounding Linux, networking, storage, authentication, database, and GPU infrastructure.

What you'll be doing:

Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters.

Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.

Solve Slurm configuration and policy features, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.

Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required.

Isolate problems across Slurm and its surrounding dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.

Advise customers on Slurm configuration, upgrades, operational practices, managing system resources, and safe recovery from production incidents.

collaborate with engineering teams by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.

Develop guides, knowledge-base articles, diagnostic tools, and internal training that strengthen Slurm expertise across the support organization.
Requirements:
What we need to see:

BS.c degree in Computer Science, Engineering, or a related field, or equivalent experience.

5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments including business-critical outage incidents.

Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.

Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution.

In-depth Linux system-administration and solve experience, including systemd, cgroups, authentication, networking, and database-backed services.

Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources.

Strong analytical and research skills, showing proficiency in distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems.

Excellent written and verbal communication skills, including the ability to turn detailed technical findings into clear explanations and actionable recommendations.


Ways to stand out from the crowd:

Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs.

Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues.

Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.

Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity.

Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836131
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Tel Aviv-Yafo and Ra'anana
Job Type: Full Time
We are seeking a talented and driven Senior Software Verification Engineer to join our innovative team and tackle SW verification challenges in the domains of high-speed networking, virtualization, and security. You will play a key role in validating and testing complex software products that support Ethernet and InfiniBand protocols, delivering advanced networking, storage, and security services for cloud, compute, and AI workloads.

What Youll Be Doing:

Develop and Automate Testing: Design, implement, and maintain automated test scripts and frameworks (primarily in Python) to verify the correct functionality of our software products.

End-to-End Feature Ownership: Deep dive into feature sets, taking responsibility from test planning through to final implementation and full automation.

System & Integration Validation: Validate software functionality and performance through system-level and integration testing, utilizing Linux-based environments and virtualization tools.

Test Environment Management: Set up, maintain, and optimize test environments using Linux, Docker, virtual machines, and other modern tools.

Collaboration & Communication: Work closely with software, DevOps, architecture, and product teams to define test requirements, coordinate releases, and ensure high-quality product delivery.

Continuous Improvement: Drive design verification flows, contribute to methodology improvements, and leverage planning/tracking systems to manage release progress and build release indicators.

Defect Analysis: Analyze test results, file defects, and track issues to closure, ensuring robust and scalable solutions.
Requirements:
What We Need to See:

Bachelors/masters degree in computer science or computer engineering, or equivalent experience

5+ years of experience in software testing, QA automation, or software engineering.

Strong proficiency in Python and scripting for automation.

Solid experience with Linux-based environments, including system tools and command-line utilities.

Proven understanding of computer networking and modern Linux operating systems.

Familiarity with software testing, integration, and system validation practices.

Excellent problem-solving, critical thinking, and communication skills.

Ability to work independently, manage multiple tasks, and drive technical initiatives.

Great interpersonal skills, agility, and determination for success.

Fluent English; strong presentation and public speaking abilities.

Ways to Stand Out from the Crowd:

Deep technical know-how and familiarity with networking protocols or low-level system tools.

Experience with Docker, KVM, or other virtualization technologies.

Knowledge of CI/CD tools (e.g., Jenkins, GitLab CI) and test reporting tools (e.g., Allure, Grafana, Kibana).

Experience with large HW+SW systems and advanced Linux OS technologies.

Proficiency with GIT, Bash, and other scripting languages.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837161
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
1 ימים
Location: Yokne`am
Job Type: Full Time
We are looking for a Technical Lead to join our Network System Validation group and lead the validation of advanced networking solutions across complex AI cluster environments. This is a deeply hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence. You will develop validation methodologies and automation frameworks, while working hands-on on debugging, performance analysis, and cutting-edge AI networking technologies at scale. Join us to help push our networking technologies to their limits and shape how next-generation AI infrastructure is validated.

What youll be doing:
Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions
Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution
Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection
Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes
Requirements:
What we need to see:
B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.
12+ years of experience in networking, system validation, or related domains.
Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.
Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).
Strong scripting and automation experience using Python, Bash, and/or Ansible.
Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress.
Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.
Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.

Ways to stand out from the crowd:
Experience with large-scale clusters or distributed systems.
Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).
Background in performance analysis, Kubernetes, or cloud environments.
Background in chaos testing, fault injection, or simulation systems .
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8836095
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: Yokne`am
Job Type: Full Time
We are looking for a Technical Lead to join our Network System Validation group and lead the validation of advanced networking solutions across complex AI cluster environments. This is a deeply hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence. You will develop validation methodologies and automation frameworks, while working hands-on on debugging, performance analysis, and cutting-edge AI networking technologies at scale. Join us to help push our networking technologies to their limits and shape how next-generation AI infrastructure is validated.

What youll be doing:

Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions

Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.

Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution

Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities

Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection

Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations

Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes
Requirements:
What we need to see:

B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.

8+ years of experience in networking, system validation, or related domains.

Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.

Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).

Strong scripting and automation experience using Python, Bash, and/or Ansible.

Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress.

Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.

Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.

Ways to stand out from the crowd:

Experience with large-scale clusters or distributed systems.

Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).

Background in performance analysis, Kubernetes, or cloud environments.

Background in chaos testing, fault injection, or simulation systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837843
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
2 ימים
חברה חסויה
Location: Tel Aviv-Yafo and Ra'anana
Job Type: Full Time
We are looking for a creative and motivated Software Engineer to join our SONiC open source NOS team. Our team is responsible for the design and integration of SONiC across our Networking products. In this role, youll have the opportunity to work on next-generation, high-speed networking solutions for AI data centers, contribute to open-source projects, and collaborate with experts across us and the broader SONiC community.

Youll help lead the development of new SONiC networking features and chassis management capabilities across our Networking products. Were looking for someone who enjoys solving complex technical challenges, learning quickly, and collaborating in a dynamic and diverse team environment.

What youll be doing:

Design, integrate, and deliver new features as part of the SONiC release train across our Networking products.

Develop networking and chassis management capabilities for next-generation platforms.

Work in a Continuous Deployment environment with fast development and deployment cycles.

Collaborate with experienced teams across our Networking and with members of the broader SONiC open-source community.

Develop high-quality, production-ready code that is published and reviewed in leading open-source environments.

Take ownership of features from design and implementation through integration and delivery.
Requirements:
What we need to see:

B.Sc. in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.

3-5 years of software development experience.

Understanding of networking concepts and protocols.

Experience with C++ and Python

Experience using AI development tools or building AI agents.

Strong technical problem-solving skills and the ability to learn new technologies quickly.

A collaborative mindset with strong communication and interpersonal skills.

Ways to stand out from the crowd:

Hands-on experience with L2 and L3 networking protocols.

Experience with SONiC, SAI, or other open-source networking projects.

Experience with Rust.

Familiarity with Linux and shell scripting.

Experience working with Scrum or other Agile development methodologies.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8835703
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
2 ימים
Location: Ra'anana and Yokne`am
Job Type: Full Time
We are looking for an Engineering Manager to join our Network System Validation group. You will work on system validating advanced networking solutions across our complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. In this role you will lead the validation direction and engineering excellence of one of our technology validation teams. This is a management role for a technology leader who can own the technical roadmap and execution, mentoring a team of high-performance engineers, and push our network to its speed-of-light limits. This role combines the development of methodologies and automation tools with system validation, performance analysis, and investigation of cutting-edge AI networking technologies at scale.

What youll be doing:

Lead, mentor, and coach a team of software development and system validation engineers.

Review system and product requirements, design validation methodologies, develop comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions.

Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.

Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution.

Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection.

Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations.

Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes.

Foster a team culture centered on software quality, accountability, and technical excellence.
Requirements:
What we need to see:

B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.

8+ overall years of experience in networking, system validation, or related domains.

3+ years of experience leading software or system development team.

Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.

Strong scripting and automation experience using Python, Bash, and/or Ansible.

Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).

Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.

Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.


Ways to stand out from the crowd:

Experience with large-scale clusters or distributed systems.

Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).

Background in performance analysis, Kubernetes, or cloud environments.

Background in chaos testing, fault injection, or simulation systems.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8835840
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 2 שעות
חברה חסויה
Location: Tel Aviv-Yafo and Ra'anana
Job Type: Full Time
As a Senior DevOps Engineer, youll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIAs AI infrastructure.


What You'll Be Doing:
Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations.
Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints.
Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory.
Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity.
Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations.
Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows.
Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening.
Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes.
Requirements:
What We Need to See:
Bachelors degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices.
Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting.
Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible.
Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices.
Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting.
Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting.
Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment.


Ways To Stand Out From the Crowd:
Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost.
Familiarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage.
Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana.
Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers.
Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837920
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
2 ימים
חברה חסויה
Location: Ra'anana and Yokne`am
Job Type: Full Time
We are looking for a Senior Software Engineer to join our team developing the core technologies behind our SmartNIC platform. Your work will directly impact next-generation AI infrastructure, cloud networking, and high-performance data centers used by customers around the world.

As part of the team, you will contribute to the DOCA SDK (Our DOCA Software Framework), enabling developers to build hardware-accelerated, cloud-native networking and security applications. You will also collaborate on the Linux Foundation's DPDK project, helping deliver industry-leading packet processing performance and scalable networking solutions.

What you'll be doing:

Architect, design, and develop next-generation networking acceleration technologies.

Build and enhance core software components for programmable, high-performance pipeline execution.

Develop software abstractions and infrastructure that expose sophisticated platform capabilities in a scalable and maintainable way.

Optimize execution flow, data movement, and software performance across complex features and workloads.

Work closely with architects, customers, and engineering teams to understand requirements and deliver high-quality solutions.

Collaborate across the software stack, including applications, SDK and runtime layers, drivers, Linux kernel, firmware, and hardware teams.

Drive technical decisions and play a role in crafting the build of performant, maintainable software.
Requirements:
What we need to see:

B.Sc. in Computer Science, Software Engineering, or equivalent practical experience.

5+ years of strong C/C++ software development experience.

Experience with Linux development tools, debugging, and performance analysis.

Proven understanding of networking concepts and protocols such as Ethernet, TCP/IP, or similar.

Strong understanding of operating systems, computer architecture, and low-level systems software.

Experience building or debugging multi-layer software systems.

Strong analytical and problem-solving skills.


Ways to stand out from the crowd:

Experience developing high-performance or low-latency software.

Background in networking, packet processing, or data-plane software.

Experience designing developer-facing APIs, libraries, or software frameworks.

Experience with compiler, code-generation, or runtime infrastructure.

Contributions to open-source systems or networking projects.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8835856
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: More than one
Job Type: Full Time
We are looking for a talented Senior Software infrastructure and tools Engineer to join the Data Processing Unit (DPU) SW Group. As a part of the team, you will lead complicated integrations which combining new technologies, automating processes, bringing up full software stacks, developing system scripts for cutting edge networking technologies with the quality requirements of industry-leading customers. You will work closely with our SDK development and gain an understanding of our products and technologies.

What youll be doing:

Crafting efficiency and usability improvements across our proprietary products that will help streamline various release pipelines and processes across us.

Definition and development of SDKs to be used for DPU SW development.

POC of new technologies that require additional development/integration efforts.

Creating and maintaining build system for complex SW products.

Developing tools using our various proprietary and other cutting-edge technologies.

Collaborate with team members, Architects, design, QA teams, and verification.
Requirements:
What we need to see:

B.Sc. or equivalent experience in Computer Science, Computer/Software Engineering or related field.

5+ years work experience in a software development.

Strong programming skills in Python, Go, and Bash.

Strong understanding of Linux and networking.

Solid expertise in Linux build systems, encompassing RPM and Makefiles.

Experience with Docker, Ansible, and Jenkins pipelines.

Proven object-oriented programming skills and design patterns.

Motivated, responsive, and keen on process improvement.

Excellent Analytical, debugging, and problem-solving skills.

Ways to stand out from the crowd:

Experience with bootloaders.

Contribution to open-source projects.

Experience working with SoC.

Familiar with other SDK such as CUDA SDK.

Strong interest towards groundbreaking technologies and ability to take initiatives and drive them across multiple functional teams.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837802
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
לפני 3 שעות
Location: Tel Aviv-Yafo and Yokne`am
Job Type: Full Time
We are looking for a Senior Software Engineer to join the DOCA SDK Verification team. The DOCA SDK enables developers to rapidly create applications and services on top of our BlueField data processing units (DPUs), leveraging industry-standard APIs. With DOCA, developers can deliver breakthrough networking, security, and storage performance by harnessing the power of our DPUs.

What you will be doing:
As a Senior Software Engineer in the DOCA verification team, you will play a key role in designing and developing the verification infrastructure for the DOCA SDK. This infrastructure is a complex system that executes thousands of tests every night across multiple hardware platforms and configurations. It includes mechanisms for error handling and fault recovery, while test results are stored, analyzed using advanced tools, and presented through live dashboards and reports. Your expertise in building robust and efficient verification systems will be critical to ensuring the reliability and quality of our software.
Requirements:
What we need to see:

Bachelor's or Master's degree in Computer Science, Software Engineering, or a related field.

5+ years of experience as a software engineer.

Proficiency in programming languages such as Python, Java, or similar.

Deep understanding of software development methodologies and engineering best practices.

Excellent problem solving skills and the ability to address complex technical challenges.

Strong communication and collaboration skills, with the ability to work effectively in a team environment.

Proven track record of delivering high quality work on time and meeting project deadlines.

Ways to stand out from the crowd:

Expert level knowledge of the Python programming language.

Experience with Kubernetes networking.

Strong knowledge of the Linux operating system.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8837836
סגור
שירות זה פתוח ללקוחות VIP בלבד