In this role, you will be responsible for several teams of architects, engineers and software developers, all working together to conduct state-of-the-art R&D in system and network architecture. As the group lead, you will guide and mentor the individual team leads, and also conduct hands-on work leading architecture, technology innovation and technical planning and of high-performance computing cluster network, which oriented at AI, HPC, and big data.
Responsibilities
You will perform a wide range of duties including:
Architecture Innovation:
Deeply analyzing the advantages and disadvantages of mainstream network systems, to find opportunities for network architecture innovation;
Insight into the technology developing trend of the high-performance computing network field, and leading the corresponding technology planning.
Exploring new architectures of high-performance computing network systems and efficiently integrating communication library, topology, and network protocol to solve performance bottlenecks.
Technical breakthroughs in networking and cluster routing algorithm:
Analyzes computing cluster network performance and leads the development of computing cluster network technologies
Research and optimize the heterogeneous interconnection topology of key computing chips to continuously improve the key competitiveness of Huawei computing heterogeneous chipsets.
Responsible for the research of data center network technologies, and guide network topology design and routing algorithm development
Group leadership:
Lead the development of a comprehensive system architecture for AI Fabric and HPC Fabric solutions.
Manage and mentor highly skilled team leaders, to ensure that the group operates together in pursuit of common goal.
Foster a collaborative and innovative work environment.
Provide technical guidance and support to team members.
Collaborate closely with cross-functional teams internationally, including hardware, software, and ucode design teams, to ensure alignment of architectural decisions with product and platform common objectives.
Initiate and supervise collaborations with top academic researchers in Israel and abroad.
Stay up to date with emerging technologies and industry trends in AI, HPC and big data industries.
Evaluate and recommend technologies and next generation projects.
Occasional travel related to ongoing projects, seminars, conferences etc.
Requirements: Requirements:
At least 10 years of hands-on experience in system architecture design, or equivalent research experience.
Demonstrated experience in leading R&D team.
Familiarity with high-performance computing cluster services and system architectures, such as AI, HPC and big data.
In-depth understanding of computer networks, communication libraries, and design of AI or HPC cluster networks
Key qualifications you hold include:
Ability to work in a team environment; actively seek out resolution to issues with the team and work co-operatively to build a coherent system; Team-working and excellent inter-personal communication skills.
High level of self-reliance and an autonomous target-oriented work style, can do attitude, eager to learn new things and ready to think outside the box
Ability to work on a schedule, even for un-schedule-able issues such as inventiveness.
Experience with network system of AI, HPC or big data cluster.
Demonstrated capability in several items from the list below:
o Deep understanding of computer network protocol and network topology of computing cluster
o Familiarity with communication library, such as OpenMPI/NCCL
o Extensive experience in leading the network architecture design of computing cluster
o Experience with implementing routing algorithm
o Familiarity with network specifications of CPU/Network Processing Unit/ Neural Processing Unit/Switch chip
Fluent in written and spoken English.
This position is open to all candidates.