We are looking for an Engineering Manager to join our Network System Validation group. You will work on system validating advanced networking solutions across our complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. In this role you will lead the validation direction and engineering excellence of one of our technology validation teams. This is a management role for a technology leader who can own the technical roadmap and execution, mentoring a team of high-performance engineers, and push our network to its speed-of-light limits. This role combines the development of methodologies and automation tools with system validation, performance analysis, and investigation of cutting-edge AI networking technologies at scale.
What youll be doing:
Lead, mentor, and coach a team of software development and system validation engineers.
Review system and product requirements, design validation methodologies, develop comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions.
Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution.
Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection.
Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations.
Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes.
Foster a team culture centered on software quality, accountability, and technical excellence.
Requirements: What we need to see:
B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.
8+ overall years of experience in networking, system validation, or related domains.
3+ years of experience leading software or system development team.
Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.
Strong scripting and automation experience using Python, Bash, and/or Ansible.
Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).
Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.
Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.
Ways to stand out from the crowd:
Experience with large-scale clusters or distributed systems.
Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField).
Background in performance analysis, Kubernetes, or cloud environments.
Background in chaos testing, fault injection, or simulation systems.
This position is open to all candidates.