Required ML Hardware Achitect
About the job
In this role, youll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to drive cutting-edge TPU (Tensor Processing Unit) technology that powers our most demanding AI/ML applications. Youll be part of a team that pushes boundaries, developing custom silicon solutions that power the future of our TPU. You'll contribute to the innovation behind products loved by millions worldwide, and leverage your design and verification expertise to verify complex digital designs, with a specific focus on TPU architecture and its integration within AI/ML-driven systems.
In this role, you will help shape the future of Clouds next-generation AI infrastructure, architecting high-performance Machine Learning silicon designed to power hyperscale AI inference. You will have an opportunity to drive accelerator technology that powers Generative AI models, large language models (LLMs), and emerging agentic workloads where throughput, latency, memory bandwidth, and energy efficiency are mission-critical.
You will be part of a silicon architecture team pushing the boundaries of custom computing. Leveraging your deep expertise in hardware-software co-design, machine learning algorithms, and computer architecture, you will define and optimize custom compute engines and memory hierarchies that accelerate the world's most advanced AI models across Cloud datacenters.
Responsibilities
Lead the architectural definition, modeling, and specification of next-generation, high-performance ML compute IP and acceleration blocks for Cloud AI silicon.
Own the ML IP architecture specification throughout the entire product lifecycle: concept exploration, cycle-accurate modeling, implementation, silicon bring-up, and production.
Partner closely with leading AI research and algorithm teams (e.g., Google DeepMind, Gemini research teams) and software compiler teams (XLA, PyTorch) to explore architectural trade-offs and define hardware requirements for emerging model architectures.
Drive comprehensive architecture studies, evaluating compute dataflows, numerical formats, sparsity, and specialized acceleration mechanisms such as key-value (KV) cache optimization.
Drive performance, latency, power efficiency, and silicon area projections across model topologies and workload configurations.
Requirements: Minimum qualifications:
Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, a related field, or equivalent practical experience.
15 years of experience in computer architecture, ML accelerator design, or high-performance processor architecture.
Experience leading architectural definition and authoring architecture specifications for silicon or compute IP blocks.
Experience with performance modeling, workload profiling, and hardware-software co-design.
Preferred qualifications:
Master's degree or PhD in Electrical Engineering, Computer Engineering, or Computer Science with an emphasis on computer architecture or ML hardware systems.
5 years of experience leading the architectural definition and microarchitecture of AI/ML accelerators from concept through production.
Deep knowledge of modern deep learning workloads (Transformers, MoE, Diffusion, Generative AI inference) and their system bottlenecks (memory capacity, KV cache bandwidth, interconnect scaling).
Strong understanding of high-performance memory subsystems (custom SRAM architectures, high-bandwidth memory hierarchies, caching schemes).
Experience working with modern ML frameworks (PyTorch, JAX, TensorFlow) and ML compilers/runtimes (XLA, TVM, Triton).
This position is open to all candidates.