Skip to main content

Director Engineering, Trainium Systems

Amazon Web Services (AWS) is redefining the future of artificial intelligence infrastructure. At the heart of this transformation is Trainium — AWS's family of custom-designed AI accelerator chips and server systems that power some of the world's most demanding machine learning workloads. In 2025, the Trainium Systems organization delivered over one million custom AI chips into production, launched Trainium3 at AWS re:Invent, delivering over 4.4x the compute performance of its predecessor, and activated Project Rainier, one of the world's most powerful AI compute clusters. Built in partnership with Anthropic, Project Rainier features nearly more than one million Trainium2 chips deployed across a dedicated campus in Indiana, delivering more than ten times the compute power used to train Anthropic's previous- AI models. With billions of dollars in annualized revenue and investment, and a product roadmap spanning multiple silicon , Trainium is one of the fastest growing and most strategically important programs at Amazon.

We are seeking a Director of Trainium Servers and Systems to lead all aspects of server hardware and firmware delivery and operations for this rapidly scaling portfolio. This role owns the end-to-end lifecycle of Trainium server products — from baseboard and accelerator card design, through chip-to-chip interconnect and rack-level architecture, to system firmware, fleet operations, and datacenter deployment. The scope spans multiple concurrent product lines including air-cooled and liquid-cooled UltraServers housing up to 144 accelerator chips, high-speed NeuronLink switched fabrics, and next- inference-optimized systems — all designed to compete directly with the industry's leading AI accelerator platforms.

This leader will manage teams of engineers across hardware, firmware, software, and systems development, fleet operations, and new product introduction, with a global footprint spanning Austin, San Jose, Seattle, and Taiwan, plus indirect oversight of 1,000+ engineers across design and manufacturing partners worldwide. The organization is responsible for driving reduction in time-to-production, from months to weeks, while simultaneously scaling server deliveries to thousands per week while maintaining availability targets across millions of chips.

This is a uniquely high-impact role at the intersection of cutting-edge hardware engineering and unprecedented operational scale. You will shape the systems that power AWS's custom silicon strategy — including massive UltraClusters like Project Rainier — enable frontier AI model training and inference for customers like Anthropic, OpenAI, and Amazon Bedrock, and help position AWS Trainium among the top three most widely used AI accelerators globally. If you are passionate about building world-class hardware at massive scale, thrive in fast-moving environments with multiple parallel product launches, and want to directly influence the trajectory of AI infrastructure, we'd love to talk.

Key job responsibilities
Deliver Multi- Trainium Server Products at Unprecedented Speed and Scale

Lead the simultaneous execution of multiple concurrent Trainium product lines while reducing chip-to-sellable cycle time, with delivery of millions of chips annually. Success means AWS maintains its competitive position against Google TPU and NVIDIA Blackwell/Rubin while meeting customer commitments.

Achieve Fleet Operations Excellence and Industry-Leading Sellable Rates

Increase sellable server rates across the Trainium fleet by building world-class fleet operations, test, and repair capabilities. This includes eliminating the current backlog of thousands of unhealthy servers through AI-powered visual inspection, automated diagnostics, and root-cause analysis of subtle connector damage modes — while simultaneously establishing the operational playbooks, tooling, and telemetry infrastructure needed to maintain these rates as the fleet scales to millions of chips across multiple server and cooling architectures. Every percentage point of sellable rate improvement at this scale translates to hundreds of millions of dollars in recovered CapEx.

Build and Scale a World-Class Hardware Engineering Organization Across Global Sites

Rapidly grow the organization to deliver on the rapidly scaling Trainium portfolio, integrating diverse talent pools across Austin, San Jose, Seattle, and Taiwan. Establish deep technical bench strength in emerging disciplines critical to future product , including liquid cooling operations, high-speed interconnects (PCIe/UAL/NVL), signal integrity, and near-package optics. Build succession depth among senior leaders, fostering a unified engineering culture across a team that is quickly multiplying. This objective is foundational: the organization's ability to execute on every other strategic priority depends on having the right leaders and engineers in place, operating with high trust and clear accountability across a complex, globally distributed hardware development and operations charter.