812,549open jobs
52,293companies
130,976added this week
Browse all
Salary
$157k – $213k per year
Location
In office (Austin)
Seniority
Junior · 2+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
Amazon is an American technology and retail conglomerate founded by Jeff Bezos in 1994 as an online bookstore and headquartered in Seattle, Washington. It operates the world's largest online marketplace together with a global logistics network, physical grocery stores and a third-party seller platform that accounts for most units sold. Amazon Web Services, launched in 2006, is the leading public cloud provider and generates the majority of the group's operating profit, while advertising, Prime Video, Alexa devices and Kuiper satellite broadband round out the business.

AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms - from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.

What You Will Do

You will define the hardware that runs the world's largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.

Why You Will Love It

The world's most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust - you own platforms end to end from architecture definition through fleet operations.

The Ideal Candidate

You think across the full hardware stack - from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.

Key job responsibilities

Architecture & Design

* Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale

* Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs

* Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)

Validation & Bring-up

* Define and execute validation strategies from PCBA bring-up through server and rack integration - covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance

* Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems

* Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions

Fleet Quality & Continuous Improvement

* Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes

* Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms

* Partner with test and automation teams to improve manufacturing yield and reduce test dwell times

Cross-Team Collaboration

* Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs

* Drive ODM/JDM design partners through development milestones and production ramp

* Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready

May require occasional (<10/>

A day in the life

You start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal - correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.

About the team

The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end - from server conception through fleet-scale operations.

Basic qualifications

- Bachelor's degree in electrical engineering, computer engineering, or equivalent

- Experience in server technologies such as, thermal, mechanical, power, and signal integrity

- Experience in developing functional specifications, design verification plans and functional test procedures

- 2+ years of hardware design, development and validation experience for server or compute platforms

Preferred qualifications

- Master's degree in Electrical Engineering, Computer Engineering, or a related technical field

- Experience with analog, digital, and high-speed circuit design

- Experience working in data centers or critical infrastructure

- 2+ years of experience working with ODMs through product development and manufacturing lifecycle

- 2+ years experience working with hardware bring-up, debug, or validation of GPU/accelerator platforms

- Familiarity with PCIe topology, NVMe, and accelerator interconnects

- Experience developing and executing test procedures for electrical or mechanical systems

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 157,300.00 - 212,800.00 USD annually

USA, TX, Austin - 136,000.00 - 184,000.00 USD annually

USA, WA, Seattle - 136,000.00 - 184,000.00 USD annually

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
812,549 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Austin
$172k – $222k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Seattle
Python
Java
C++
C++
PyTorch C++
AI/ML
JAX
PyTorch
AWS Trainium
Machine Learning
DevOps
AWS
Apply
$143k – $193k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Austin
Python
Java
C++
AI/ML
AI Agents
Machine Learning
DevOps
Linux
Unix
Apply
$172k – $222k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Sunnyvale
Python
Java
C++
AI/ML
Multimodal AI
Computer Vision
TensorFlow
Synthetic Data
Machine Learning
Robotics
Sim-to-Real
Digital Twin
Apply
Research Intern 3 days ago
≈ $135k – $363k per year (Estimated) • In office • Internship • Palo Alto
AI/ML
Fine-tuning
Reinforcement Learning
SFT
LLM Evaluation
World Models
Apply
≈ $148k – $287k per year (Estimated) • Equity • Hybrid • Full-Time • 3+ years exp • Master's Degree • Menlo Park
Python
C++
AI/ML
Reinforcement Learning
Tool Use
Physical AI
Robotics
Isaac Sim
MuJoCo
Isaac Lab
Sim-to-Real
Imitation Learning
Reinforcement Learning
Diffusion Policy
Apply
≈ $50k – $100k per year (Estimated) • In office • Full-Time • 3+ years exp • Manchester
Python
Ruby
Perl
AI/ML
Prompt Engineering
NLP
DevOps
Rest API
CloudFormation
Azure
AWS
AWS Lambda
Amazon EC2
Amazon S3
IAM
Cybersecurity
GDPR
HIPAA
Marketing
Salesforce
Apply
$185k – $250k per year • Equity • In office • Full-Time • 8+ years exp • Seattle
DevOps
AWS
Amazon EC2
Amazon S3
Apply
≈ $25k – $53k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Go
SQL
MATLAB
SAS
Go
Chi
Databases
Oracle
DynamoDB
Amazon Redshift
DevOps
AWS
Amazon EC2
Amazon S3
Analytics
Tableau
ETL/ELT
Apply
≈ $90k – $214k per year (Estimated) • Equity • In office • Full-Time • Bachelor's Degree • Seattle
Python
Ruby
AI/ML
Machine Learning
DevOps
AWS
Amazon EC2
Analytics
Tableau
Cognos
Apply
$207k – $280k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
AI/ML
AI Agents
AWS Bedrock
DevOps
AWS
Apply
$143k – $193k per year • Equity • In office • Full-Time • PhD • Boston
AI/ML
Reinforcement Learning
Computer Vision
Physical AI
Machine Learning
Robotics
Sensor Fusion
Motion Planning
Reinforcement Learning
Apply
$143k – $193k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Seattle
Python
Java
C++
AI/ML
Reinforcement Learning
Function Calling
LLM
Post-training
Tool Use
DevOps
AWS
Linux
Unix
Apply
$157k – $213k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • New York
Python
MATLAB
Management
Agile
Apply
$148k – $200k per year • Equity • In office • Full-Time • 5+ years exp • Seattle
AI/ML
AI Agents
Machine Learning
DevOps
AWS
Apply
$144k – $194k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
Java
C#
C++
Perl
AI/ML
Machine Learning
Apply
$165k – $224k per year • Equity • In office • Full-Time • 2+ years exp • Austin
Python
Java
C++
Perl
SystemC
AI/ML
AWS Trainium
Machine Learning
DevOps
AWS
Windows
QA
Pytest
Apply
$140k – $210k per year • Equity 0–0.2% • Remote (United States) • Full-Time • 6+ years exp • Master's Degree • Austin
Python
JavaScript
SQL
Databases
DynamoDB
AI/ML
Cursor
Claude Code
Frontend
React.js
DevOps
Rest API
AWS
AWS Lambda
Amazon EC2
IAM
Amazon ECS
Apply
$215k per year • In office • Full-Time • 3+ years exp • Austin
Management
Slack
Marketing
Salesforce
Apply
$90k – $300k per year • Equity 0–0.2% • In office • Full-Time • 1+ year exp • Austin
Apply
≈ $97k – $257k per year (Estimated) • Equity 0–0.2% • In office • Full-Time • 1+ year exp • Austin
AI/ML
AI Agents
Cybersecurity
FedRAMP
Apply
See all jobs
This is one of many
812,549 more open roles from verified company boards, updated every day.