795,157open jobs
50,783companies
124,735added this week
Browse all
Salary
$193k – $262k per year
Location
In office (Cupertino)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 22, 2026.

Overview
Company
Impact
Profile match
Amazon is an American technology and retail conglomerate founded by Jeff Bezos in 1994 as an online bookstore and headquartered in Seattle, Washington. It operates the world's largest online marketplace together with a global logistics network, physical grocery stores and a third-party seller platform that accounts for most units sold. Amazon Web Services, launched in 2006, is the leading public cloud provider and generates the majority of the group's operating profit, while advertising, Prime Video, Alexa devices and Kuiper satellite broadband round out the business.

The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on AWS Trainium, Amazon's custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate with PyTorch and JAX, so customers can train frontier-scale models on Trainium without rewriting their stack.

The Distributed Training team enables the training of a wide range of models, from large-scale pretraining through post-training and reinforcement learning, on AWS's custom ML accelerators. As more customer workloads shift toward RLHF, PPO/GRPO, and other fine-tuning methods, we are building the distributed training infrastructure, parallelism techniques, numerics, and high-performance kernels that these methods depend on. As part of the broader Neuron organization, we work across frameworks, kernels, compiler, runtime, and collectives - a true hardware and software co-design in practice. We not only optimize current performance but also contribute to future architecture designs, since the gaps we characterize today become requirements for the next generation of Trainium.

We are looking for passionate technical leaders who can help us build and fine tune these distributed training solutions. This role offers a rare opportunity to work at the intersection of machine learning, high-performance computing, and distributed systems, where you will help shape the direction of AI acceleration technology.

Key job responsibilities

You will lead the effort to build distributed training and post-training support into PyTorch and JAX for Trainium accelerators. You will work across PyTorch and Neuron software stack with the Neuron compiler and runtime teams to enable and fine tune large-scale training, post training, and reinforcement learning workloads on the latest Trainium instances. You will own the parallelism strategies these models depend on, spanning data, tensor, pipeline, expert, and context parallelism, and apply reduced-precision formats where they measurably pay off. You will profile end to end to determine whether a workload is bound by compute, memory, collectives, or host overhead, then drive the fix to the layer that owns it, working with compiler, runtime, and collectives engineers to land it. You will translate the performance gaps you characterize into requirements that influence frameworks, and contribute upstream to the open source frameworks our customers train on.

About the team

Inclusive Team Culture

Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon’s culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust.

Work/Life Balance

Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.

Basic qualifications

- 5+ years of non-internship professional software development experience

- 5+ years of programming with at least one software programming language experience

- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience

- Experience as a mentor, tech lead or leading an engineering team

- Bachelor's degree or above in computer science or equivalent

- * Familiarity with LLM/transformer fundamentals such as attention mechanisms, autoregressive decoding, V-cache behavior, forms of parallelism

- * Experience in machine learning, data mining, information retrieval, statistics or natural language processing

Preferred qualifications

- Master's degree or above in computer science or equivalent

- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture

- * Hands-on SFT and/or RLHF/PPO/DPO experience

- * Experience with ML frameworks such as Pytorch/Jax, Distributed libraries and Frameworks, RL frameworks or End-to-end Model Training

- * Experience with performance engineering: workload profiling, characterization (compute bound, memory bound, network bound), and optimization

- * Contribution to open source projects

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 193,300.00 - 261,500.00 USD annually

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
795,157 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Cupertino
$150k – $190k per year • Hybrid • 7+ years exp • Bachelor's Degree • Chicago
Python
PowerShell
Bash
DevOps
Ansible
GCP
OpenShift
Azure DevOps
Prometheus
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Configuration Management
Amazon EKS
Google GKE
Azure AKS
Incident Management
Linux
Apply
$160k – $200k per year • Hybrid • Chicago
Python
TypeScript
SQL
AI/ML
LangGraph
LangChain
Claude
Claude Code
Model Context Protocol
Vertex AI
AI Agents
Google ADK
OpenAI
Anthropic
OpenAI Agents SDK
A2A
Structured Outputs
LLM Evaluation
LLM Guardrails
DevOps
Rest API
GCP
Azure
CI/CD
Apply
$155k – $215k per year • In office • 10+ years exp • New York
Python
Java
AI/ML
LangGraph
LangChain
MLFlow
Embeddings
Prompt Engineering
Function Calling
AI Agents
RAG
OpenAI
LLM Guardrails
Machine Learning
DevOps
Rest API
GCP
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Platform Engineering
Management
Agile
Apply
$135k – $220k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Los Angeles
Python
SQL
AI/ML
Scikit-learn
TensorFlow
PyTorch
Anomaly Detection
Time Series Forecasting
Apply
≈ $158k – $287k per year (Estimated) • In office • Full-Time • 2+ years exp • Boise
Databases
Apache Kafka
AI/ML
AI Agents
Machine Learning
DevOps
Terraform
Kubernetes
Platform Engineering
Unix
Apply
≈ $32k – $68k per year (Estimated) • Hybrid • Full-Time • 2+ years exp • Moscow
Python
Python
FastAPI
AI/ML
LangChain
Qwen
LlamaIndex
LoRA
Fine-tuning
Embeddings
RLHF
Multimodal AI
NLP
PEFT
QLoRA
Llama
Mistral
Transformers
PyTorch
LLM
RAG
Hallucination
OpenAI
Anthropic
Hugging Face
DPO
LLM Guardrails
Constitutional AI
DevOps
CI/CD
Docker
Kubernetes
Apply
In office
Python
SQL
AI/ML
TensorFlow
PyTorch
Machine Learning
Apply
≈ $22k – $60k per year (Estimated) • Hybrid • Full-Time • Moscow
Python
AI/ML
RLHF
GigaChat
LLM
SFT
Management
Telegram
Apply
≈ $39k – $105k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Madrid
Python
SQL
MATLAB
Databases
ClickHouse
AI/ML
Scikit-learn
TensorFlow
NumPy
PyTorch
Machine Learning
DevOps
GCP
Azure
AWS
Kubernetes
Wi-Fi
Analytics
Tableau
Power BI
ETL/ELT
Apply
≈ $22k – $61k per year (Estimated) • Remote (EAEU) • Full-Time • Moscow
Python
SQL
AI/ML
LangGraph
LangChain
Qwen
DeepSeek
LoRA
RLHF
Prompt Engineering
Function Calling
NLP
PEFT
Llama
Transformers
LLM
RAG
DPO
SFT
PPO
GRPO
Structured Outputs
Tool Use
Apply
$157k – $213k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • New York
Python
MATLAB
Management
Agile
Apply
$148k – $200k per year • Equity • In office • Full-Time • 5+ years exp • Seattle
AI/ML
AI Agents
Machine Learning
DevOps
AWS
Apply
$144k – $194k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
Java
C#
C++
Perl
AI/ML
Machine Learning
Apply
$192k – $260k per year • Equity • In office • Full-Time • 6+ years exp • Master's Degree • Sunnyvale
Python
Java
C++
AI/ML
RLHF
Reinforcement Learning
Multimodal AI
Speech Recognition
SFT
Post-training
Pre-training
Text-to-Speech
RLAIF
Reward Modeling
Machine Learning
Apply
$143k – $193k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Seattle
Python
Java
C++
AI/ML
Reinforcement Learning
Computer Vision
AI Agents
NLP
Machine Learning
Apply
≈ $194k – $396k per year (Estimated) • Equity • In office • 10+ years exp • Bachelor's Degree • Cupertino
AI/ML
Fine-tuning
Quantization
Knowledge Distillation
AI Agents
Transformers
TensorFlow
PyTorch
LLM
Federated Learning
SFT
Pre-training
Recommender Systems
Model Distillation
Machine Learning
Management
Agile
Apply
$165k – $224k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Cupertino
Java
C++
C++
CMake
TensorFlow C++
PyTorch C++
LLVM
AI/ML
DeepSeek
Stable Diffusion
JAX
Llama
TensorFlow
PyTorch
AWS Trainium
MLIR
Apache TVM
DevOps
AWS
Amazon EC2
GitHub
Amazon S3
Apply
≈ $169k – $306k per year (Estimated) • Equity • In office • 6+ years exp • Bachelor's Degree • Cupertino
Python
AI/ML
Reinforcement Learning
Computer Vision
PyTorch
Anomaly Detection
Image Segmentation
Machine Learning
Apply
≈ $167k – $304k per year (Estimated) • Equity • In office • 8+ years exp • Master's Degree • Cupertino
Python
Objective-C
AI/ML
TensorFlow
PyTorch
Machine Learning
Apply
$157k – $213k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Cupertino
Python
DevOps
AWS
Unix
Apply
See all jobs
This is one of many
795,157 more open roles from verified company boards, updated every day.