1,466,394open jobs
87,697companies
230,126added this week
Browse all
Salary
$213k – $288k per year
Location
In office (Santa Clara)
Seniority
Staff · 7+ years exp
Visa
H-1B filings in 12 months: 16,768 · green card filings: 56
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 10, 2026. Amazon scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Amazon is an American technology and retail conglomerate founded by Jeff Bezos in 1994 as an online bookstore and headquartered in Seattle, Washington. It operates the world's largest online marketplace together with a global logistics network, physical grocery stores and a third-party seller platform that accounts for most units sold. Amazon Web Services, launched in 2006, is the leading public cloud provider and generates the majority of the group's operating profit, while advertising, Prime Video, Alexa devices and Kuiper satellite broadband round out the business.

Interested in building the distributed systems that let customers customize foundation models at scale? The SageMaker Training and Model Customization team builds the services customers use to fine-tune and post-train models against their own accuracy, reliability, and cost targets. We are hiring a Software Development Manager to lead the team that owns those capabilities.

Model customization is the fastest-moving layer of the machine learning stack. Post-training has moved from supervised fine-tuning to reinforcement learning against verifiable rewards, and from single-turn tasks to agentic training where a model learns by acting in an environment over long trajectories. Each shift changes the shape of the workload. Reinforcement learning puts an inference engine inside the training loop. Agentic training adds environments and tool calls, and moves the bottleneck from one run to the next. Your team turns each new technique into a capability customers can use, without rebuilding the platform every time the research moves.

Underneath the techniques, this stays a hard distributed systems problem. Training jobs run across large accelerator fleets for days at a time, and a single node failure can stall the job. The team works in PyTorch and across frameworks including Verl, FSDP, Megatron and vLLM, and contributes upstream to the open source projects the platform depends on.

As the manager you own the team and its charter. You will hire and grow engineers, set technical direction with your senior engineers and product managers, and decide what the team does and does not build. You own the roadmap and defend its tradeoffs with leadership. You own the service in production, including on-call health, operational metrics, and the requests that come with sitting in the critical path of customer training workloads.

Key job responsibilities

- Own the team's charter and roadmap, and defend its tradeoffs with senior leadership.

- Hire, develop and retain engineers in a specialized field with a scarce talent pool.

- Set technical direction with your senior engineers, applied scientists and product managers.

- Decide what the team builds, and what it does not.

- Turn new post-training techniques into supported customer capabilities on a predictable schedule.

- Own the service in production, including on-call health, availability and operational metrics.

- Improve training throughput and cost per run across large accelerator fleets.

A day in the life

Your morning starts in a design review. The team is working out how to keep an inference engine and a training engine in the same loop for reinforcement learning without leaving accelerators idle, and the answer will shape the future of the platform. You deep dive into possible scenarios, because at this scale a node dies routinely and recovery has to be normal rather than exceptional.

Then a charter call. A new post-training technique is three months out of the research literature and two customers are already asking for it. With your senior engineers and product manager you decide whether it becomes a first-class capability, a recipe on top of what you already have, or a no. That decision is yours to make, and the industry does not wait for it.

You spend an hour in 1:1s, including a career conversation with an engineer who wants to own the parallelism work. After lunch you sit a hiring debrief and make the call. Later you meet the applied scientists to understand what they have validated, and you argue about sequencing.

You close the day writing a page for leadership: what you are not building this half, and why that is the right trade.

About the team

We build the managed services customers use to customize foundation models. A customer brings a task, a dataset, and a way to score a good answer. Our services run the post-training loop on their behalf. We own that entire path, from the customer-facing API down to the training session infrastructure and the accelerator fleet it runs on.

That makes this a foundational charter. Model customization is how a customer turns a general model into one that is theirs, and every customization experience on SageMaker is built on what we deliver. When we add a technique, it becomes available to every customer on the platform. When we make the training loop faster, every customer's job gets cheaper.

Basic qualifications

- 3+ years of engineering team management experience

- 7+ years of working directly within engineering teams experience

- 5+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience

- Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations

- Experience partnering with product or program management teams

- Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers

Preferred qualifications

- Experience managing teams, or experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution

- Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution, or experience with PyTorch, JIT compilation, and AOT tracing

- Experience with deep learning libraries such as PyTorch, TensorFlow, MxNet Research publications in computer vision, deep learning or machine learning at peer-reviewed workshops, conferences or journals

- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in debugging, profiling, and implementing software engineering best practices in large-scale systems

- Experience with training and deploying machine learning systems to solve large-scale optimizations

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Santa Clara - 212,700.00 - 287,700.00 USD annually

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,466,394 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Management
Similar stack
Same company
Santa Clara
≈ $72k – $173k per year (Estimated) • In office • 1+ year exp • High School Diploma • Shawnee
Management
Microsoft Office
Apply
≈ $93k – $190k per year (Estimated) • Remote (United States)
Apply
Team Lead Sr 1 hour ago
$58k – $72k per year • In office • Full-Time • 3+ years exp • High School Diploma • Golden
Apply
≈ $72k – $172k per year (Estimated) • Equity • In office • 5+ years exp • Bachelor's Degree • West Chester
SAS
DevOps
GCP
Apply
≈ $62k – $149k per year (Estimated) • Equity • In office • 2+ years exp • High School Diploma • Cincinnati
Apply
In office • Internship • San Jose
Python
SQL
Databases
PostgreSQL
Apache Kafka
AI/ML
Fine-tuning
AI Agents
Flink
Transformers
PyTorch
LLM Guardrails
Machine Learning
Apply
GTM Engineer 1 day ago
In office • Full-Time • San Francisco • New York
Python
TypeScript
SQL
AI/ML
Claude Code
AI Agents
Analytics
A/B Testing
Management
n8n
Zapier
Marketing
GA4
Apply
≈ $123k – $226k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Sunnyvale
Python
JavaScript
Java
Node JS
Databases
Apache Kafka
Google BigQuery
BigQuery
AI/ML
AI Agents
Kubeflow
Flink
TensorFlow
Frontend
GraphQL
React.js
DevOps
GCP
AWS
Analytics
A/B Testing
Apply
$200k – $275k per year • Equity • Remote (United States) • Full-Time • Bachelor's Degree
SQL
AI/ML
Copilot
Claude
PyTorch
Ignite
Analytics
Cognos
Microsoft Excel
Management
Outlook
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • India
Python
TypeScript
SQL
AI/ML
LangGraph
AutoGen
LangChain
Embeddings
Prompt Engineering
AI Agents
AgentOps
CrewAI
RAG
AWS Bedrock AgentCore
Machine Learning
DevOps
OpenTelemetry
AWS
Grafana
AWS Lambda
Amazon S3
Amazon CloudWatch
API Gateway
Management
Agile
Apply
$194k – $263k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Boston
DevOps
GCP
Azure
AWS
Apply
$80k per year • In office • Internship • Bachelor's Degree • Seattle
Python
Java
SQL
C++
Apply
$127k – $185k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Apply
$89k – $155k per year • Equity • In office • Full-Time • 4+ years exp • Austin
Management
Airtable
Apply
$242k – $328k per year • Equity • In office • Full-Time • 10+ years exp • New York
Management
Agile
Apply
≈ $219k – $415k per year (Estimated) • In office • Santa Clara
Apply
≈ $220k – $485k per year (Estimated) • In office • Santa Clara
AI/ML
Embeddings
Machine Learning
Apply
≈ $267k – $548k per year (Estimated) • In office • Santa Clara
AI/ML
Machine Learning
Apply
≈ $116k – $273k per year (Estimated) • In office • Santa Clara
Apply
≈ $228k – $504k per year (Estimated) • In office • Santa Clara
AI/ML
Multimodal AI
LLM
Apply
See all jobs
This is one of many
1,466,394 more open roles from verified company boards, updated every day.