1,268,139open jobs
73,604companies
208,392added this week
Browse all
Salary
$144k – $194k per year
Location
In office (Seattle)
Seniority
Middle · 3+ years exp
Visa
H-1B filings in 12 months: 16,768 · for this role: 10,545 · green card filings: 56
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 6, 2026. First seen by Alion on Oct 6, 2026. Amazon scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Amazon is an American technology and retail conglomerate founded by Jeff Bezos in 1994 as an online bookstore and headquartered in Seattle, Washington. It operates the world's largest online marketplace together with a global logistics network, physical grocery stores and a third-party seller platform that accounts for most units sold. Amazon Web Services, launched in 2006, is the leading public cloud provider and generates the majority of the group's operating profit, while advertising, Prime Video, Alexa devices and Kuiper satellite broadband round out the business.

AWS Neuron is the complete software stack for AWS Inferentia and Trainium - AWS purpose-built accelerators for cloud-scale machine learning. Join the Machine Learning Inference Applications team to build the serving technology that lets customers run large-scale model inference fast and efficiently on Neuron chips.

As an engineer on this team, you'll work on core serving technologies within open-source frameworks such as vLLM and SGLang, optimizing model serving performance on Neuron and broadening the range of models we support out of the box. Your work directly accelerates how quickly new models are enabled, shipped, and delivered to customers.

Key job responsibilities

- Contribute core serving features to open-source inference frameworks such as vLLM and SGLang - implementing and upstreaming support for continuous batching, paged attention, quantization, and distributed inference on Neuron.

- Broaden the range of models supported out of the box, and build tooling and automation that shortens the path from a new model to a production-ready deployment.

- Improve model development and shipping velocity by reducing enablement time for new models and strengthening the test, benchmarking, and release workflows the team relies on.

- Collaborate with model development, performance, compiler, and runtime engineers to deliver end-to-end model performance - production-ready accuracy, scalability, and efficiency across a broad range of models and customer workloads.

- Apply strong engineering practices - code reviews, testing, and operational excellence - to ship reliable, high-performance inference that customers depend on.

About the team

Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.

Basic qualifications

- 3+ years of non-internship professional software development experience

- 3+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience

- 2+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience

- Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field

- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques

Preferred qualifications

- Experience in debugging, profiling, and implementing software engineering best practices in large-scale systems

- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware

- Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution

- Experience with vLLM, SGLang, TensorRT or similar platforms in production environments, or experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution

- Kernel development experience (e.g., CUDA, Triton)

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,268,139 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Seattle
≈ $124k – $237k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Phoenix
Python
Go
JavaScript
TypeScript
AI/ML
Prompt Engineering
AI Agents
LLM
RAG
Tokenization
Agentic Workflows
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Service Mesh
SOAP
Cybersecurity
HashiCorp Vault
PKI
Apply
$149k – $216k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • San Jose
Python
PowerShell
DevOps
Rest API
Azure DevOps
GitHub Actions
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Configuration Management
Bitbucket
GitHub
GitLab
IAM
Cybersecurity
Burp Suite
Snyk
OWASP ZAP
SonarQube
Checkmarx
Semgrep
CodeQL
NIST CSF
OWASP Top 10
OWASP ASVS
OWASP SAMM
Threat Modeling
SBOM
SLSA
Veracode
Fortify
Invicti
Mend
PKI
OWASP
Analytics
Tableau
Power BI
Management
Jira
ServiceNow
Apply
$101k – $188k per year • In office • 2+ years exp • Bachelor's Degree • Sunnyvale
C++
MATLAB
MATLAB
Simulink
Apply
≈ $134k – $238k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp • Overland Park
Python
TypeScript
Python
FastAPI
Asyncio
Pydantic
Databases
PostgreSQL
Firestore
AI/ML
Claude
Vertex AI
Prompt Engineering
Multimodal AI
Function Calling
AI Agents
Llama
Mistral
Gemini
LLM
RAG
Google ADK
OpenAI
A2A
GPT-4
Edge AI
Prompt Caching
Agentic Workflows
Multi-Agent Systems
Tool Use
Frontend
Tailwind CSS
React.js
React Router
DevOps
Rest API
Terraform
GCP
Azure
CI/CD
AWS
QA
Playwright
Pytest
Vitest
Apply
≈ $157k – $312k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Palo Alto
AI/ML
DeepSpeed
Multimodal AI
NLP
Transformers
LLM
Post-training
Pre-training
FSDP
Recommender Systems
Interpretability
Machine Learning
Apply
≈ $157k – $326k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Seattle
AI/ML
Fine-tuning
Machine Learning
DevOps
GCP
Apply
Superintendent 1 day ago
≈ $91k – $180k per year (Estimated) • In office • Full-Time • 5+ years exp • Seattle
Apply
≈ $89k – $172k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Seattle
Python
Java
SQL
C++
Management
Outlook
Apply
Area Sales Manager 1 day ago
≈ $209k – $442k per year (Estimated) • In office • 10+ years exp • Seattle
Management
Outlook
Apply
$116k – $163k per year • Remote (United States) • Full-Time • 7+ years exp • Seattle
AI/ML
AI Agents
DevOps
SLI/SLO/SLA
Management
Intercom
Apply
See all jobs
This is one of many
1,268,139 more open roles from verified company boards, updated every day.