Salary
≈ $29k – $80k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
Amazon is an American technology and retail conglomerate founded by Jeff Bezos in 1994 as an online bookstore and headquartered in Seattle, Washington. It operates the world's largest online marketplace together with a global logistics network, physical grocery stores and a third-party seller platform that accounts for most units sold. Amazon Web Services, launched in 2006, is the leading public cloud provider and generates the majority of the group's operating profit, while advertising, Prime Video, Alexa devices and Kuiper satellite broadband round out the business.
As an LLM Inference Engineer on our AI Platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observablethe last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps, and requires strong intuition for how model architecture, runtimes, and hardware interact.
Responsibilities:
- Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
- Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
- Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e. g., schedulers, KV cache, batching, memory).
- Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
- Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
- Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e. g., quantization, paging, kernel/runtimes improvements).
- Partner cross-functionally: Work with data science and product teams to translate business requirements into performance and availability SLOs.
Requirements:
- Experience deploying and operating LLM inference services in production.
- Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).
- Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
- Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
- Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.
- 3+ years of hands-on experience in performance optimization and systems programming for AI/ML workloads.
- Demonstrated ability to deliver measurable production improvements (e. g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
- Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Bengaluru
≈ $26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
≈ $26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
≈ $22k – $57k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Hyderabad
Python
TypeScript
JavaScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
Claude
Claude Code
LLM
Model Context Protocol
Frontend
Angular
Next.js
React.js
DevOps
AWS
AWS Lambda
CI/CD
Docker
GitHub Actions
Terraform
Amazon ECS
Amazon S3
GitHub
IAM
Apply
Инженер-разработчик C# / Unity (Middle)
2 days ago
$17k – $20k per year (net) • In office • Full-Time • Krasnoyarsk
C#
Python
DevOps
Docker
Git
Zabbix
GitHub
Apply
$100k – $500k per year • In office • Full-Time • 1+ year exp • Fort Collins
C++
Python
SystemVerilog
Cython
C++
CMake
Cython
PyBind11
AI/ML
ChatGPT
Claude
Copilot
Edge AI
Chips/EDA
Synopsys ZeBu
Apply
Sr. Product Software Engineer I - 484
5 days ago
≈ $37k – $76k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Bengaluru
PowerShell
Python
Databases
Databricks
Snowflake
DevOps
AIOps
AWS
AWS Lambda
CI/CD
Datadog
GitHub Actions
Jenkins
Prometheus
Terraform
Amazon CloudWatch
Amazon S3
GitHub
GitLab
IAM
Analytics
QlikSense
Apply
Sr. Engineering Manager
6 days ago
≈ $42k – $93k per year (Estimated) • In office • 3+ years exp • Bengaluru
Java
Java
Spring Boot
Spring Framework
AI/ML
AI Agents
Apply
React Native Developer
7 days ago
≈ $19k – $63k per year (Estimated) • In office • 5+ years exp • Pune
JavaScript
TypeScript
Java
Java
Gradle
Frontend
GraphQL
Redux
Redux Toolkit
React.js
Mobile
Adaptive UI
Bitrise
CocoaPods
Fastlane
Material Design
React Native
State Management
DevOps
CI/CD
Git
GitHub Actions
GitHub
QA
Detox
Jest
Apply
SDET II
8 days ago
≈ $15k – $45k per year (Estimated) • In office • 3+ years exp • Bengaluru
JavaScript
TypeScript
AI/ML
AI Agents
Claude
DevOps
CI/CD
Docker
Git
GitHub Actions
Jenkins
Kubernetes
Rest API
GitHub
QA
Cucumber
Playwright
Postman
Apply
Software Development Engineer in Test
8 days ago
≈ $15k – $45k per year (Estimated) • In office • 3+ years exp • Bengaluru
JavaScript
TypeScript
AI/ML
AI Agents
DevOps
CI/CD
Docker
Git
GitHub Actions
Jenkins
Kubernetes
Rest API
GitHub
QA
Cucumber
Playwright
Postman
Selenium
Apply
Custom Software Engineering Lead
2 hours ago
≈ $31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
Analog Layout Engineer
2 hours ago
≈ $31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
Application Support Analyst II
3 hours ago
≈ $16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
Senior Data Platform Engineer
3 hours ago
≈ $37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
≈ $41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
This is one of many
368,657 more open roles from verified company boards, updated every day.

