Salary
≈ $32k – $80k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
eBay is a multinational e-commerce corporation that connects buyers and sellers in more than 190 markets worldwide thus enabling economic opportunity for individuals, entrepreneurs, businesses, and organizations.
As an LLM Inference Engineer on our AI Platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable, the last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps and requires strong intuition for how model architecture, runtimes, and hardware interact.
Responsibilities:
- Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
- Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
- Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e. g., schedulers, KV cache, batching, and memory).
- Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
- Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
- Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e. g., quantization, paging, and kernel/runtime improvements).
- Partner cross-functionally: Work with data science and product teams to translate business requirements into performance and availability SLOs.
Requirements:
- Experience deploying and operating LLM inference services in production.
- Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).
- Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
- Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
- Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.
- 4-5+ years of hands-on experience in performance optimization and systems programming for AI/ML workloads.
- Demonstrated ability to deliver measurable production improvements (e. g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
- Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Bengaluru
Gen AI Engineering and Scaled AI Transformation
10 hours ago
$145k – $218k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Mississauga
Python
Python
FastAPI
Databases
FAISS
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Anthropic
Claude
Embeddings
Fine-tuning
Gemini
Hallucination
Hugging Face
Human-in-the-Loop
Hybrid Search
LangChain
LangGraph
Llama
LlamaIndex
LLM
LLM Guardrails
OpenAI
Prompt Engineering
PyTorch
RAG
TensorFlow
Mobile
Clean Architecture
DevOps
Docker
Incident Management
Vector
Apply
Golang developer (Новосибирск)
11 hours ago
≈ $19k – $46k per year (Estimated) • Remote/Hybrid • Full-Time • Novosibirsk
Go
JavaScript
Databases
PostgreSQL
Redis
AI/ML
Claude
Claude Code
Cursor
Frontend
GraphQL
Vue.js
DevOps
Docker
Drone
Grafana
gRPC
Kibana
Kubernetes
Nginx
Prometheus
Apply
$16k per year (net) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Petrozavodsk
AI/ML
LLM
Apply
MLOps Engineer
11 hours ago
≈ $20k – $57k per year (Estimated) • In office • 2+ years exp • Ahmedabad
C++
Python
Databases
FAISS
Pinecone
Weaviate
AI/ML
Airflow
Computer Vision
Kubeflow
MLFlow
NLP
Ray
Ray Serve
TensorRT
Triton
Triton Inference Server
vLLM
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
Grafana
Jenkins
Kubernetes
Prometheus
Terraform
Vector
Apply
Full Stack Engineer
11 hours ago
≈ $47k – $102k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • Hyderabad • Gurgaon
Java
Java
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
PostgreSQL
Redis
AI/ML
AI Agents
LLM
OpenAI
RAG
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Platform Engineering
Apply
Full Stack Engineer
12 days ago
≈ $21k – $91k per year (Estimated) • In office • Bengaluru
JavaScript
Kotlin
Python
TypeScript
Java
Java
Spring Boot
AI/ML
AI Agents
Model Context Protocol
Frontend
Angular
React.js
DevOps
CI/CD
Kubernetes
Rest API
Tekton
Apply
Software Engineer - Kubernetes / GPU
18 days ago
≈ $30k – $124k per year (Estimated) • In office • Bengaluru
Java
Node JS
Python
JavaScript
AI/ML
CUDA Toolkit
Ray
InfiniBand
NCCL
DevOps
Gateway API
GitOps
Kubernetes
Apply
Software Engineer - Kubernetes
18 days ago
≈ $27k – $70k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Bengaluru
Java
Python
AI/ML
CUDA Toolkit
Ray
InfiniBand
NCCL
NVLink
DevOps
ArgoCD
etcd
Gateway API
GitOps
Grafana
Kubernetes
OpenTelemetry
Prometheus
Robotics
NVIDIA Omniverse
Apply
Senior Software Engineer - Hadoop
24 days ago
≈ $25k – $67k per year (Estimated) • In office • 5+ years exp • Bengaluru
Java
Scala
AI/ML
Hadoop
Apply
Data Platform Engineer
24 days ago
≈ $18k – $69k per year (Estimated) • In office • Bengaluru
Go
Java
Python
Databases
Apache Kafka
AI/ML
Flink
Spark
DevOps
CI/CD
Kubernetes
Apply
Custom Software Engineering Lead
1 hour ago
≈ $31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
Analog Layout Engineer
1 hour ago
≈ $31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
Application Support Analyst II
2 hours ago
≈ $16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
Senior Data Platform Engineer
3 hours ago
≈ $37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
≈ $41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
This is one of many
368,657 more open roles from verified company boards, updated every day.

