380,196open jobs
9,948companies
48,279added this week
Browse all
Salary
$27k – $74k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
eBay is a multinational e-commerce corporation that connects buyers and sellers in more than 190 markets worldwide thus enabling economic opportunity for individuals, entrepreneurs, businesses, and organizations.

As an LLM inference engineer on our AI platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable - the last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps and requires strong intuition for how model architecture, runtimes, and hardware interact.

Responsibilities:

  • Own production inference: take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimisation, and incident response.
  • Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
  • Optimise runtimes and servers: Scale inference across heterogeneous GPU fleets; optimise stacks such as vLLM, Triton, and related components (e. g., schedulers, KV cache, batching, and memory).
  • Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilisation, memory, and cost.
  • Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
  • Apply and ship new optimisations: Evaluate research and implement pragmatic inference optimisations (e. g., quantisation, paging, and kernel/runtime improvements).
  • Partner cross-functionally: Work with data science and product teams to translate business requirements into performance and availability SLOs.

Requirements:

  • Experience deploying and operating LLM inference services in production.
  • Strong production coding skills in Python, plus Go or Rust (systems-level implementation and debugging).
  • Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
  • Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
  • Solid understanding of LLM inference and optimisation techniques: continuous batching, KV cache management, quantisation, speculative decoding (nice-to-have), etc.
  • 3+ years' hands-on experience in performance optimisation and systems programming for AI/ML workloads.
  • Demonstrated ability to deliver measurable production improvements (e. g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
  • Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
380,196 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$27k – $67k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
Ruby
AI/ML
LLM
DevOps
AWS
Azure
GCP
Apply
$28k – $55k per year (Estimated) • Remote • Full-Time • 5+ years exp
Python
SQL
Python
pySpark
Databases
Snowflake
AI/ML
Claude
Claude Code
Cursor
dbt
Embeddings
LLM
Spark
Structured Outputs
DevOps
Amazon ECS
Amazon S3
AWS
AWS Lambda
Vector
Apply
$46k – $101k per year (Estimated) • In office • Bengaluru
Go
JavaScript
Python
Apply
$28k – $66k per year (Estimated) • Remote • 5+ years exp • Saint Petersburg
Go
Python
SQL
Python
Django
Databases
Apache Kafka
ElasticSearch
Kafka
PostgreSQL
RabbitMQ
Redis
AI/ML
Airflow
Claude
Claude Code
Copilot
Cursor
DevOps
CI/CD
Docker
GitLab
GitLab CI
Grafana
Kubernetes
Prometheus
Rest API
WebSockets
Yandex Cloud
Cybersecurity
Keycloak
Analytics
ETL/ELT
Management
Telegram
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Guangzhou
Python
QA
Appium
JMeter
Selenium
Apply
$32k – $80k per year (Estimated) • In office • 5+ years exp • Bengaluru
Go
Python
Rust
AI/ML
Computer Vision
CUDA
CUDA Toolkit
LLM
PyTorch
Quantization
SGLang
TensorFlow
TensorRT
Triton
vLLM
DevOps
HPC
Apply
Full Stack Engineer 14 days ago
$21k – $91k per year (Estimated) • In office • Bengaluru
JavaScript
Kotlin
Python
TypeScript
Java
Java
Spring Boot
AI/ML
AI Agents
Model Context Protocol
Frontend
Angular
React.js
DevOps
CI/CD
Kubernetes
Rest API
Tekton
Apply
$30k – $124k per year (Estimated) • In office • Bengaluru
Java
Node JS
Python
JavaScript
AI/ML
CUDA Toolkit
Ray
InfiniBand
NCCL
DevOps
Gateway API
GitOps
Kubernetes
Apply
$27k – $70k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Bengaluru
Java
Python
AI/ML
CUDA Toolkit
Ray
InfiniBand
NCCL
NVLink
DevOps
ArgoCD
etcd
Gateway API
GitOps
Grafana
Kubernetes
OpenTelemetry
Prometheus
Robotics
NVIDIA Omniverse
Apply
$25k – $67k per year (Estimated) • In office • 5+ years exp • Bengaluru
Java
Scala
AI/ML
Hadoop
Apply
AI Solutions Architect 10 hours ago
$33k – $80k per year (Estimated) • Remote/Hybrid • Full-Time • Bengaluru
Apply
$18k – $80k per year (Estimated) • In office • Full-Time • Bengaluru
Bash
Python
DevOps
CI/CD
GitHub
GitHub Actions
Jenkins
Marketing
LinkedIn
Apply
$20k – $50k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
DevOps
CI/CD
Apply
AI / ML Engineer 11 hours ago
$31k – $79k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • Bengaluru • Hyderabad
AI/ML
PyTorch
TensorFlow
Apply
$29k – $66k per year (Estimated) • In office • Full-Time • 13+ years exp • Bengaluru
Apply
See all jobs
This is one of many
380,196 more open roles from verified company boards, updated every day.