368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$22k – $66k per year (Estimated)
Location
In office (Bengaluru, Chennai)
Seniority
Junior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Sarvam AI is a leading Indian artificial intelligence company focused on building full-stack sovereign generative AI infrastructure, foundational large language models (LLMs), and speech technologies tailored for India’s diverse languages and enterprise requirements.

Performance Engineer, Inference

Part of Sarvam's Performance Engineering team. We are hiring two specialized performance roles - Inference (this posting) and Kernels (companion posting). They are a vertical stack: the kernels team authors the µs-level GPU code, and the inference team integrates it into a running serving stack and owns the system-level numbers. If your depth genuinely spans both, apply to either and tell us - but most candidates are strongest in one, and we hire for that depth.

Location: [Bengaluru / Chennai / Hybrid / On-site] · Team: Performance Engineering · Level: Senior

About the team

Sarvam serves multiple model families - small and large LLMs, Mixture-of-Experts, Indic ASR & TTS, streaming models and multimodal models - across a multi-node, multi-tenant fleet of Hoppers and Blackwells. The Performance Engineering team owns the numbers the rest of the company plans against: how fast we serve, how much it costs, and how much we get out of every GPU. This team works at the intersection of the serving runtime, the kernel layer, and the SRE org that keeps the fleet alive.

About the role

You will own Sarvam's production serving path for large distributed models end to end. You should be source-level fluent in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM - able to read and modify it where stock behavior does not fit our workloads - and you should operate and extend a distributed-serving stack at depth: disaggregated prefill-decode across nodes, distributed KV/cache transfer, and the routing and scheduling that span them. You will integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack, and you will build and train your own speculators - draft models, distillation from the target, acceptance-rate tuning against the live serving distribution - rather than only wiring in stock implementations. You will produce, and defend, the latency and throughput numbers the company plans against, and you will spend significant time in cross-team work with architecture co-design, the kernels team, the model team, and SRE.

Your scoreboard: TTFT (p50 / p95 / p99), TPOT, throughput, GPU utilization, and cost per million tokens.

What we're looking for

  • 5+ years in ML systems, with 2+ years on inference serving at production scale. Your record shows concrete outcomes - tokens per day, throughput wins, p99 reductions - rather than "deployed a model."

  • Experience serving 100B+ parameter models in production across multi-node tensor, pipeline, or expert parallelism.

  • Source-level fluency in one of SGLang, vLLM, Dynamo, or TensorRT-LLM - you have modified the scheduler, the KV allocator, or the disaggregation path - and reading-level familiarity with the other three.

  • Distributed serving at operating-and-extending depth: a disaggregated prefill-decode stack, distributed KV/cache transfer (Mooncake or equivalent), and cross-node routing and scheduling. You have run one of these in production and modified it where it didn't fit.

  • Speculative decoding as a build-and-train competency: you have trained your own draft models or speculators (EAGLE / DFlash or otherwise), distilled them from a target model, measured and tuned acceptance rate against a real serving distribution, and composed speculation with the rest of the stack - not only integrated a stock implementation.

  • Deep understanding of KV cache internals: block tables, copy-on-write, prefix sharing, and fragmentation.

  • Working command of TP / PP / EP, NCCL primitives, and how they interact with the scheduler.

  • Multi-tenant serving: model co-location and MIG / MPS isolation.

  • Profiling fluency with Nsight Systems, framework tracing, and py-spy / perf.

  • C++ and CUDA at a read-and-modify level.

  • On-call ownership of an inference SLO.

Strong pluses

  • Upstream contributions to SGLang, vLLM, Dynamo, llm-d, TensorRT-LLM, or LMDeploy on non-trivial code paths.

  • Direct production experience with Dynamo or llm-d at scale.

  • Having operated a forked runtime in production.

  • Published or shipped speculator work - a draft model or speculative-decoding technique you trained and measured.

  • MoE serving at scale, long-context (128K+), multi-model serving, or Indic and multilingual workloads.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$19k – $51k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
C++
C
C++
STL
C
Valgrind
DevOps
RTOS
IoT
FreeRTOS
Apply
$120k – $220k per year • Equity 0.2–0.8% • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
C++
Go
JavaScript
Python
TypeScript
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
$19k – $28k per year (net) • Remote • Full-Time • Moscow
C#
C++
C++
CMake
DevOps
CI/CD
Git
Management
Jira
Slack
Apply
$17k – $47k per year (Estimated) • Remote/Hybrid • Internship • 5+ years exp • Minsk
Go
JavaScript
Python
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
Prompt Engineering
RAG
Anthropic
Function Calling
OpenAI
Frontend
React.js
DevOps
CI/CD
Docker
Git
Terraform
GitHub
Apply
DevOps Engineer 4 days ago
$18k – $81k per year (Estimated) • In office • Full-Time • Bengaluru
Python
DevOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
Azure
Blue-Green Deployment
CI/CD
Crossplane
GitHub Actions
GitLab CI
Grafana
Helm
Kubernetes
Kustomize
Loki
Prometheus
Terraform
GitHub
GitLab
IAM
Apply
$29k – $66k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Python
SQL
AI/ML
Fine-tuning
LLM
Multimodal AI
Speech Recognition
Text-to-Speech
Apply
Visual Designer 7 days ago
$16k – $48k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Design
Adobe After Effects
Adobe Photoshop
Blender
Figma
Apply
Motion Designer 7 days ago
$16k – $49k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
JavaScript
Frontend
Three.JS
Mobile
Lottie
Game Dev
GLSL
Houdini
Design
Adobe After Effects
Adobe Photoshop
Blender
Cinema 4D
Figma
Apply
$26k – $60k per year (Estimated) • In office • Full-Time • 4+ years exp • Delhi
AI/ML
Fine-tuning
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.