368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$11k – $44k per year (Estimated)
Location
In office (Bengaluru, Chennai)
Seniority
Junior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Sarvam AI is a leading Indian artificial intelligence company focused on building full-stack sovereign generative AI infrastructure, foundational large language models (LLMs), and speech technologies tailored for India’s diverse languages and enterprise requirements.

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

Sarvam runs a large, multi-vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs and must run uninterrupted for weeks, and inference services that must hold a flat p99 under production load. Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve.

This is not a Kubernetes administration role. We assume Kubernetes fluency as a baseline. The difficulty lies above and below it - in parallel filesystems under heavy checkpoint load, in RDMA fabrics that degrade quietly, in NCCL hangs whose root cause may be the network or the kernel, in driver and firmware drift across heterogeneous hardware, and in distributed training failures that masquerade as infrastructure faults.

We are hiring a team of specialists rather than a set of identical generalists. This posting covers five areas of focus. We expect candidates to bring genuine depth in one and working fluency across the others, because on a shared fleet a storage problem often first appears as a training hang, and the engineer on call must route an incident correctly before anyone can resolve it.

When you apply, please indicate the area of focus that best matches your experience. Strong generalists are welcome; we will place you where your depth is most useful.

What You’ll Do

  • Operate the GPU fleet end to end across training and serving - provisioning, observability, capacity, and fleet health.

  • Hold a meaningful on-call rotation, write runbooks that hold up under pressure, and drive postmortems that produce durable fixes.

  • Build the internal tooling the team relies on, rather than operating off-the-shelf systems alone.

  • Partner with ML and platform teams to keep large runs alive and serving latency predictable.

What We're Looking For

  • 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.*

  • Demonstrated on-call ownership of infrastructure that mattered, with a track record of postmortems that led to real change.

  • Proficiency in Python or Go, used to build and maintain internal tooling.

  • Working fluency across all five areas of focus below - enough to recognize, triage, and route a problem outside your specialty, even if the fix belongs to a teammate.

* For the Storage and Fabric areas of focus, we will weigh deep domain expertise against the GPU-cluster requirement; exceptional specialists with less direct GPU-fleet time are encouraged to apply.

Bring depth in one of the five areas below; expect to be conversational across the rest.

  • Distributed high-performance storage - operate a parallel filesystem (Lustre, GPFS, WEKA, or BeeGFS) at scale and keep it from stalling under checkpoint-write storms.

  • Fabric & RDMA networking - InfiniBand or RoCE health, NVLink/NVSwitch topology, and RDMA debugging that catches degradation before the workload feels it.

  • GPU systems reliability - NCCL debugging, driver and firmware lifecycle across a mixed fleet, and DCGM-based node health at scale.

  • Kubernetes platform reliability - the GPU operator stack, scheduling (Slurm-on-k8s or pure k8s), multi-tenant isolation, and cost/SLO primitives.

  • Training & inference workload reliability - hang and straggler detection, checkpoint/restart, and protecting serving p99 with HA and DR.

Bonus Points

  • Slurm and Kubernetes hybrid environments.

  • On-premise GPU deployment, including coordination with datacenter operations on power, cooling, and InfiniBand cabling.

  • Experience with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud etc.

  • Multi-tenant GPU isolation (MIG, MPS, time-slicing) in production.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.

Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

High ownership and high impact, from day one

Everything we do is AI-first, from the way we build and ship to the way we think about problems

You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$32k – $64k per year (Estimated) • Remote • Contractor • Saint Petersburg
Python
DevOps
CI/CD
GitLab CI
Kubernetes
SLI/SLO/SLA
GitLab
Apply
$34k – $83k per year (Estimated) • Remote/Hybrid • Full-Time • Bengaluru
Node JS
JavaScript
Databases
Apache Kafka
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Datadog
FinOps
GitHub Actions
Grafana
Istio
Kubernetes
Platform Engineering
Prometheus
SLI/SLO/SLA
Terraform
GitHub
IAM
Cybersecurity
GDPR
Microsoft Defender
Microsoft Defender for Cloud
Okta
PCI DSS
Management
ServiceNow
Apply
$42k – $91k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Bengaluru
Kotlin
Node JS
Swift
TypeScript
JavaScript
Node JS
Nest.JS
Databases
Apache Kafka
Frontend
Angular
React.js
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub Actions
Kubernetes
GitHub
QA
Appium
BrowserStack
Playwright
Selenium
Apply
Data Engineer 3 1 day ago
$21k – $52k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bengaluru
Java
Python
Scala
SQL
Java
Maven
Python
pySpark
Databases
Apache Kafka
Databricks
Snowflake
AI/ML
Hadoop
Spark
DevOps
Azure
Analytics
Tableau
ETL/ELT
Apply
$16k – $40k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Gurgaon
Python
SQL
Analytics
Power BI
Tableau
ETL/ELT
Apply
DevOps Engineer 4 days ago
$18k – $81k per year (Estimated) • In office • Full-Time • Bengaluru
Python
DevOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
Azure
Blue-Green Deployment
CI/CD
Crossplane
GitHub Actions
GitLab CI
Grafana
Helm
Kubernetes
Kustomize
Loki
Prometheus
Terraform
GitHub
GitLab
IAM
Apply
$29k – $66k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Python
SQL
AI/ML
Fine-tuning
LLM
Multimodal AI
Speech Recognition
Text-to-Speech
Apply
Visual Designer 7 days ago
$16k – $48k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Design
Adobe After Effects
Adobe Photoshop
Blender
Figma
Apply
Motion Designer 7 days ago
$16k – $49k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
JavaScript
Frontend
Three.JS
Mobile
Lottie
Game Dev
GLSL
Houdini
Design
Adobe After Effects
Adobe Photoshop
Blender
Cinema 4D
Figma
Apply
$26k – $60k per year (Estimated) • In office • Full-Time • 4+ years exp • Delhi
AI/ML
Fine-tuning
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.