368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$26k – $69k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Sarvam AI is a leading Indian artificial intelligence company focused on building full-stack sovereign generative AI infrastructure, foundational large language models (LLMs), and speech technologies tailored for India’s diverse languages and enterprise requirements.

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Team

Sarvam's research teams build our own vision-language models for OCR and structured extraction. This team builds everything around them - the serving harness that turns a 3B or 30B in-house model into a production document intelligence platform.

The bet is specific: with the right harness - routing, decomposition, retries, verification, ensembling, layout awareness, confidence calibration - a small sovereign model should match or beat what teams today get from frontier hosted models like Gemini Flash, at a fraction of the cost and fully within India. Closing that gap is an engineering problem, and it is this team's problem.

We run against the full messiness of Indian documents at population scale: PAN and Aadhaar, bank statements, GST filings, insurance and medical reports, 60-page

contracts, legal filings and RFPs - across languages, scan quality, and layouts that were never designed to be machine-read.

Stack: Go, Python, Temporal, REST, Kubernetes, PostgreSQL, Redis, object storage, OpenTelemetry-based observability.

About the Role

You will own the architecture of the serving harness for Sarvam's vision models - the system that has to deliver frontier-grade extraction quality out of 3B and 30B in-house models, at national scale, with cost and latency budgets that actually close.

This means owning the hard trade-off surface directly: accuracy versus latency versus rupees per page. Multi-pass inference, model routing and cascades, self-consistency and verification passes, confidence-driven escalation, batching and caching strategy, GPU utilisation. These are the levers that decide whether the product works, and you will be the person deciding how to pull them.

You will also set the reliability bar. These pipelines process documents that customers cannot afford to lose - KYC, loan underwriting, claims, contracts. Durability, idempotency, backpressure and graceful degradation are the baseline, not the roadmap.

The architecture you set will be inherited by everything the team builds after you.

What You'll Do

Own the end-to-end architecture of the OCR and extraction serving harness: API layer, orchestration, inference layer, post-processing, delivery

Design the accuracy harness - multi-pass extraction, ensembling, cross verification, schema-constrained decoding, confidence calibration, targeted re runs - and prove its gains against held-out evaluation sets

Architect durable, resumable document workflows in Temporal: fan-out across pages, partial failure recovery, exactly-once side effects, long-running jobs measured in minutes to hours

Own the inference serving layer alongside infra: batching strategy, GPU pool management, autoscaling on real signals, queue depth and admission control, multi-model routing

Drive latency, throughput and unit economics down deliberately - profile, measure, and defend cost-per-page targets as volume scales

Build the observability substrate: distributed tracing across the pipeline, per-stage cost and quality metrics, SLOs, alerting, and post-incident rigour

Design for multi-tenancy, tenant isolation, rate limiting and fair scheduling across enterprise customers with very different load shapes

Support on-prem and constrained deployments where the whole harness has to run inside a customer's environment

Set technical direction and raise the bar through design review and mentorship of SDE 1-2 engineers

What We're Looking For

  • 5-6+ years in backend engineering, with meaningful time spent operating high throughput production systems you were on-call for

  • Deep proficiency in Go and/or Python, and the judgement to know which belongs where

  • Strong distributed systems design: queues, workflow orchestration, idempotency, backpressure, retry and timeout semantics, consistency trade-offs, graceful degradation

  • Production experience with Temporal or an equivalent durable execution engine, on workflows that mattered

  • Kubernetes in production - autoscaling, resource management, rollouts, debugging under load; GPU workload scheduling is a strong plus

  • Demonstrable experience serving ML or LLM inference in production: batching, caching, model versioning, A/B rollout, latency budgeting

  • Rigour about observability and reliability - you have designed SLOs, run incidents, and shipped the fixes

  • Hard-won cost intuition: you have made a system materially cheaper without giving up quality

Bonus Points

Direct experience with OCR, IDP, or document AI systems - Textract, Document AI, Azure DI, or something you built yourself

GPU inference stacks: vLLM, TensorRT-LLM, Triton, SGLang, Ray Serve Evaluation infrastructure for ML systems - golden sets, regression gates, human in-the-loop review loops

Experience with BFSI, healthcare, or public-sector compliance and data-residency constraints

On-prem or air-gapped deployment experience

Note

We are looking for people who can own the outcomes described here, not people who match every line of this specification. If this problem excites you and you believe you can do this work, we want to hear from you.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers ofAI with real population-scale impact.

Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

High ownership and high impact, from day one

Everything we do is AI-first, from the way we build and ship to the way we think about problems

You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier ofAI in India, Sarvam is the place to be.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$28k – $58k per year (Estimated) • In office • Moscow
AI/ML
LLM
RAG
Management
n8n
Apply
$18k – $53k per year (Estimated) • Remote/Hybrid • Moscow
AI/ML
Fine-tuning
LLM
PyTorch
RAG
vLLM
VLM
Apply
$18k – $46k per year (Estimated) • Remote • Full-Time • Tula
C#
JavaScript
Node JS
SQL
TypeScript
C#
.NET
Node JS
InversifyJS
Databases
DynamoDB
MySQL
AI/ML
Claude
Copilot
Cursor
OpenAI Codex
Frontend
Angular
React.js
Tailwind CSS
Mobile
Dependency Injection
DevOps
AWS
AWS Lambda
CI/CD
OpenTelemetry
Rest API
Terraform
Amazon CloudWatch
Amazon S3
API Gateway
GitHub
Cybersecurity
HIPAA
Apply
Java Developer 1 day ago
$31k – $54k per year (Estimated) • Remote • 5+ years exp • Tula
C#
JavaScript
TypeScript
Java
C#
.NET
Java
Spring Boot
Databases
Apache Kafka
ElasticSearch
PostgreSQL
RabbitMQ
Frontend
Angular
Bootstrap
React.js
DevOps
Docker
Git
Jenkins
Kubernetes
Rest API
GitLab
QA
Swagger
Apply
$27k – $58k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Moscow
C#
JavaScript
SQL
TypeScript
C#
ASP.NET Core
Databases
Apache Kafka
MS SQL
Frontend
Angular
React.js
DevOps
Docker
Kubernetes
Apply
DevOps Engineer 4 days ago
$18k – $81k per year (Estimated) • In office • Full-Time • Bengaluru
Python
DevOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
Azure
Blue-Green Deployment
CI/CD
Crossplane
GitHub Actions
GitLab CI
Grafana
Helm
Kubernetes
Kustomize
Loki
Prometheus
Terraform
GitHub
GitLab
IAM
Apply
$29k – $66k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Python
SQL
AI/ML
Fine-tuning
LLM
Multimodal AI
Speech Recognition
Text-to-Speech
Apply
Visual Designer 7 days ago
$16k – $48k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Design
Adobe After Effects
Adobe Photoshop
Blender
Figma
Apply
Motion Designer 7 days ago
$16k – $49k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
JavaScript
Frontend
Three.JS
Mobile
Lottie
Game Dev
GLSL
Houdini
Design
Adobe After Effects
Adobe Photoshop
Blender
Cinema 4D
Figma
Apply
$26k – $60k per year (Estimated) • In office • Full-Time • 4+ years exp • Delhi
AI/ML
Fine-tuning
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.