368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$26k – $71k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Sarvam AI is a leading Indian artificial intelligence company focused on building full-stack sovereign generative AI infrastructure, foundational large language models (LLMs), and speech technologies tailored for India’s diverse languages and enterprise requirements.

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

You will own the data infrastructure that feeds our next family of foundational models. This means building petabyte-scale curation and filtering pipelines, designing the systems that decide what goes into a training run and in what proportion, and treating data quality with the same rigor a research team would treat an architectural choice.

This is not a glue-code role. The data work at a serious pretraining lab is engineering- and research-heavy: deduplication at scale, quality models, contamination detection, mixture design, curriculum and annealing, attribution and debugging. You should care deeply about all of it.

What You’ll Do

  • Design and build large-scale data pipelines for pre-training and post-training - ingestion, parsing, normalization, filtering, deduplication, tokenization, and packing - at petabyte scale.

  • Develop and continually improve quality filtering systems, including model-based quality classifiers and contamination detection.

  • Own data mixture design, curriculum, and annealing strategies in partnership with the research team. The question "what data did this model see, in what proportion, in what order" should always have a precise answer because of work you did.

  • Build the tooling that lets researchers and engineers analyze, slice, attribute, and debug the data.

  • Scale the pipeline to handle multilingual corpora, code, math, multi-source web data, and licensed datasets, while keeping provenance and licensing tracked end-to-end.

  • Partner with the training infrastructure team so that data is never the bottleneck of a production training run

What We’re Looking For

  • BS or MS in Computer Science or a closely related technical field (or equivalent demonstrated experience).

  • 3+ years of experience building large-scale data systems - petabyte-scale processing, distributed data pipelines, or comparable. Exceptional early-career candidates with a strong systems background will be considered.

  • Hands-on experience with data curation and filtering for LLM training. You should be able to walk through a pre-training corpus you helped build, end to end, and defend the choices that went into it.

  • Deep familiarity with distributed data processing frameworks - Spark, Ray, Beam, Dask, or equivalent - and the storage systems that sit underneath them.

  • Strong Python; comfort with the low-level pieces of the data path (tokenization, sharding, packing, IO patterns) and the performance tradeoffs they imply.

  • Meaningful open-source contributions in the data tooling ecosystem - datasets, dedup libraries, filtering frameworks, or substantive work on widely-used open data releases.

Bonus Points

  • Direct experience building or working with large open pretraining corpora.

  • Work on multilingual data - collection, normalization, quality scoring, and mixing across many languages.

  • Hands-on experience with model-based data quality classifiers, contamination detection, or data attribution research.

  • Familiarity with tokenization research and the practical implications of tokenizer choices on training.

  • First-author papers or technical reports on data curation, quality, or pretraining mixtures.

Why this role?

The frontier is moving towards data being the dominant lever in model quality, and the labs that get this right will define the next generation of models. You will be the person at Sarvam most responsible for that lever.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.

  • Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

  • High ownership and high impact, from day one

  • Everything we do is AI-first, from the way we build and ship to the way we think about problems

  • You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$33k – $78k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp • Bachelor's Degree • India
Apex
JavaScript
Python
TypeScript
Apex
Copado
Lightning Web Components
AI/ML
AutoGen
CrewAI
Fine-tuning
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
RAG
Semantic Kernel
Semantic Search
Synthetic Data
Vertex AI
Agentforce
AWS Bedrock AgentCore
Semantic Search
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
GitHub Actions
Jenkins
Vector
GitHub
Cybersecurity
Crowdstrike
Management
Slack
Marketing
Salesforce
Apply
Remote/Hybrid • Full-Time • 5+ years exp • Kenya
C#
Python
SQL
Python
pySpark
Databases
Databricks
MS SQL
AI/ML
Copilot
Cursor
Spark
DevOps
CI/CD
Git
GitHub
Analytics
ETL/ELT
Apply
Lead AI Engineer 9 hours ago
$30k – $73k per year (Estimated) • In office • Full-Time • Pune
Python
AI/ML
Fine-tuning
LLM
Reinforcement Learning
LLM Guardrails
AI Agents
DevOps
CI/CD
Docker
GitOps
Helm
Kubernetes
OpenShift
Platform Engineering
Vector
Apply
$44k – $99k per year (Estimated) • Equity • Remote • Full-Time • Tokyo • Nagoya • Osaka
Bash
Perl
PowerShell
Python
DevOps
Splunk
Cybersecurity
Crowdstrike
Apply
$111k – $216k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp • Bachelor's Degree • United States
Python
Ruby
Databases
PostgreSQL
DevOps
AWS
Chef
CI/CD
Configuration Management
Datadog
Git
Jenkins
Nagios
Amazon CloudWatch
GitLab
Cybersecurity
Crowdstrike
FedRAMP
Analytics
Tableau
ETL/ELT
Apply
DevOps Engineer 4 days ago
$18k – $81k per year (Estimated) • In office • Full-Time • Bengaluru
Python
DevOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
Azure
Blue-Green Deployment
CI/CD
Crossplane
GitHub Actions
GitLab CI
Grafana
Helm
Kubernetes
Kustomize
Loki
Prometheus
Terraform
GitHub
GitLab
IAM
Apply
$29k – $66k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Python
SQL
AI/ML
Fine-tuning
LLM
Multimodal AI
Speech Recognition
Text-to-Speech
Apply
Visual Designer 6 days ago
$16k – $48k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Design
Adobe After Effects
Adobe Photoshop
Blender
Figma
Apply
Motion Designer 6 days ago
$16k – $49k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
JavaScript
Frontend
Three.JS
Mobile
Lottie
Game Dev
GLSL
Houdini
Design
Adobe After Effects
Adobe Photoshop
Blender
Cinema 4D
Figma
Apply
$26k – $60k per year (Estimated) • In office • Full-Time • 4+ years exp • Delhi
AI/ML
Fine-tuning
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
Data Architect 1 hour ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
$26k – $69k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Bengaluru • Hyderabad • Chennai • Noida
Databases
Db2
IMS
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.