985,271open jobs
58,916companies
160,722added this week
Browse all
Salary
≈ $18k – $45k per year (Estimated)
Location
In office (Delhi)
Seniority
Senior · 5+ years exp

First seen by Alion on Sep 30, 2026.

Overview
Company
Impact
Profile match
Chegg is an online textbook rental company where students can rent, buy and sell text books from its platform. Headquartered in the US with offices in India and several other countries, Chegg became a public company by listing in 2013.

Senior Software Engineer - AI Evals

About the Role :

Chegg is building the next generation AI Evaluation Platform to keep improving learner experience across product and services from copilots to autonomous agents. As we scale our use of large language models and agentic systems, the quality, safety, and reliability of these systems depend on rigorous, well-designed evaluations. We're hiring a Senior Engineer to own the design and implementation of evaluation frameworks ("evals") that measure model and agent performance across accuracy, reasoning, safety, and pedagogical quality and that directly inform how we train, fine-tune, and ship AI systems.

This is a high-leverage, cross-functional role sitting at the intersection of ML research, data engineering, and product. You'll own evaluation solutions end-to-end - from framework design through deployment and production monitoring - partnering closely with ML researchers, applied scientists, and product teams to define what "good" looks like for models and agents and turn measurements into actionable outcomes.

What You'll Do :

- Own evaluation solutions end-to-end: design the methodology, build the pipeline and tooling, deploy it into training and production workflows, and maintain/iterate on it over time - with minimal hand-off to other teams.

- Design and build evaluation frameworks and harnesses for LLMs and multi-step AI agents, covering offline benchmarks, online/production evals, and human-in-the-loop review.

- Define rubrics and scoring methodologies for open-ended tasks (e.g., tutoring quality, step-by-step reasoning, citation accuracy) where correctness isn't binary.

- Build automated, model-graded (LLM-as-judge) and rule-based evaluators, and validate them against human judgment for reliability and bias.

- Develop agent-specific evaluations: tool-use correctness, multi-turn task completion, planning/trajectory quality, failure recovery, and cost/latency tradeoffs.

- Create golden datasets, adversarial test sets, and regression suites that catch quality and safety regressions before they reach production.

- Partner with research teams to translate eval results into training signal - informing SFT/RLHF/RLAIF data curation, reward modeling, and fine-tuning priorities.

- Instrument production systems to collect real-world interaction data and feed it back into the evaluation and training loop.

- Build dashboards and reporting that give researchers, PMs, and leadership a clear, trustworthy view of model quality trends across releases.

- Drive eval methodology rigor: statistical significance, inter-rater reliability, sampling strategy, and avoiding metric gaming or overfitting to benchmarks.

- Collaborate with Trust & Safety and Legal/Compliance stakeholders to build evals for bias, hallucination, academic integrity, and other responsible-AI dimensions relevant to an education product.

- Mentor other engineers on eval best practices and help establish evaluation as a first-class part of the model development lifecycle.

What We're Looking For :

- 5+ years of software/ML engineering experience, including hands-on work building or maintaining evaluation, testing, or measurement infrastructure for ML systems.

- Direct experience with foundation model evaluation and benchmarking beyond using APIs - experience gained at a foundation model lab or similarly frontier research environment is strongly preferred.

- Demonstrated ability to own an evaluation solution end-to-end - from initial design and dataset/methodology creation through pipeline build, deployment, and production monitoring - with minimal hand-off.

- Direct experience designing evals for LLMs and/or AI agents - e.g., benchmark design, LLM-as-judge pipelines, human annotation platforms, or A/B and offline/online eval frameworks.

- Strong programming skills in Python, PyTorch / Tensorflow and experience building production-grade data/ML pipelines.

- Hands-on knowledge of key model training and evaluation platforms, particularly AWS (e.g., SageMaker, Bedrock) and Databricks (e.g., MLflow, Unity Catalog, Delta Lake).

- Solid grounding in applied statistics - comfortable reasoning about sample size, variance, significance testing, and the limitations of aggregate metrics.

- Working knowledge of how LLMs are trained and adapted (pretraining, SFT, RLHF/RLAIF/DPO) and how eval signal feeds into that lifecycle.

- Experience with agentic architectures - tool calling, multi-step planning, memory, orchestration frameworks - and the unique evaluation challenges they introduce.

- Familiarity with eval and observability tooling (e.g., internal or open-source frameworks for tracing, dataset versioning, experiment tracking).

- Excellent cross-functional collaboration skills; able to translate ambiguous product/research questions into concrete, measurable eval criteria.

- A bias toward rigor and skepticism - you instinctively question whether a metric is actually measuring what it claims to.

Bonus Points :

We prioritize candidates with direct experience in the following areas :

- Post-training & evals : hands-on experience with post-training techniques (SFT, RLHF, RLAIF, DPO) and the evaluation methodologies used to validate them, ideally gained at a frontier model labs.

- Model benchmarking : experience building, running, or maintaining benchmark suites used to track and compare frontier model capabilities across training runs and releases.

- Data quality : experience with the data quality side of model training - curation, filtering, deduplication, and quality scoring of pretraining and post-training datasets.

Why Chegg :

You'll shape how Chegg measures and improves the AI systems millions of students rely on - with direct influence on model training decisions, product quality, and responsible AI practices. This role offers high visibility across research, engineering, and product leadership, and the opportunity to help define evaluation as a discipline within the company's AI strategy.

Skills

Artificial Intelligence, Agentic AI, LLM, Machine Learning, Python, Tensorflow, PyTorch, AWS, AWS SageMaker

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
985,271 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Delhi
$12k – $19k per year (gross) • In office • Full-Time • Hanoi
JavaScript
Node JS
Databases
PostgreSQL
Supabase
Frontend
React.js
Apply
Software Developer 26 min ago
$93k – $110k per year • In office • 2+ years exp • Bachelor's Degree • Syracuse
JavaScript
TypeScript
SQL
C#
Node JS
C#
ASP.NET Core
Entity Framework Core
Frontend
Webpack
Bootstrap
React.js
JQuery
npm
DevOps
Azure DevOps
Azure
CI/CD
Git
Configuration Management
QA
Jest
Apply
Java Team Lead 22 days ago
≈ $39k – $111k per year (Estimated) • Remote (Ukraine) • Full-Time • 8+ years exp • Bachelor's Degree • Kyiv
Java
SQL
Java
Spring Boot
Databases
PostgreSQL
Mobile
Clean Architecture
DevOps
Rest API
Terraform
Prometheus
CI/CD
AWS
Docker
Kubernetes
Grafana
Amazon EKS
Amazon EC2
Amazon S3
IAM
Amazon CloudWatch
Linux
Apply
Hybrid • 7+ years exp
Scala
Databases
DynamoDB
Apache Kafka
Frontend
GraphQL
DevOps
Terraform
CI/CD
AWS
Kubernetes
Management
Agile
Apply
$200k – $230k per year • In office • Full-Time • 7+ years exp • United States
DevOps
Terraform
Istio
CI/CD
ArgoCD
AWS
Kubernetes
Cloudflare
DNS
Apply
≈ $21k – $42k per year (Estimated) • In office • 5+ years exp • Delhi • Hyderabad • Bengaluru
Python
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
LangChain
Spark
LlamaIndex
MLFlow
Fine-tuning
Embeddings
Scikit-learn
Prompt Engineering
AI Agents
NLP
Semantic Kernel
Transformers
TensorFlow
PyTorch
LLM
RAG
OpenAI
Hugging Face
LLM Guardrails
Agentic Workflows
Machine Learning
DevOps
Azure DevOps
GitHub Actions
Azure
CI/CD
Docker
Kubernetes
Azure AKS
Apply
≈ $7k – $18k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • India
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
Incident Management
GitHub
Apply
≈ $12k – $24k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • India
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
Incident Management
GitHub
Apply
$62k – $91k per year • In office • Full-Time
Python
SQL
Databases
Databricks
AI/ML
Machine Learning
DevOps
Terraform
Azure
CI/CD
AWS
Analytics
ETL/ELT
Apply
≈ $20k – $47k per year (Estimated) • In office • 5+ years exp • Hyderabad
Python
Java
AI/ML
Embeddings
AI Agents
LLM
RAG
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Management
ServiceNow
Agile
Apply
≈ $25k – $59k per year (Estimated) • In office • 7+ years exp • Delhi
Python
DevOps
AWS
Docker
Kubernetes
Incident Management
IAM
Linux
Windows
Cybersecurity
SentinelOne
MITRE ATT&CK
Exabeam
SIEM
Apply
In office • Delhi
Apply
In office • 3+ years exp • Bengaluru • Delhi
Go
JavaScript
Rust
SQL
C++
Databases
PostgreSQL
Redis
Frontend
React.js
DevOps
Rest API
WebSockets
Docker
Kubernetes
Linux
Apply
≈ $16k – $41k per year (Estimated) • In office • 2+ years exp • Delhi • Gurgaon
JavaScript
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Databases
MySQL
PostgreSQL
Frontend
Bootstrap
JQuery
DevOps
Rest API
Git
GitHub
GitLab
Apply
≈ $17k – $43k per year (Estimated) • In office • 3+ years exp • Delhi • Gurgaon
Python
Python
FastAPI
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
AWS Lambda
Apply
PHP Developer 1 day ago
In office • 2+ years exp • Delhi • Gurgaon • Surat
JavaScript
PHP
PHP
Yii
Laravel
Symfony
CodeIgniter
Databases
MySQL
Frontend
JQuery
DevOps
Rest API
Git
Linux
Apply
In office • 4+ years exp • Hyderabad • Delhi • Gurgaon
JavaScript
Java
Kotlin
TypeScript
SQL
Java
Maven
Spring Boot
Hibernate
Gradle
Kotlin
Mockito
Databases
MySQL
MS SQL
Frontend
Angular
Bootstrap
React.js
Mobile
JUnit
DevOps
Rest API
GitLab CI
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Management
Agile
Apply
See all jobs
This is one of many
985,271 more open roles from verified company boards, updated every day.