372,956open jobs
9,661companies
50,233added this week
Browse all
Salary
$144k – $255k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Clera is a San Francisco-based AI recruiting and talent-matching platform designed as an AI talent agent for candidates and hiring teams. Acting as a tech-driven alternative to traditional headhunting, Clera directly connects job seekers to open roles at top startups backed by venture firms like Andreessen Horowitz (a16z), Y Combinator, Index Ventures, and General Catalyst.

This a Full Remote job, the offer is available from: California (USA)

About the Role

We're an early-stage AI company building autonomous agents that handle real work - email, calendar, browser, business software, and more. We're hiring a Data Scientist focused on Agent Evaluations & Quality to measure, understand, and continuously improve the quality of our agent capabilities.

Your mission is to translate ambiguous product behavior into measurable definitions of success, build representative evaluation datasets, design reliable graders and metrics, analyze failures, and create the feedback loops that guide engineering and product decisions. This is applied data science at the intersection of evaluation design, statistics, experimentation, production Python, and deep understanding of how LLM agents behave in real products.

This is a full-time, on-site role based in Palo Alto, CA. Visa sponsorship is not available.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate capabilities into explicit success criteria - including pass, partial-pass, and failure definitions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, ambiguous requests, long-tail behavior, edge cases, and adversarial scenarios.

  • Define and track metrics such as task success, partial completion, tool-selection accuracy, tool-use correctness, instruction adherence, factual consistency, user corrections, latency, cost, and reliability.

  • Design deterministic graders, model-based graders, and human-review processes; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement.

  • Analyze traces, tool calls, model outputs, user context, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

  • Build dashboards, reports, and release-quality signals that make evaluation results understandable and actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

What We're Looking For

Required - Dealbreakers:

  • 5+ years of experience in data science, machine learning, or analytics roles building or delivering evaluation systems, metrics frameworks, or quality measurement solutions for production systems.

  • Demonstrated experience designing and implementing evaluation frameworks, metrics, and grading systems for ML/AI systems in production.

  • Production-quality Python and SQL proficiency with the ability to build automated data pipelines and analysis code at scale.

Required Skills & Experience:

  • Experience designing evaluation methodologies: success criteria definition, dataset construction, metric selection, and distinguishing useful benchmarks from misleading ones.

  • Statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems.

  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance as product behavior evolves.

  • Working knowledge of LLM behavior - including model-based graders, tool use, retrieval systems, multi-step execution, partial completion, and practical failure modes.

  • Analytical debugging ability: connecting quantitative patterns to individual system traces and identifying failure origins across model, prompt, context, tools, data, and application logic.

  • Experience building dashboards, reports, and communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.

Nice to Have:

  • Experience with LLM-as-a-judge systems, calibration, and measurement of grader agreement, false positives, and false negatives.

  • Prior work on evaluation or benchmarking platforms for AI systems.

  • Experience with agentic systems, multi-step task execution, or tool-use evaluation.

  • Experience working on customer-facing consumer software or production ML systems with real-world user outcomes.

Location & Work Arrangement

  • On-site in Palo Alto, CA

  • Visa sponsorship is not available

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
372,956 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
In office • Full-Time • Canberra
Java
Python
SQL
Python
pySpark
Databases
Amazon Redshift
Apache Kafka
Databricks
Google BigQuery
Microsoft Fabric
Snowflake
AI/ML
dbt
Spark
DevOps
AWS
Azure
CI/CD
GCP
Platform Engineering
Analytics
ETL/ELT
Power BI
Tableau
Apply
$26k – $50k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
TypeScript
JavaScript
Python
FastAPI
Databases
Amazon Aurora
Amazon Neptune
Databricks
GraphDB
Neo4j
OpenSearch
pgvector
Pinecone
PostgreSQL
AI/ML
AI Agents
Airflow
Knowledge Graph
LangChain
LangGraph
NLP
Prefect
RAG
Reranking
Spark
Frontend
Angular
DevOps
Amazon S3
AWS
AWS Lambda
AWS Step Functions
CI/CD
Docker
Git
IAM
Terraform
Vector
Analytics
ETL/ELT
Apply
In office • Full-Time • 1+ year exp • Bachelor's Degree • Subang Jaya
Python
SQL
Apply
$17k – $56k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Hyderabad
Python
TypeScript
JavaScript
Python
Django
FastAPI
Databases
PostgreSQL
Frontend
AG Grid
Angular
Chart.js
D3.js
GraphQL
React.js
RxJS
Sass
Vue.js
DevOps
Amazon S3
AWS
GitHub
GitLab
Grafana
IAM
Apply
$93k – $151k per year • In office • Full-Time • Munich
Python
TypeScript
JavaScript
Databases
PostgreSQL
AI/ML
Embeddings
NLP
RLHF
Frontend
Next.js
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
$93k – $151k per year • In office • Full-Time • Munich
Python
TypeScript
JavaScript
Databases
PostgreSQL
AI/ML
Embeddings
NLP
RLHF
Frontend
Next.js
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
AI/LLM Engineer 1 day ago
$93k – $151k per year • In office • Full-Time • 2+ years exp • Munich
TypeScript
JavaScript
Databases
Pinecone
PostgreSQL
Weaviate
AI/ML
AI Agents
DPO
Embeddings
LLM
NLP
Reranking
RLHF
Frontend
Next.js
React.js
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Terraform
Apply
Staff Engineer 1 day ago
$93k – $151k per year • In office • Full-Time • 8+ years exp • Munich
Python
TypeScript
JavaScript
AI/ML
LLM
Frontend
React.js
DevOps
AWS
Azure
Docker
GCP
Terraform
Apply
$93k – $151k per year • In office • Full-Time • Munich
Python
SQL
TypeScript
AI/ML
AI Agents
Edge AI
Prompt Engineering
Frontend
GraphQL
Apply
up to $160k per year • In office • Full-Time • 10+ years exp
Databases
Databricks
AI/ML
Spark
DevOps
AWS
Azure
CI/CD
GCP
Apply
$87k – $173k per year (Estimated) • Remote • Contractor • 6+ years exp • Palo Alto
AI/ML
AI Agents
Model Context Protocol
Apply
In office • Internship • Master's Degree • Palo Alto
Python
SQL
AI/ML
AI Agents
Analytics
Power BI
Tableau
Management
Confluence
Jira
Microsoft Project
Apply
In office • Internship • Bachelor's Degree • Palo Alto
Python
SQL
AI/ML
AI Agents
Analytics
Power BI
Tableau
Management
Confluence
Jira
Microsoft Project
Apply
Senior ML Engineer 1 day ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
$110k – $240k per year (Estimated) • In office • Bachelor's Degree • Palo Alto
Java
Python
Scala
AI/ML
AI Agents
Fine-tuning
LLM Guardrails
DevOps
AWS
Azure
GCP
Git
GitHub
Marketing
Salesforce
Apply
See all jobs
This is one of many
372,956 more open roles from verified company boards, updated every day.