744,295open jobs
44,695companies
107,333added this week
Browse all
Salary
≈ $141k – $242k per year (Estimated)
Location
Remote (Palo Alto, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026. Clera scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Clera is a San Francisco company that runs an AI recruiting platform positioned as a talent agent rather than a job board, matching engineers and other technical candidates to roles at venture-backed startups. Its system reads a candidate's background and preferences, then represents them to hiring teams at companies funded by firms such as Andreessen Horowitz, Y Combinator, Index Ventures and General Catalyst, compressing the introduction step that traditional headhunting handles manually. The model is aimed at the segment where recruiter fees are highest and candidate supply is thinnest, and the platform handles screening, scheduling and pipeline tracking for both sides of the match.

About the Role

This role sits at the intersection of applied data science and AI product quality for a small, fast-moving AI productivity startup building autonomous agents that handle email, calendar, browser, and business software tasks. You will own the measurement of agent quality end-to-end: turning ambiguous product behavior into rigorous, actionable evaluation systems that directly guide engineering and product decisions.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, edge cases, ambiguous requests, and adversarial scenarios.

  • Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure grader agreement, false positives, and false negatives.

  • Analyze traces, tool calls, model outputs, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

  • Build dashboards and release-quality signals that make evaluation results understandable and actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

What We're Looking For

  • 5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems.

  • Demonstrated experience designing and implementing evaluation frameworks, grading systems, and success criteria for ML or AI systems in production.

  • Strong Python and SQL proficiency with the ability to build automated data pipelines and production-quality analysis code at scale.

  • Solid statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems.

  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance.

  • Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes of language model systems.

  • Ability to connect quantitative patterns to individual system traces and identify failure origins across model, prompt, context, tools, data, and application logic.

  • Experience communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.

  • Comfort operating with high ownership in ambiguous, fast-moving environments, independently turning open-ended quality questions into evaluation systems.

  • Experience with LLM-as-a-judge systems, agentic or multi-step task evaluation, or benchmarking platforms for AI systems is a strong plus.

Location

On-site in Palo Alto, California, United States. Visa sponsorship is not available for this role.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
744,295 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Palo Alto
$225k – $250k per year • In office • 6+ years exp • Bachelor's Degree • New York
Python
Java
SQL
Databases
Snowflake
AI/ML
Dagster
dbt
Apply
Senior Data Scientist 3 hours ago
≈ $121k – $234k per year (Estimated) • In office • 3+ years exp • Master's Degree • Cambridge
Python
AI/ML
arXiv
Machine Learning
Apply
Data Engineer 3 hours ago
≈ $122k – $211k per year (Estimated) • Remote (United States) • 5+ years exp • Bachelor's Degree
Python
Java
SQL
Databases
PostgreSQL
Snowflake
Databricks
Apache Kafka
AI/ML
Spark
Airflow
DevOps
AWS
Amazon Kinesis
Linux
Unix
Analytics
AWS Glue
Management
Agile
Apply
Sr. Data Scientist 6 hours ago
$150k – $180k per year • Equity • Remote (United States) • Full-Time • 5+ years exp
Python
SQL
AI/ML
XGBoost
Scikit-learn
LightGBM
NumPy
PyTorch
Anomaly Detection
Time Series Forecasting
Machine Learning
DevOps
GCP
CI/CD
Git
AWS
Analytics
A/B Testing
Apply
Staff Data Scientist 6 hours ago
$190k – $230k per year • Equity • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree
Python
SQL
AI/ML
XGBoost
Scikit-learn
LightGBM
NumPy
PyTorch
Anomaly Detection
Time Series Forecasting
Machine Learning
DevOps
GCP
CI/CD
AWS
Analytics
A/B Testing
Apply
$162k – $243k per year • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Boston
Python
SQL
SAS
Databases
Snowflake
AI/ML
AI Agents
Analytics
Power BI
Microsoft Excel
Apply
$192k – $288k per year • Equity • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Boston
Python
SQL
SAS
Databases
Snowflake
AI/ML
AI Agents
Analytics
Power BI
Microsoft Excel
Apply
$94k – $142k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Mississauga
Python
JavaScript
TypeScript
SQL
AI/ML
Copilot
Edge AI
Machine Learning
DevOps
GitHub Actions
GitLab CI
CI/CD
Jenkins
Git
Docker
Kubernetes
Shift-Left
Self-Healing
Pipeline as Code
Tekton
Cybersecurity
Shift-Left Security
Management
Agile
Scrum
QA
Selenium
Apply
$83k – $149k per year • Equity • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
SQL
AI/ML
Physical AI
Analytics
Tableau
Power BI
Apply
≈ $129k – $241k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Master's Degree • Uxbridge
Python
SQL
Databases
Databricks
AI/ML
Multimodal AI
NLP
Explainable AI
Machine Learning
DevOps
Azure
AWS
Apply
≈ $139k – $239k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp • Palo Alto
Python
SQL
AI/ML
Function Calling
AI Agents
LLM
Machine Learning
Apply
Data Scientist 6 days ago
$80k – $125k per year • In office • Full-Time • 5+ years exp • PhD • Berlin
AI/ML
Cursor
Claude Code
AI Agents
Apply
$100k – $200k per year • In office • Full-Time • 3+ years exp • Munich
Python
AI/ML
Reinforcement Learning
Post-training
Reward Modeling
Machine Learning
Apply
$100k – $200k per year • In office • Full-Time • 3+ years exp • Munich
AI/ML
AI Agents
Post-training
Machine Learning
Apply
$100k – $200k per year • In office • Full-Time • 3+ years exp • Munich
Rust
TypeScript
AI/ML
AI Agents
Apply
Senior ML Engineer 9 hours ago
$188k – $200k per year • Hybrid • 5+ years exp • Master's Degree • Palo Alto
Python
C++
AI/ML
Fine-tuning
AI Agents
Amazon SageMaker
Machine Learning
DevOps
GCP
AWS
Apply
Workplace Manager 1 day ago
≈ $68k – $143k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Palo Alto
Apply
≈ $145k – $282k per year (Estimated) • In office • 3+ years exp • Palo Alto
JavaScript
Java
Java
Spring Boot
Databases
RabbitMQ
Apache Kafka
Frontend
React.js
DevOps
CI/CD
Kubernetes
Management
Agile
Apply
Corporate Counsel 1 day ago
$260k per year • Remote (United States) • Full-Time • Palo Alto
Apply
$145k per year • In office • Palo Alto
Apply
See all jobs
This is one of many
744,295 more open roles from verified company boards, updated every day.