368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$155k – $233k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Harvey is an artificial intelligence technology company headquartered in San Francisco, California, and founded in 2022. The company develops a generative AI platform specifically designed for the legal industry to automate tasks such as contract analysis, due diligence, litigation strategy, and regulatory compliance. It serves global law firms and in-house legal departments, operating as a venture-backed enterprise with strategic partnerships with organizations like OpenAI and major professional service networks.

Why Harvey

At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.

This is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth - personal, professional, and financial - is unmatched.

Our team moves fast, takes ownership, and is deeply committed to the mission - operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you.

At Harvey, the future of professional services is being written today - and we’re just getting started.

Role Overview

We’re looking for a senior operator to own the quality bar behind Harvey’s human evaluations. As we scale globally, the volume of eval work is growing 10x, but volume only matters if the output is trusted. This role makes Human Data’s signal decision-grade: rigorous, calibrated, and reproducible enough that Product, Engineering, and AI Research act on it to ship.

As a member of our Evaluation Operations team, you’ll work alongside our Evaluation Operations Manager (who runs throughput and coordination) and partner closely with Applied Legal Researchers, Product, Engineering, and AI Research. You'll set the standard for what "good" looks like across eval data and methodology, and own the data analyses to make conclusions that teams will rely on, building the stakeholder trust that lets EPD act on the signal.

What You'll Do

  • Own the quality bar for Harvey’s human evaluations: define what “good” looks like for eval methodology and data analysis, and produce decision-grade outputs for EPD

  • Author and maintain the evaluation guidelines, instructions, and databases that contract attorneys work from

  • Standardize and streamline rubric and evaluation design into repeatable templates and one documented methodology, partnering with Applied Legal Research (ALR), who supplies feature-specific legal depth

  • Own contract-attorney quality: onboarding, calibration training, inter-rater reliability, and the feedback loop (including benchmarks and gold references) that keeps judgment consistent across attorneys and over time

  • Conducting quantitative and qualitative statistical analyses, diagnosing error states, investigating root causes, and turning raw eval results into a structured, prioritized signal Product and ALR can act on

  • Run QA on vendor and contract-attorney deliverables against a defined bar before results inform a launch decision

  • Ensure the quality bar holds across jurisdictions and non-English geographies as coverage expands

  • Establish one standard, documented way to analyze eval results, and build lightweight operational dashboards to track rater capability and eval-program health

  • Support ALR in a review step that certifies an evaluation is sound before it scales to contract attorneys

  • Partner with ALR and Analytics to determine where human eval aligns with online signal and where it can provide expanded insights

What You Have

  • 6+ years in product operations, research operations, evaluation/QA operations, or quality program management

  • A track record of owning quality inputs (guidelines, instructions, benchmarks, QA procedures) for complex, expert-driven or human-in-the-loop work

  • Experience onboarding, training, and calibrating a distributed pool of expert raters, annotators, or reviewers, and running the feedback loop that improves their quality over time

  • Enough grounding in measurement concepts (calibration, inter-rater reliability, sampling, rubric design) to independently set up and own the quality of our evaluation loop yourself

  • Experience with running quantitative and qualitative data analyses, interpreting and running statistical tests on evaluation data (natively or with AI tool support), and communicating conclusions to various stakeholders

  • A record of scaling and streamlining evaluation quality processes under shipping pressure, with a bias toward documentation and reproducibility over one-off analysis

  • Ability to work deeply with domain experts (e.g., ALR / lawyers) and translate nuanced judgment into repeatable, documented standards

  • Strong cross-functional coordination across Product, Engineering, Research, ALR, and data providers/vendors

  • Clear communicator who can build credibility and trust with stakeholders

  • Bias to action and high ownership, from writing the guideline to auditing a vendor batch line by line

Bonus Points

  • Experience in legal tech or working with domain experts in regulated industries

  • Experience owning quality across multiple markets, languages, or jurisdictions

  • Built calibration, inter-rater reliability, or capability-tracking systems for annotation or evaluation pipelines

  • Experience transitioning evaluation work in-house or otherwise improving evaluation ROI

  • Familiarity with LLM-as-judge / automated evaluation used alongside human eval

  • Early employee at a hyper-growth startup, or experience at a world-class product or platform operations org

Compensation

$155,400 - $233,200 USD

Depending on your location, an Applicant Privacy Notice may apply to you. You can find all of our Applicant Privacy Noticeshere.

Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made by emailing [email protected]

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$140k – $150k per year • Remote • Full-Time • 5+ years exp • United States
TypeScript
JavaScript
AI/ML
AI Agents
Claude
Claude Code
LLM
LLM Guardrails
Frontend
GraphQL
React.js
Mobile
Twilio
DevOps
AWS
Cybersecurity
HIPAA
SOC 2
Apply
$211k – $428k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • San Francisco
Go
Python
TypeScript
AI/ML
Claude
Claude Code
Cursor
LLM
OpenAI Codex
Vapi
AI Agents
Model Context Protocol
DevOps
CI/CD
Apply
$32k – $90k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Barcelona • Madrid
Python
AI/ML
LightGBM
LLM
NumPy
Pandas
PyTorch
Scikit-learn
TensorFlow
XGBoost
AI Agents
Analytics
Matplotlib
ETL/ELT
Apply
AI Architect 6 days ago
$72k – $157k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree • Barcelona • Madrid
AI/ML
AI Agents
AutoGen
CrewAI
Kubeflow
LangChain
LlamaIndex
LLM
MLFlow
RAG
Semantic Search
Vertex AI
EU AI Act
Knowledge Graph
Semantic Search
DevOps
AWS
Azure
Docker
GCP
Kubernetes
OpenShift
Vector
Cybersecurity
GDPR
Apply
$127k – $253k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree
Python
AI/ML
AI Agents
Embeddings
Function Calling
LangChain
LangGraph
LLM
Model Context Protocol
RAG
Semantic Kernel
Semantic Search
Anthropic
LLM Guardrails
OpenAI
Semantic Search
Structured Outputs
DevOps
Azure
CI/CD
Vector
Management
ServiceNow
Apply
$231k – $340k per year • Remote/Hybrid • Full-Time • 10+ years exp • New York
Python
SQL
Databases
Amazon Redshift
Apache Kafka
Databricks
Delta Lake
Google BigQuery
Snowflake
Trino
AI/ML
Dagster
dbt
Flink
Spark
AI Agents
DevOps
AWS
Azure
GCP
Kubernetes
Pulumi
Terraform
Apply
$231k – $340k per year • Remote/Hybrid • Full-Time • 10+ years exp • San Francisco
Python
SQL
Databases
Amazon Redshift
Apache Kafka
Databricks
Delta Lake
Google BigQuery
Snowflake
Trino
AI/ML
Dagster
dbt
Flink
Spark
AI Agents
DevOps
AWS
Azure
GCP
Kubernetes
Pulumi
Terraform
Apply
$193k – $290k per year • Remote/Hybrid • Full-Time • 5+ years exp • New York
Python
SQL
Databases
Amazon Redshift
Apache Kafka
Databricks
Delta Lake
Google BigQuery
Snowflake
Trino
AI/ML
Dagster
dbt
Flink
Spark
AI Agents
DevOps
AWS
Azure
GCP
Kubernetes
Pulumi
Terraform
Apply
$193k – $290k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Python
SQL
Databases
Amazon Redshift
Apache Kafka
Databricks
Delta Lake
Google BigQuery
Snowflake
Trino
AI/ML
Dagster
dbt
Flink
Spark
AI Agents
DevOps
AWS
Azure
GCP
Kubernetes
Pulumi
Terraform
Apply
$97k – $194k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Dublin
JavaScript
Python
SQL
Apex
Apex
MuleSoft
Databases
Google BigQuery
Snowflake
AI/ML
AI Agents
ChatGPT
Claude
LLM
DevOps
SLI/SLO/SLA
Management
Zapier
Marketing
HubSpot
Marketo
Salesforce
Zendesk
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$293k – $385k per year • In office • Full-Time • San Francisco
AI/ML
OpenAI
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.