Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

Clera

Clera is a San Francisco-based AI recruiting and talent-matching platform designed as an AI talent agent for candidates and hiring teams. Acting as a tech-driven alternative to traditional headhunting, Clera directly connects job seekers to open roles at top startups backed by venture firms like Andreessen Horowitz (a16z), Y Combinator, Index Ventures, and General Catalyst.

About the Role

Join a small, technically elite team - including International Olympiad medalists and published AI researchers - building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks, you'll own the design and implementation of evaluations that frontier labs and enterprise customers trust. This role is central to ensuring our benchmarks are rigorous, credible, and tightly aligned with real-world agent performance.

This is an on-site role based in San Francisco, CA. Visa sponsorship is available.

What You'll Do

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.

  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.

  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.

  • Develop metrics and analyses that measure benchmark difficulty, reliability, and failure modes.

  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.

  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.

What We're Looking For

Required

  • 2-4 years of experience in software engineering, ML engineering, or research roles.

  • Strong proficiency in Python, Docker, and Linux environments.

  • Experience building environments, evaluations, or benchmarks for AI systems.

  • Published research or technical writing on topics such as public benchmarks, model failure modes, or evaluation methodology.

  • Deep understanding of what makes a benchmark realistic, reliable, and practically useful.

  • Curiosity and genuine ability to understand how real-world workflows operate across diverse domains.

  • Strong attention to detail - a habit of spotting subtle inconsistencies and edge cases in task design.

  • Ability to reason from first principles about task design, scoring, and failure modes.

  • Comfort thriving in unstructured problem spaces and working independently in fast-paced, early-stage environments.

  • Excellent communication skills for collaborating across time zones and with technical teams.

Compensation & Benefits

  • Salary: $150,000 - $250,000 USD annually, depending on experience.

  • Visa sponsorship available.

  • Opportunity for significant early-stage equity and career growth within a high-impact, research-driven team.

Location

This is a full-time, on-site position in San Francisco, CA. Candidates must be willing and able to work in-office.

Recommended for you based on this role

Similar stack
Same company
In your city
In office • Full-Time • 5+ year exp
AI/ML
AI Agents
DevOps
AWS
Azure
CI/CD
CloudFormation
Datadog
GCP
Kubernetes
New Relic
Prometheus
Pulumi
Terraform
Apply
$185k – $235k per year • Remote • Full-Time • New York
Databases
Apache Kafka
PostgreSQL
DevOps
AWS
CloudFormation
Terraform
Apply
$130k – $250k per year • Remote • Full-Time • 5+ year exp • Austin
Design
SolidWorks
Apply
$82k per year • Equity • In office • Full-Time
AI/ML
ChatGPT
Perplexity
Apply
$140k – $200k per year • Remote • Full-Time • New York
Marketing
HubSpot
Salesforce
Apply
$130k – $170k per year • In office • Full-Time • 2+ year exp • San Francisco
TypeScript
Node JS
JavaScript
Node JS
Prisma
Databases
Supabase
Typesense
AI/ML
Cursor
Fine-tuning
LLM
Apply
$170k – $230k per year • In office • Full-Time • 5+ year exp • Master's Degree • New York
Go
Python
SQL
AI/ML
Anomaly Detection
LangChain
LLM
PyTorch
RAG
Scikit-learn
Sentiment Analysis
TensorFlow
Time Series Forecasting
Apply
$160k – $200k per year • In office • Full-Time • New York
Python
DevOps
Amazon EKS
AWS
CI/CD
Docker
GitOps
Jenkins
Kubernetes
SLI/SLO/SLA
Terraform
Terragrunt
Cybersecurity
GDPR
SOC 2
Apply
Java Developer 7 days ago
In office • Full-Time • Los Angeles
Java
DevOps
Kubernetes
Apply
Founding Engineer 7 days ago
$70k – $100k per year • In office • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
GCP
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
San Francisco
Remote work
In office
Employment
Full-Time

Compensation

Salary
$150k – $250k per year
Equity
Equity stake in a tech company