795,157open jobs
50,783companies
124,735added this week
Browse all
Salary
$100k – $120k per year
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
Independent benchmarks for build-vs-buy decisions across inference APIs, web search, company data, voice agent latency, and speech models. Measured first-hand and published in full.

Founding Engineer (AI + Backend)

About OpenBenchmarks

Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.

Agents increasingly prefer open, independent and grounded benchmarks to make decisions.

Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.

Our mission is to be the trusted evaluation layer for agents.

Founders previously led AI research and Infra teams at Oracle and Appfolio;

We started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.

We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.

About the role

We're a team of researchers, engineers and theorists, in person in SF. You'll be building the benchmarks and the systems that run them.

Here are the broad themes that you’ll be working on

Benchmarks and evals design - designing domain-specific benchmarks from scratch: what to measure, how to ground it, and what makes a benchmark that an agent will continuously trust and pick from.

Continuously running large-scale autonomous eval systems - infrastructure that runs without a human in the loop: vendor APIs at scale, LLM-as-judge pipelines, scoring and metric computation, drift detection as models and products change underneath.

Research - understanding model behavior - why agents choose what they choose, and what survives their scrutiny. Adjacent to this you'll work on synthetic data generation, and problems like self-improvement code/software.

You'll build benchmarks & evals that the fastest-growing AI-first companies pay close attention to and produce continuous evals that their engineering teams care about.

If you want to do your life's work, reach out.

  • Intro Call
  • Spend a day in office with the team in SF
  • 2 day work trial
  • Accepted
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
795,157 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$150k – $190k per year • Hybrid • 7+ years exp • Bachelor's Degree • Chicago
Python
PowerShell
Bash
DevOps
Ansible
GCP
OpenShift
Azure DevOps
Prometheus
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Configuration Management
Amazon EKS
Google GKE
Azure AKS
Incident Management
Linux
Apply
$160k – $200k per year • Hybrid • Chicago
Python
TypeScript
SQL
AI/ML
LangGraph
LangChain
Claude
Claude Code
Model Context Protocol
Vertex AI
AI Agents
Google ADK
OpenAI
Anthropic
OpenAI Agents SDK
A2A
Structured Outputs
LLM Evaluation
LLM Guardrails
DevOps
Rest API
GCP
Azure
CI/CD
Apply
$200k – $250k per year • Hybrid • 10+ years exp • Chicago
Python
TypeScript
AI/ML
LangGraph
LangChain
Claude
Claude Code
Model Context Protocol
Vertex AI
Embeddings
AI Agents
RAG
Google ADK
OpenAI
Anthropic
OpenAI Agents SDK
A2A
LLM Evaluation
LLM Guardrails
DevOps
gRPC
GCP
Azure
CI/CD
Docker
Kubernetes
Apply
$80k – $110k per year • Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • New York
AI/ML
AI Agents
Apply
$147k – $220k per year • In office • 2+ years exp • Bachelor's Degree • Seattle
Python
AI/ML
JAX
TPU
Machine Learning
DevOps
GCP
GitHub
Apply
$150k – $250k per year • Equity 0.2–0.5% • In office • Full-Time • 3+ years exp • San Francisco
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
AI/ML
AI Agents
LLM
Explainable AI
Frontend
React.js
DevOps
Docker
Cybersecurity
Okta
Crowdstrike
Management
Slack
ServiceNow
Apply
Founding GTM 1 day ago
$30k – $37k per year • Equity 0.2–0.6% • Remote (India) • Full-Time
AI/ML
Claude
Claude Code
Firecrawl
Marketing
LinkedIn
Apply
Founding GTM 1 day ago
$100k – $150k per year • Equity 0.2–0.6% • In office • Full-Time
AI/ML
Claude
Claude Code
Firecrawl
Marketing
LinkedIn
Apply
$100k – $200k per year • Equity 0.5–3% • In office • Full-Time • San Francisco
Python
AI/ML
Reinforcement Learning
PyTorch
LLM
Machine Learning
Apply
$100k – $200k per year • Equity 0.5–3% • In office • Full-Time • San Francisco
Python
JavaScript
Rust
TypeScript
AI/ML
AI Agents
LLM
Frontend
Next.js
React.js
DevOps
AWS
Apply
Founding GTM 1 day ago
$30k – $37k per year • Equity 0.2–0.6% • Remote (India) • Full-Time
AI/ML
Claude
Claude Code
Firecrawl
Marketing
LinkedIn
Apply
Founding GTM 1 day ago
$100k – $150k per year • Equity 0.2–0.6% • In office • Full-Time
AI/ML
Claude
Claude Code
Firecrawl
Marketing
LinkedIn
Apply
$125k – $175k per year • Equity 0.2–2% • In office • Full-Time • 3+ years exp • PhD • San Francisco
Python
AI/ML
Multimodal AI
Time Series Forecasting
Machine Learning
Robotics
Sensor Fusion
Apply
$180k – $250k per year • Remote (United States) • Full-Time • San Francisco
Python
TypeScript
Python
Hypothesis
AI/ML
AI Agents
Post-training
Machine Learning
Apply
≈ $134k – $359k per year (Estimated) • In office • Internship • San Francisco
Python
TypeScript
Python
Hypothesis
AI/ML
AI Agents
Post-training
Machine Learning
Apply
Founding Engineer 1 day ago
$110k – $180k per year • Equity 0.1–1% • In office • Full-Time • San Francisco
Python
JavaScript
Node JS
AI/ML
Vertex AI
OpenAI
Anthropic
Frontend
Next.js
React.js
DevOps
Azure
Kubernetes
Apply
$60k – $84k per year • In office • Internship • San Francisco
Python
JavaScript
TypeScript
AI/ML
Copilot
Cursor
Claude
Claude Code
Model Context Protocol
Vertex AI
AI Agents
LLM
OpenAI
Anthropic
LLM Guardrails
Tool Use
Frontend
Next.js
React.js
DevOps
Azure
AWS
Kubernetes
Apply
See all jobs
This is one of many
795,157 more open roles from verified company boards, updated every day.