368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$160k – $210k per year
Location
In office (San Francisco)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
OfficeHours is an expert network company headquartered in San Francisco, California, and founded in 2020 by Erez Arnon, Joe Kim, and Patrick Reynolds. The platform connects investors, consultants, and operators with vetted subject matter experts for paid calls, handling scheduling, compliance screening, and payment. It is used by venture capital firms, consultancies, and startups doing diligence in technology, software, and healthcare markets.

Software Engineer, Benchmarking (SF, NYC, or Remote)

About Us

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. Experts earn income by sharing their knowledge through advisory work, projects, and AI model training. Our platform handles the complexities behind the scenes- screening, compliance, scheduling, and payments-so knowledge sharing stays focused on meaningful insights and real impact.

We're a hyper-growth and profitable company, quickly expanding our expert network, launching new offices, and new products. We are headquartered in San Francisco, with offices in Brooklyn and Bangalore. Our customers include the fastest-growing digital health companies, technology companies, institutional investment firms, consulting firms and AI Labs. We are backed by top marketplace investors and operators of companies like DoorDash, Airbnb, Affirm.

What we believe

Human knowledge is the world's most valuable asset. And yet, despite being more interconnected than ever, most knowledge still remains stuck in our heads, inaccessible and underutilized. Our vision is to make human knowledge easily accessible and infinitely scalable by building tools for the new age knowledge economy.

About the role

We're looking for a Software Engineer to build and run the platform behind our AI model evaluations. You'll work closely with our research team to prepare benchmark datasets, build the pipelines and environments our evaluations run in, and turn results into published output.

Our researchers design the methodology. You'll turn it into systems that run consistently and reproducibly, so results stay comparable across models, agent scaffolds, and time. Evaluations draw on the knowledge domains our expert network covers.

What you'll do

  • Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results.

  • Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable.

  • Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use.

  • Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance.

  • Develop the scoreboard and leaderboard: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance.

  • Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time.

  • Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish.

  • Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.

What you bring

  • Solid engineering skills: 4+ years of professional experience building and maintaining complex systems, with strong Python. You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure.

  • Data rigor: Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable.

  • Comfort with containers and environments: Experience with Docker and building reproducible execution environments.

  • Collaborative: You work well alongside researchers and scientists and can translate their methodology into working systems.

Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus.

Tech Stack

  • Evaluations: Python, model APIs, agent/evaluation frameworks, custom evaluation tooling

  • Models: APIs from the major AI providers, terminal agents, and open-source models via the Hugging Face ecosystem and PyTorch

  • Environments: Docker

  • Publishing: React, Next.js, Tailwind

  • Workflow: GitHub, Slack, Notion, Linear

Bonus Experience

  • Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work

  • Experience with agentic, multi-turn, long-context, or tool-use evaluation

  • Experience validating LLM-as-judge or rubric-based grading setups

  • Background or strong interest in a scientific or technical domain

  • Experience building data-heavy dashboards, leaderboards, or visualizations

  • Open-source contributions or published work related to benchmarks and measurement

Benefits + Perks

  • Competitive salary and equity

  • Medical, dental, and vision coverage

  • 401(k)

  • Monthly wellness and fitness stipend

  • Paid time off policy, along with company holidays

  • Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City)

  • Parent-friendly policies, remote flexibility, and paid family leave

Pay Transparency Notice

Full-time offers include base salary, equity, and benefits.

Pay range: $160,000-$210,000, based on seniority, relevant experience and location

This role can be fully remote or hybrid out of our SF or NYC offices.

Don't meet every single requirement? Studies have shown that some candidates, especially underrepresented groups such as women and people of color, are less likely to apply to jobs unless they meet every single qualification. At Office Hours we believe in building a diverse and inclusive workplace, so if you're excited about this role but don't meet every qualification in the job description, we still encourage you to apply. You could still be the right candidate for this or other roles at Office Hours!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
up to $48k per year (net) • Remote/Hybrid • Moscow
Node JS
Python
TypeScript
JavaScript
Node JS
Nest.JS
Python
Django
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
PostgreSQL
Qdrant
RabbitMQ
Redis
AI/ML
Chain-of-Thought
Claude
Claude Code
Copilot
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
RAG
Anthropic
Function Calling
OpenAI
Structured Outputs
Frontend
GraphQL
Next.js
React.js
Redux
Redux Toolkit
Zustand
Mobile
State Management
DevOps
AWS
CI/CD
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Kubernetes
Prometheus
Yandex Cloud
GitHub
GitLab
Apply
$24k – $64k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Chennai
Python
Scala
SQL
TypeScript
JavaScript
Java
Python
pySpark
Java
Spring Boot
Databases
Apache Kafka
Databricks
AI/ML
AI Agents
Copilot
Google ADK
LLM
NLP
Prompt Engineering
Spark
Devin
Model Context Protocol
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
OpenShift
GitHub
Analytics
ETL/ELT
Apply
Backend Engineer 1 day ago
$45k – $57k per year • In office • Full-Time • 3+ years exp • Tokyo
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
Frontend
Next.js
React.js
DevOps
Amazon EC2
AWS
AWS CDK
CI/CD
Docker
GitHub Actions
Vercel
Amazon CloudWatch
Amazon S3
GitHub
IAM
Management
Linear
Apply
$23k – $62k per year (Estimated) • In office • Full-Time • 10+ years exp • Gurgaon
C#
SQL
Databases
Apache Kafka
Redis
DevOps
Azure
CI/CD
Docker
gRPC
Kubernetes
Apply
$11k – $24k per year (Estimated) • Remote/Hybrid • Saint Petersburg
Bash
PowerShell
DevOps
CI/CD
Docker
Git
GitLab CI
Hyper-V
kubectl
Kubernetes
Proxmox VE
VMWare
GitLab
Apply
$165k – $185k per year • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • New York
Node JS
TypeScript
JavaScript
Databases
OpenSearch
PostgreSQL
AI/ML
Knowledge Graph
Frontend
Next.js
React.js
shadcn/ui
Tailwind CSS
Radix UI
DevOps
AWS
Datadog
Docker
Kubernetes
GitHub
Design
Figma
Management
Linear
Notion
Slack
Apply
$170k – $220k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • San Francisco
Node JS
TypeScript
JavaScript
Databases
ElasticSearch
RabbitMQ
AI/ML
Embeddings
LLM
Reranking
Semantic Search
AI Agents
Knowledge Graph
Recommender Systems
Semantic Search
Frontend
Next.js
React.js
Storybook
Tailwind CSS
DevOps
AWS
Docker
Kibana
Kubernetes
Terraform
Vector
GitHub
Design
Figma
Management
Notion
Slack
Marketing
Amplitude
QA
Sentry
Apply
$170k – $220k per year • In office • Full-Time • 6+ years exp • San Francisco
Node JS
TypeScript
JavaScript
Databases
ElasticSearch
RabbitMQ
AI/ML
Embeddings
LLM
Reranking
Semantic Search
AI Agents
Knowledge Graph
Recommender Systems
Semantic Search
Frontend
Next.js
React.js
Storybook
Tailwind CSS
DevOps
AWS
Docker
Kibana
Kubernetes
Terraform
Vector
GitHub
Design
Figma
Management
Notion
Slack
Marketing
Amplitude
QA
Sentry
Apply
Senior IOS Engineer 1 month ago
$180k – $200k per year • Equity • In office • Full-Time • San Francisco
Node JS
Swift
TypeScript
JavaScript
Swift
Swift Concurrency
Databases
OpenSearch
AI/ML
Claude
Claude Code
LLM
AI Agents
OpenAI Codex
Mobile
Fastlane
SwiftUI
DevOps
CI/CD
GitHub Actions
GitHub
Design
Figma
Management
Linear
Notion
Slack
Apply
Platform Engineer 1 month ago
$160k – $180k per year • Remote/Hybrid • Full-Time • San Francisco
Bash
Node JS
Python
TypeScript
JavaScript
Databases
ElasticSearch
OpenSearch
PostgreSQL
DevOps
Amazon EKS
ArgoCD
AWS
CI/CD
Datadog
Docker
Git
GitHub Actions
GitOps
Helm
Kubernetes
Kustomize
Terraform
Vector
GitHub
IAM
Cybersecurity
ISO 27001
SOC 2
Least Privilege
QA
Sentry
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.