413,239open jobs
14,163companies
59,792added this week
Browse all
Salary
$150k – $350k per year
Location
In office (San Francisco, New York)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Overview

About Arcanum Labs

Arcanum Labs works on the data and infrastructure problems that sit behind frontier model training and deployed AI agents. We partner directly with labs on post-training research & infrastructure, and we deploy models and agents for large enterprises and government customers.

We're under four months into operation and already profitable: over seven figures in realized revenue to date, with more contracted. We haven't taken on a priced round yet but have early capital from a small group of institutional and angel investors including senior researchers from Anthropic & OpenAI. We're a team of ten including engineers and early employees from Glean, Palantir, MSL, Mercor, Nvidia, and other top startups, and we're hiring rapidly.

About the Role

As an Evals Researcher at Arcanum, you'll build the evaluation systems that determine which data is worth training on and whether our models and agents are actually getting better. This role sits at the center of our post-training loop: before we spend compute on a dataset, you're the one deciding if it's worth it; after we ship a model change, you're the one who can say whether it actually helped.

Evals work here spans reward design, automatic scoring at scale, and staying ahead of drift as customer traffic and task shapes evolve - this isn't a fixed benchmark you maintain once and forget.

In this role, you will:

  • Rank training value before spending compute. Work out which tasks in a dataset are worth training on, and build a system for ranking every task in a dataset by expected training value - before compute gets spent on it.

  • Build scoring that runs at scale, cheaply. Design automatic scoring cheap enough to run constantly, tuned to each customer's definition of a good outcome rather than generic correctness.

  • Catch drift before it's a problem. Notice when real customer requests have moved far enough from your existing test set that it no longer describes the job, and rebuild evals to match.

Requirements

Must Haves

  • You've built RL data or done RL research - hands-on, not adjacent.

  • Real, substantive opinions on reward design: what makes a checker trustworthy, and how models learn to game them.

  • Comfortable owning ambiguous evaluation problems end-to-end - building the first version, inspecting the data, and iterating until the system is actually useful.

  • Able to work 6 days a week, in-person or hybrid in SF.

Nice to Haves

We're not expecting one person to check every box below - these are areas that make a candidate a stronger fit, not requirements.

  • Experience with LLM-as-judge systems, automated graders, or rubric design for model outputs.

  • Background in applied ML research, research engineering, or data science at a lab or fast-moving startup.

  • Experience building eval infrastructure that's tied directly to a training pipeline (not just offline benchmarking).

Compensation & Benefits

Salary Range: $150,000 - $350,000 USD, depending on experience and seniority

Equity: Meaningful equity grants for early team members

Benefits: Health, dental, and vision coverage

Our Process

  • Behavioral and technical screen

  • Work trial

  • References

  • Offer

We move fast - most candidates go from first conversation to offer in about four weeks when there's mutual interest and momentum.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
413,239 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$34k – $85k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Hyderabad
JavaScript
TypeScript
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
AI/ML
Model Context Protocol
AI Agents
OpenAI
Frontend
Angular
Bootstrap
Mobile
Material Design
DevOps
Azure
CI/CD
Docker
Kubernetes
Analytics
ETL/ELT
Apply
SR AI ENGINEER, SMAI 10 hours ago
$34k – $85k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Hyderabad
JavaScript
TypeScript
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
AI/ML
Model Context Protocol
AI Agents
OpenAI
Frontend
Angular
Bootstrap
Mobile
Material Design
DevOps
Azure
CI/CD
Docker
Kubernetes
Analytics
ETL/ELT
Apply
AI Platform Engineer 10 hours ago
$48k – $60k per year • Remote/Hybrid • 4+ years exp • Milan • Padua
Python
Python
FastAPI
AI/ML
Model Context Protocol
AI Agents
LLM
DevOps
Platform Engineering
Apply
$82k – $110k per year • In office • Full-Time • 6+ years exp • PhD • Rome
Python
JavaScript
TypeScript
Apex
Databases
Snowflake
Databricks
AI/ML
Cursor
LangChain
Claude
LlamaIndex
Prompt Engineering
AI Agents
Agentforce
LLM Guardrails
Multi-Agent Systems
DevOps
CI/CD
Apply
$39k – $104k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Guadalajara
Java
TypeScript
SQL
Java
Spring Boot
AI/ML
Copilot
Cursor
Claude
Claude Code
AI Agents
Agentic Workflows
DevOps
Azure
CI/CD
Docker
Kubernetes
GitHub
Apply
$150k – $250k per year • Equity • In office • Full-Time • San Francisco • New York
AI/ML
AI Agents
OpenAI
Anthropic
Post-training
Apply
$60k – $100k per year • In office • Full-Time • 2+ years exp • San Francisco
Apply
$77k – $173k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Apply
$140k – $210k per year • Equity • In office • Full-Time • 4+ years exp • San Francisco
Python
Go
TypeScript
DevOps
GCP
CI/CD
Kubernetes
Management
Stripe
Apply
$200k – $250k per year • Equity 1–2% • In office • Full-Time • 3+ years exp • San Francisco
Apply
$160k – $250k per year • Equity 1–2% • In office • Full-Time • 3+ years exp • San Francisco
Python
AI/ML
PyTorch
DevOps
AWS
Apply
See all jobs
This is one of many
413,239 more open roles from verified company boards, updated every day.