Overview
About Arcanum Labs
Arcanum Labs works on the data and infrastructure problems that sit behind frontier model training and deployed AI agents. We partner directly with labs on post-training research & infrastructure, and we deploy models and agents for large enterprises and government customers.
We're under four months into operation and already profitable: over seven figures in realized revenue to date, with more contracted. We haven't taken on a priced round yet but have early capital from a small group of institutional and angel investors including senior researchers from Anthropic & OpenAI. We're a team of ten including engineers and early employees from Glean, Palantir, MSL, Mercor, Nvidia, and other top startups, and we're hiring rapidly.
About the Role
As an Evals Researcher at Arcanum, you'll build the evaluation systems that determine which data is worth training on and whether our models and agents are actually getting better. This role sits at the center of our post-training loop: before we spend compute on a dataset, you're the one deciding if it's worth it; after we ship a model change, you're the one who can say whether it actually helped.
Evals work here spans reward design, automatic scoring at scale, and staying ahead of drift as customer traffic and task shapes evolve - this isn't a fixed benchmark you maintain once and forget.
In this role, you will:
Rank training value before spending compute. Work out which tasks in a dataset are worth training on, and build a system for ranking every task in a dataset by expected training value - before compute gets spent on it.
Build scoring that runs at scale, cheaply. Design automatic scoring cheap enough to run constantly, tuned to each customer's definition of a good outcome rather than generic correctness.
Catch drift before it's a problem. Notice when real customer requests have moved far enough from your existing test set that it no longer describes the job, and rebuild evals to match.
Requirements
Must Haves
You've built RL data or done RL research - hands-on, not adjacent.
Real, substantive opinions on reward design: what makes a checker trustworthy, and how models learn to game them.
Comfortable owning ambiguous evaluation problems end-to-end - building the first version, inspecting the data, and iterating until the system is actually useful.
Able to work 6 days a week, in-person or hybrid in SF.
Nice to Haves
We're not expecting one person to check every box below - these are areas that make a candidate a stronger fit, not requirements.
Experience with LLM-as-judge systems, automated graders, or rubric design for model outputs.
Background in applied ML research, research engineering, or data science at a lab or fast-moving startup.
Experience building eval infrastructure that's tied directly to a training pipeline (not just offline benchmarking).
Compensation & Benefits
Salary Range: $150,000 - $350,000 USD, depending on experience and seniority
Equity: Meaningful equity grants for early team members
Benefits: Health, dental, and vision coverage
Our Process
Behavioral and technical screen
Work trial
References
Offer
We move fast - most candidates go from first conversation to offer in about four weeks when there's mutual interest and momentum.

