368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$175k – $300k per year
Location
In office (San Francisco)
Seniority
Staff
Overview
Company
Impact
Profile match
Patronus AI develops simulation research and infrastructure to accelerate progress toward human-aligned AGI.

About Patronus AI

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation likeFinanceBench,Lynx,SimpleSafetyTests,CopyrightCatcher,Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

As a Researcher at Patronus AI, you will own and drive foundational research that defines how agentic AI systems are trained, evaluated, and improved. You will work at the intersection of reinforcement learning, simulations, and scalable oversight, building systems that directly influence how frontier models are developed, stress-tested, and deployed.

This is a highly autonomous role. You will tackle open-ended research questions surrounding agent simulations and translate them into rigorous experiments, benchmarks, environments, and production systems. You will work across areas including reward design, tool simulations, agent cognition, behavior analysis, and scalable oversight, helping shape the industry standard for robust, high-quality environments.

Your work will inform how frontier labs design, train, evaluate, and improve the next generation of agents for complex, long-horizon tasks, advancing our path toward safe, human-aligned general intelligence.

In this role, you will: 

  • Own ambitious research projects end-to-end, from identifying and formulating open-ended problems through experiment design, execution, analysis, and production impact.
  • Advance research in agent simulation, reinforcement learning, and scalable oversight, including agent cognition, behavior analysis, reward design, and new training methods.
  • Design state-of-the-art simulation and RL environments for training and evaluating frontier agents, spanning tools and actions, observations and state, trajectories, curricula, and reward systems.
  • Train models and experiment with post-training algorithms, including GRPO and SFT. Run ablations with open source models and understand the impact of distillation, COT reasoning, sparse and dense rewards and hyperparameters.
  • Develop methods to understand and improve agent behavior across complex, long-horizon tasks, including reasoning, planning, adaptation, generalization, and reward hacking.
  • Run rigorous experiments and turn findings into measurable outcomes, including new techniques, benchmarks, datasets, environments, platform capabilities, and research publications.
  • Build high-quality, reproducible research systems, writing production-level code and partnering closely with engineering and product to translate research into real-world systems.
  • Contribute to Patronus AI’s research direction and thought leadership, staying at the frontier of the field, collaborating with the research community, and publishing or open sourcing our work.

Qualifications

"The number one qualification to succeed in this machine learning course is gumption” - John Lafferty, CS Professor at Yale

Above all, we look for an eagerness to learn, passion for research, creativity in problem solving and a proactive mindset. You are a great fit if you have a background in the following:

  • An MS or PhD in Computer Science, Machine Learning, Statistics, Mathematics, or a related quantitative field.
  • Experience conducting independent research in reinforcement learning, NLP, agentic systems, evaluation, alignment, or related areas.
  • Demonstrated ability to take open-ended research problems from 0→1 and deliver high-impact outcomes.
  • Strong experimental skills, including experiment design, analysis, and interpretation of results.
  • Experience writing clean, reproducible research code in Python and modern machine learning frameworks.
  • Ability to execute quickly and independently with minimal guidance while maintaining a high bar for research quality.
  • Experience collaborating cross-functionally with research, engineering, and product teams.
  • Clear written and verbal communication skills, including the ability to explain complex technical ideas succinctly.
  • Strong integrity, good judgment, and respect for others.

To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance five days a week.

The expected base salary range for this role is $175,000 - $300,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

Benefits

  • Competitive salary and equity packages
  • 15 days of paid vacation per annum
  • Parental & sick leave
  • Health, dental, and vision insurance plans
  • 401(k) plan + matching
  • In-office lunch & dinner
  • Sponsored personal tax accounting 
  • Whoop band
  • Monthly meal stipend
  • Monthly health and wellness stipend
  • Equinox membership
  • Fun global offsites!

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

By clicking ‘Apply’, you agree to Greenhouse's Terms of Service  and Privacy Policy.

By clicking 'Apply', you agree to Patronus AI, Inc. Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$47k – $102k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
C++
Java
Python
SQL
C#
TypeScript
JavaScript
Java
Maven
C#
.NET
Databases
Apache Kafka
MySQL
AI/ML
ChatGPT
Copilot
Frontend
Angular
DevOps
CI/CD
Docker
Jenkins
Kubernetes
Prometheus
Apply
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
$163k – $434k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
$32k – $83k per year (Estimated) • In office • Full-Time • 5+ years exp • Gurgaon
Python
Python
pySpark
Databases
Microsoft Fabric
AI/ML
Spark
DevOps
Azure
Apply
Remote • Full-Time • 3+ years exp • Prague
Bash
PowerShell
Python
DevOps
Splunk
Cybersecurity
Crowdstrike
IBM QRadar
Apply
$125k – $250k per year • Equity • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Claude
Reinforcement Learning
Apply
$200k – $300k per year • Equity • In office • 8+ years exp • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Reinforcement Learning
Apply
$150k – $300k per year • Equity • In office • 4+ years exp • Bachelor's Degree • San Francisco
Go
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
Reinforcement Learning
Function Calling
OpenAI Codex
Post-training
Frontend
Next.js
React.js
DevOps
Kubernetes
Platform Engineering
QA
Playwright
Selenium
Apply
$125k – $200k per year • Equity • In office • 3+ years exp • Bachelor's Degree • San Francisco
Python
TypeScript
JavaScript
Databases
SQLite
AI/ML
AI Agents
Claude
Claude Code
Reinforcement Learning
OpenAI Codex
Frontend
Next.js
React.js
QA
Playwright
Selenium
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$293k – $385k per year • In office • Full-Time • San Francisco
AI/ML
OpenAI
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.