368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$125k – $250k per year
Location
In office (San Francisco)
Overview
Company
Impact
Profile match
Patronus AI develops simulation research and infrastructure to accelerate progress toward human-aligned AGI.

About Patronus AI

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation likeFinanceBench,Lynx,SimpleSafetyTests,CopyrightCatcher,Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

As a Technical Program Manager at Patronus AI, you will run the production line that turns frontier-lab demand into delivered RL environments. Our customers commission simulations of real applications and workflows, and you own the path from signed work order through SME sourcing, environment build, task generation, QA, and delivery.

Like all cutting edge fields, the state of the industry and the work changes shape constantly. App lists get swapped after a work order is signed, difficulty definitions get renegotiated a week before delivery etc. Your job is to keep the plan honest through all of it: who owns what, what is due when, and whether a delivery is actually ready to ship. When scope moves, you move the plan with it.

This is not a coordination-only role. We believe you should not manage work you don't understand. You will stay in the details, reading QA feedback rows, poking at environments, and sanity-checking task quality, and you are expected to pitch in directly where it unblocks the team. You will own identifying opportunities for automation and build them yourself.

Your work will help frontier labs stress-test and improve the next generation of AI agents, advancing progress toward safe, human-aligned general intelligence.

In this role, you will:

  • Own delivery programs end-to-end: track work orders, deliverables, due dates, and owners across engineering, QA, and SME teams, and maintain the consolidated view the rest of the company plans against.
  • Manage scope changes mid-program. A previously committed app list doubles in size, or a customer redefines task difficulty on an active work order, and you update the plan, the pricing inputs, and our commitments without losing the thread.
  • Run the handoffs between GTM, SME sourcing, environment engineering, task generation, and QA. This is where work gets stranded today. Build the process that stops that: clear entry and exit criteria, and visibility into who is staffed on what.
  • Define what "ready to ship" means and hold the line on it. QA tickets are green before anything is marked complete, and tasks pass in the customer's harness, not just locally. We deliver work that already passes rather than delivering and triaging after.
  • Work directly with customers alongside account leads. Turn their asks into plans with clear outcomes, owners, and dates, and keep them current on progress, risks, and timeline confidence. Respond quickly; customers notice when we don't.
  • Get teams to time-bound their work and make estimates explicit ("expect this to take X hours, tell me if it takes longer"), then hold everyone, including yourself, to them.
  • Identify opportunities for automation and process improvements across the delivery pipeline and build them yourself.
  • Stay hands-on. Run agents in environments, review QA feedback and trajectories, and build small tools (trackers, scripts, Claude skills) that scale your own function.

Qualifications

"The number one qualification to succeed in this machine learning course is gumption" - John Lafferty, CS Professor at Yale

Above all, we look for a proactive mindset, willingness to learn, unlimited energy, and relentless optimism. You are a great fit if you have a background in the following:

  • BS, MS, or equivalent experience in Computer Science, Engineering, or another technical / quantitative field, with 3+ years as a technical program manager, delivery lead, engineering program manager, or in a similar role on complex, multi-team technical programs.
  • A track record of shipping when requirements move after kickoff, without letting quality or the customer relationship slip.
  • Strong technical fluency, including comfort using AI tools, reading or reviewing code, and analyzing data or model outputs. You should be able to go deep enough into the work to have credibility with engineers.
  • Excellent organization and execution skills, with the ability to manage tasks, timelines, quality reviews, customer requirements, and cross-functional stakeholders. You are the person who always knows the current state.
  • Clear written and verbal communication skills, including the ability to translate customer asks into concrete plans with owners and dates.
  • Strong eye for quality and detail, with a bias toward catching edge cases, inconsistencies, and subtle failure modes before the customer does.

Nice to have:

  • Experience with reinforcement learning, agent evaluation, RL environments, or human-data / SME-sourced data pipelines.
  • Experience delivering to frontier labs or other research-driven customers with short turnaround expectations.
  • Experience standing up QA or review processes for software or data deliverables.

To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance 5 days a week.

The expected base salary range for this role is $125,000 - $250,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

Benefits

  • Competitive salary and equity packages
  • 15 days of paid vacation per annum
  • Parental & sick leave
  • Health, dental, and vision insurance plans
  • 401(k) plan + matching
  • In-office lunch & dinner
  • Sponsored personal tax accounting 
  • Whoop band
  • Monthly meal stipend
  • Monthly health and wellness stipend
  • Equinox membership
  • Fun global offsites!

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

By clicking ‘Apply’, you agree to Greenhouse's Terms of Service  and Privacy Policy.

By clicking 'Apply', you agree to Patronus AI, Inc. Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$142k – $213k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Jersey City
Java
Python
SQL
TypeScript
JavaScript
Java
Spring Boot
Python
Asyncio
FastAPI
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Devin
Fine-tuning
Gemini
Google ADK
Hybrid Search
Knowledge Graph
LangChain
LangGraph
RAG
Frontend
Angular
DevOps
CI/CD
Docker
Kubernetes
Rest API
Apply
$121k – $171k per year • Remote/Hybrid • Full-Time • 8+ years exp • Mississauga
Java
SQL
TypeScript
JavaScript
Java
Spring Boot
Databases
Apache Kafka
Oracle
RabbitMQ
Trino
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Model Context Protocol
Frontend
Angular
React.js
DevOps
CI/CD
GitHub
WebSockets
Analytics
ETL/ELT
Management
Jira
Apply
Senior AI Architect 7 hours ago
$138k – $304k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
LangGraph
OpenAI
RAG
Spark
LangChain
DevOps
Azure
Apply
$118k – $168k per year • In office • Full-Time • Ottawa • Halifax
Databases
Snowflake
AI/ML
AI Agents
AutoGen
CrewAI
Human-in-the-Loop
LangChain
LLM
DevOps
AWS
Azure
Apply
$73k – $183k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Zug
Python
Python
FastAPI
Databases
PostgreSQL
Redis
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
DevOps
Amazon EKS
AWS
AWS CDK
CI/CD
Datadog
Kubernetes
OpenTelemetry
Platform Engineering
Apply
$175k – $300k per year • Equity • In office • Master's Degree • San Francisco
Python
AI/ML
Knowledge Distillation
NLP
Reinforcement Learning
GRPO
Post-training
SFT
AI Agents
Apply
$200k – $300k per year • Equity • In office • 8+ years exp • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Reinforcement Learning
Apply
$150k – $300k per year • Equity • In office • 4+ years exp • Bachelor's Degree • San Francisco
Go
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
Reinforcement Learning
Function Calling
OpenAI Codex
Post-training
Frontend
Next.js
React.js
DevOps
Kubernetes
Platform Engineering
QA
Playwright
Selenium
Apply
$125k – $200k per year • Equity • In office • 3+ years exp • Bachelor's Degree • San Francisco
Python
TypeScript
JavaScript
Databases
SQLite
AI/ML
AI Agents
Claude
Claude Code
Reinforcement Learning
OpenAI Codex
Frontend
Next.js
React.js
QA
Playwright
Selenium
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 1 hour ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.