368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$200k – $300k per year
Location
In office (San Francisco)
Seniority
Senior · 8+ years exp
Overview
Company
Impact
Profile match
Patronus AI develops simulation research and infrastructure to accelerate progress toward human-aligned AGI.

About Patronus AI

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation likeFinanceBench,Lynx,SimpleSafetyTests,CopyrightCatcher,Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

As a Senior Product Manager at Patronus AI, you will own the product behind our production line: the platform, infrastructure, and internal tools that let a small team turn frontier-lab demand into delivered RL environments. That surface includes our hosted environment platform, agentic tooling for building and QA-ing application clones, expert onboarding, delivery tracking, and the spec pipeline that feeds engineering.

We run on a simple operating rule: if an agent can do it, don't assign it to anyone else. Your users are as often AI agents as they are people, and the products you spec should default to agentic workflows with humans as the exception. PMing for agents as first-class users is most of what makes this role interesting.

This is not a backlog-administration role. We believe you should not manage a product you don't use. Today, too many of our PRDs go stale the moment they're written, and engineers route around them; the product function sits on the critical path for the company's OKR roadmap. Your job is to make specs the thing engineers reach for first: versioned, frozen at build kickoff, scoped to an MVP with an explicit cut list, and grounded in real usage rather than opinion.

Your work will help frontier labs stress-test and improve the next generation of AI agents, advancing progress toward safe, human-aligned general intelligence.

In this role, you will:

  • Own the roadmap for the internal platform and tools end-to-end - the environment platform, agentic build and QA tooling, expert onboarding, delivery tracking, and the spec pipeline. Sequence it against the company OKR roadmap so product is never the thing engineering is waiting on.
  • Write PRDs engineers trust and build from. Versioned and frozen at kickoff, explicit about what's in the MVP and what's deliberately cut, and honest about buildability constraints (e.g., which classes of apps our agentic build tooling can and cannot handle today).
  • Work with, identify the pain points of, and brainstorm solutions with internal customers - environment engineers, SME ops, QA, and GTM. Then turn what you learn into a ranked backlog and tell them plainly what's not getting built and why.
  • Make usage data the basis for decisions. Instrument the tools, watch what agents and people actually do in them, and kill features that don't earn their keep.
  • Close the loop between QA and spec. When automated QA keeps surfacing the same class of missing feature, that's a spec failure - fix the pipeline, not just the ticket.
  • Apply the "if an agent can do it" rule to your own function: build agentic workflows for spec generation, review, and versioning, and spend your human hours on the judgment calls only you can make.
  • Ruthlessly scope. Ship the smallest thing that changes an internal customer's week, publish the cut list, and iterate from evidence.
  • Stay hands-on. Run agents in the environments, poke at the tools, prototype with AI, and build small evals to confirm whether a feature actually worked.

Qualifications

"The number one qualification to succeed in this machine learning course is gumption" - John Lafferty, CS Professor at Yale

Above all, we look for a proactive mindset, willingness to learn, unlimited energy, and relentless optimism. You are a great fit if you have a background in the following:

  • BS, MS, or equivalent experience in Computer Science, Engineering, or another technical / quantitative field, with 8+ years as a product manager.
  • Experience PMing platform, infrastructure, or internal tools - products whose customers are engineers and internal teams rather than external end users. This is required, not preferred.
  • A track record of PRDs that engineers actually built from - and of shipping when priorities moved after the spec was written, without letting quality or trust in the spec slip.
  • Strong technical fluency, including daily use of AI tools, comfort reading or reviewing code, and the ability to pull and analyze your own usage data. You should be able to go deep enough into the work to have credibility with engineers.
  • Evidence of ruthless prioritization: MVPs shipped, cut lists published, and features you killed with the reasoning to back it up.
  • Clear written and verbal communication, including the ability to turn messy stakeholder asks into a spec with crisp requirements, owners, and dates.
  • Strong eye for quality and detail, with a bias toward catching gaps, inconsistencies, and subtle failure modes before engineering does.

Nice to have:

  • Experience with reinforcement learning, agent evaluation, RL environments, or human-data / SME-sourced data pipelines.
  • Experience building products whose primary users are AI agents rather than humans.
  • 0→1 experience at a fast-growing startup.

To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance 5 days a week.

The expected base salary range for this role is $200,000 - $300,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

Benefits

  • Competitive salary and equity packages
  • 15 days of paid vacation per annum
  • Parental & sick leave
  • Health, dental, and vision insurance plans
  • 401(k) plan + matching
  • In-office lunch & dinner
  • Sponsored personal tax accounting 
  • Whoop band
  • Monthly meal stipend
  • Monthly health and wellness stipend
  • Equinox membership
  • Fun global offsites!

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

By clicking ‘Apply’, you agree to Greenhouse's Terms of Service  and Privacy Policy.

By clicking 'Apply', you agree to Patronus AI, Inc. Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$44k – $115k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Mexico
Apex
JavaScript
Apex
Lightning Web Components
Salesforce CLI
AI/ML
AI Agents
ChatGPT
Claude
Cursor
Prompt Engineering
Agentforce
LLM Guardrails
Model Context Protocol
Frontend
Web Components
DevOps
CI/CD
Git
Rest API
GitHub
Cybersecurity
HIPAA
Marketing
Salesforce
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
In office • Internship • Master's Degree • Austin
C++
Python
C++
PyTorch C++
AI/ML
AI Agents
CUDA
CUDA Toolkit
LLM
NCCL
PyTorch
TensorRT
TensorRT-LLM
Triton
vLLM
Apply
$54k – $128k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Mexico City
C#
JavaScript
Python
TypeScript
C#
ASP.NET Core
Python
FastAPI
Databases
PostgreSQL
AI/ML
AI Agents
Anthropic
Embeddings
Function Calling
LLM
LLM Guardrails
Model Context Protocol
OpenAI
Semantic Search
Semantic Search
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
GitHub
GitHub Actions
Vector
Vercel
Apply
$129k – $231k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Atlanta
SQL
AI/ML
AI Agents
Claude
Claude Code
DevOps
Azure
Design
Figma
Apply
$175k – $300k per year • Equity • In office • Master's Degree • San Francisco
Python
AI/ML
Knowledge Distillation
NLP
Reinforcement Learning
GRPO
Post-training
SFT
AI Agents
Apply
$125k – $250k per year • Equity • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Claude
Reinforcement Learning
Apply
$150k – $300k per year • Equity • In office • 4+ years exp • Bachelor's Degree • San Francisco
Go
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
Reinforcement Learning
Function Calling
OpenAI Codex
Post-training
Frontend
Next.js
React.js
DevOps
Kubernetes
Platform Engineering
QA
Playwright
Selenium
Apply
$125k – $200k per year • Equity • In office • 3+ years exp • Bachelor's Degree • San Francisco
Python
TypeScript
JavaScript
Databases
SQLite
AI/ML
AI Agents
Claude
Claude Code
Reinforcement Learning
OpenAI Codex
Frontend
Next.js
React.js
QA
Playwright
Selenium
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 1 hour ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.