699,757open jobs
41,095companies
105,247added this week
Browse all
Salary
$175k – $300k per year
Location
In office (San Francisco)
Seniority
Staff
Overview
Company
Impact
Profile match
Patronus AI develops simulation research and infrastructure to accelerate progress toward human-aligned AGI.

About Patronus AI

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation likeFinanceBench,Lynx,SimpleSafetyTests,CopyrightCatcher,Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

As a Member of Technical Staff - Engineering at Patronus AI, you’ll build the systems, infrastructure, and products that power our simulation research and agent training work.

This is a broad engineering role for people who like operating across boundaries. Depending on the problem, you might build a realistic RL environment end-to-end, design infrastructure for running thousands of agent trajectories, ship internal platforms used by researchers, deploy and serve models, or build AI-powered developer tools that make the entire team faster.

You’ll work across the stack - frontend interfaces, backend services, infrastructure, and ML/agent integrations - and partner closely with researchers and engineers to turn ambiguous problems into robust systems.

We’re looking for engineers with high ownership, strong product and technical judgment, and the ability to learn unfamiliar areas quickly. You don't need to be an expert in every part of the stack. You should have meaningful depth somewhere and the curiosity and engineering range to work wherever the problem requires.

In this role, you will:

  • Build agent environments and simulations end-to-end, including frontend interfaces, backend services, APIs, data models, tools, and realistic workflows used to train and evaluate AI agents.
  • Build the infrastructure that powers our agent gym, including orchestration, sandboxing, packaging, benchmarking, and systems for running environments across heterogeneous targets.
  • Develop internal platforms and developer tools used by researchers and engineers, from backends and dashboards to CLIs, SDKs, review agents, codegen helpers, and workflow automations.
  • Build and operate ML infrastructure, including model deployment and serving, evaluation systems, GPU workloads, and the services that make compute accessible to the broader team.
  • Own systems from ambiguous idea through production. Define the problem, make architectural decisions, implement the solution, instrument it, and iterate based on how it performs in practice.
  • Think deeply about correctness and failure modes. Design for edge cases, adversarial agent behavior, reproducibility, observability, and the messy realities of production systems.
  • Partner closely with researchers to productionize experiments and build the software and infrastructure needed to turn research ideas into scalable systems.
  • Be a power user of AI coding tools like Claude Code, Codex, Cursor, and similar tools - and build new tooling and automations on top of them when existing workflows aren't good enough.
  • Move quickly without sacrificing judgment. Make pragmatic decisions about what needs to be robust today, what can evolve later, and where technical investment will create leverage for the team.

Qualifications

“The number one qualification to succeed in this machine learning course is gumption” - John Lafferty, CS Professor at Yale

We're looking for a hands-on generalist who ships. Above all, we value strong product instincts, the ability to operate independently in ambiguity, and genuine curiosity about agents and the frontier of AI.

The list below is broad, and we don't expect every candidate to tick every box. What matters more is that you ship, stay curious, and use AI tools fluently enough to close gaps in days or weeks, not months. The team leans on AI tooling heavily to ramp on unfamiliar areas, and we expect anyone joining to do the same.

You are a strong fit if you have:

  • A track record of shipping non-trivial software end-to-end as an individual contributor, ideally at a startup or on a small, high-velocity team.
  • Strong engineering fundamentals and meaningful depth in at least one of backend/infrastructure, frontend/product engineering, or ML systems, with the ability and desire to work across boundaries.
  • Experience building production systems in languages such as Python, Go, and/or TypeScript, and the ability to become productive quickly in an unfamiliar stack.
  • Experience working with modern LLMs and agents at the application level - including concepts like tool calling, agent loops, context management, harnesses, and evaluation.
  • Strong engineering judgment around system design, correctness, reliability, failure modes, edge cases, and operational complexity.
  • High independence. You can take an ambiguous goal, determine what needs to be built, find the people or information necessary to unblock yourself, and ship without requiring the work to be fully pre-scoped.
  • Fluency with modern AI coding tools and a strong instinct for where AI can automate or accelerate engineering workflows.
  • A BS, MS, or PhD in Computer Science, Machine Learning, Software Engineering, or a related quantitative field - or equivalent experience.

Depending on your area of depth, you may also have experience with:

  • Building complex full-stack products using technologies like React, TypeScript, Next.js, Python, relational databases, and modern API frameworks.
  • Building developer platforms, distributed systems, orchestration systems, sandboxes, or internal infrastructure.
  • Reinforcement learning environments, agent evaluation, verifiers, reward models, or benchmarking infrastructure.
  • Deploying and serving ML models using managed inference providers or self-operated GPUs.
  • GPU infrastructure, workload schedulers, Kubernetes, containers, CI/CD, and cloud infrastructure.
  • MLOps/LLMOps tooling, experiment tracking, model registries, and production observability.
  • Browser automation tools such as Playwright or Selenium.
  • Building or deeply using complex enterprise software and understanding the operational edge cases that accumulate in real-world systems.

We don't expect you to check every box. We're more interested in engineers who have demonstrated exceptional depth somewhere, can learn quickly, and are excited to take ownership of whatever technical problem matters most.

To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance 5 days a week.

The expected base salary range for this role is $175,000 - $300,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

By clicking ‘Apply’, you agree to Greenhouse's Terms of Service  and Privacy Policy.

By clicking 'Apply', you agree to Patronus AI, Inc. Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
699,757 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$100k – $180k per year • In office • Full-Time • 3+ years exp • Copenhagen
TypeScript
AI/ML
Cursor
Lovable
Design
Figma
Webflow
Apply
$100k – $180k per year • In office • Full-Time • San Francisco
Python
TypeScript
AI/ML
Computer Vision
Frontend
Tailwind CSS
Apply
$83k – $115k per year • In office • Full-Time • 2+ years exp • Berlin
Python
JavaScript
TypeScript
Node JS
AI/ML
Embeddings
Cohere SDK
Gemini
LLM
Semantic Search
OpenAI
Semantic Search
Agentic Workflows
Frontend
React.js
DevOps
Azure
Apply
Head of Engineering 2 hours ago
$167k – $241k per year • In office • Full-Time • Berlin
Python
JavaScript
TypeScript
Node JS
Node JS
Fastify
Databases
PostgreSQL
AI/ML
OpenAI
Semantic Search
DevOps
GCP
Datadog
Azure
AWS
Apply
$92k – $138k per year • In office • Full-Time • 3+ years exp • Berlin
Python
JavaScript
TypeScript
SQL
Node JS
Node JS
Fastify
Prisma
AI/ML
Prefect
Multimodal AI
AI Agents
Cohere SDK
Langfuse
Gemini
LLM
OpenAI
OCR
Semantic Search
DevOps
Datadog
Azure
AWS
Amazon S3
Analytics
ETL/ELT
Apply
$38k – $85k per year (Estimated) • Remote/Hybrid • 3+ years exp • PhD
Python
JavaScript
TypeScript
Databases
SQLite
AI/ML
Claude Code
Reinforcement Learning
AI Agents
OpenAI Codex
Reward Modeling
Machine Learning
Frontend
Next.js
React.js
QA
Selenium
Playwright
Apply
$175k – $300k per year • Equity • In office • Master's Degree • San Francisco
Python
AI/ML
Reinforcement Learning
Knowledge Distillation
AI Agents
NLP
SFT
GRPO
Post-training
Model Distillation
Machine Learning
Apply
$115k – $225k per year • Equity • In office • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Machine Learning
Apply
$200k – $300k per year • Equity • In office • 8+ years exp • Bachelor's Degree • San Francisco
AI/ML
Reinforcement Learning
AI Agents
Agentic Workflows
Machine Learning
Apply
$125k – $250k per year • Equity • In office • Bachelor's Degree • San Francisco
AI/ML
Claude
Reinforcement Learning
AI Agents
Machine Learning
Apply
$59k – $126k per year (Estimated) • Remote/Hybrid • Contractor • 1+ year exp • Bachelor's Degree • San Francisco
Management
Microsoft Office
Apply
$97k – $188k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • New York • Boston • San Francisco • Washington • Seattle
Management
Outlook
Microsoft Office
Apply
$133k – $338k per year • Remote • Full-Time • 12+ years exp • Associate's Degree • New York • Milwaukee • Dallas • Columbus • Kirkland
Python
Databases
Neo4j
Amazon Neptune
AI/ML
Airflow
Prompt Engineering
Multimodal AI
AI Agents
NLP
TensorFlow
PyTorch
LLM
Knowledge Graph
Machine Learning
DevOps
GCP
Azure
AWS
Analytics
ETL/ELT
Apache NiFi
Apply
$123k – $317k per year • In office • Full-Time • 10+ years exp • Master's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Management
Agile
Marketing
Salesforce
Apply
$99k – $205k per year (Estimated) • In office • Full-Time • 3+ years exp • High School Diploma • San Francisco
DevOps
SLI/SLO/SLA
Windows
Management
ITIL
Service Desk
Apply
See all jobs
This is one of many
699,757 more open roles from verified company boards, updated every day.