725,615open jobs
43,268companies
103,521added this week
Browse all
Salary
$149k – $311k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Aug 13, 2026.

Overview
Company
Impact
Profile match
The continuous-improvement stack for agents. Monitor and improve your agent's behavior at scale.

Product Engineer - Agents Job Description

The Role

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:

  • We ingest everything your agents do in production: traces, tool calls, decisions, outcomes

  • Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals

  • Teams close the loop, shipping agent improvements validated against real production evidence

You'll build the product experiences that make this loop legible, and you'll build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.

What You Will Accomplish

  • Judgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.

  • Verification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes.

  • Agent investigation interfaces: Design how engineers understand what their agents did and why. Long traces, tool calls, decisions, failures. What does debugging look like when the "program" is a reasoning loop? How do you make a thousand-step trajectory legible in minutes?

  • Swarm UX: A hundred parallel investigations is useless if engineers can't follow them. Design how humans watch a swarm work, redirect investigators that go down the wrong path, and consume findings without reading a hundred reports.

  • The improvement loop: Build the workflows that turn production trajectories into datasets, judges, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion.

  • The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.

  • Judgment everywhere agents are built: An SDK and terminal-first experience so Claude Code, Codex, and OpenCode sessions can summon Judgment as a subagent mid-development.

What You'll Bring

  • Experience building and scaling end-to-end production systems, from data layer to UI

  • Strong technical problem-solving skills, especially in fast-changing, ambiguous environments

  • A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship

  • Hands-on experience building with LLMs or agents, or the drive to get there fast

  • Comfort working directly with customers to understand their needs and solve real-world problems

  • Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
725,615 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$22k – $55k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Hyderabad
Python
SQL
Apex
Apex
Copado
AI/ML
AI Agents
LLM
Agentforce
DevOps
Rest API
GCP
GitHub Actions
CI/CD
AWS
SOAP
Cybersecurity
ISO 27001
Checkov
PCI DSS
SOC 2
SIEM
Analytics
Tableau
ETL/ELT
QA
Selenium
Apply
$32k – $64k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Moscow
Python
SQL
Python
pySpark
AI/ML
Spark
AI Agents
NLP
Transformers
PyTorch
Hugging Face
Machine Learning
Apply
$11k – $26k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Hyderabad
AI/ML
AI Agents
Analytics
Power BI
Management
Service Desk
Microsoft Office
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • Mumbai
AI/ML
Claude
AI Agents
Cybersecurity
Zero Trust
Defense in Depth
Apply
In office • 5+ years exp
AI/ML
Cursor
Claude
FastAI
AI Agents
PyTorch
Google AI Studio
Lovable
Frontend
Sass
Design
Figma
Apply
$65k – $142k per year (Estimated) • In office • Full-Time • 1+ year exp • San Francisco
Apply
$152k – $316k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
AI Agents
Apply
Backend/Infra Engineer 3 months ago
$141k – $272k per year (Estimated) • In office • Full-Time • San Francisco
Go
JavaScript
Go
Temporal
Databases
Databricks
ClickHouse
RabbitMQ
Apache Kafka
AI/ML
Model Context Protocol
Dagster
Prefect
AI Agents
Flink
LLM
Ray
Anomaly Detection
Context Engineering
Frontend
Next.js
React.js
Management
Slack
Apply
Applied AI Engineer 8 months ago
$176k – $379k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
Reinforcement Learning
Post-training
Machine Learning
Apply
$46k per year • In office • Full-Time • San Francisco
Management
Agile
Apply
$197k – $314k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Washington
AI/ML
Cursor
Claude Code
Function Calling
AI Agents
Langfuse
LangSmith
LLM
Braintrust
OpenAI Codex
Agentforce
Replit
LLM Evaluation
Tool Use
Management
Slack
Apply
$229k – $345k per year • Equity • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco • Milpitas • Mountain View • Los Altos • Fremont
DevOps
BGP
OSPF
Apply
$139k – $235k per year • In office • Full-Time • San Francisco
Apply
$139k – $299k per year (Estimated) • Remote • Full-Time • San Francisco
AI/ML
Cursor
Claude Code
Time Series Forecasting
Physical AI
Machine Learning
Apply
See all jobs
This is one of many
725,615 more open roles from verified company boards, updated every day.