373,573open jobs
9,691companies
51,475added this week
Browse all
Salary
$250k – $450k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Infrastructure for understanding AI Infrastructure for understanding AI. Transluce is a non-profit research lab building the public tech stack for scalable oversight of AI.

Salary range: $250,000 - $450,000/year + benefits

Description: Transluce is a fast-moving nonprofit research lab building the public tech stack for AI evaluation and oversight. We have contributed foundational research to the study of AI agents and their behaviors, and are using these to study emerging issues in the honesty and alignment of AI agents.

About the role: As an AI Behavior Researcher, you will lead projects to design and develop automated evaluations of frontier AI systems that are technically sophisticated, scientifically valid, and concretely impactful. You will conduct novel analyses of behaviors related to agentic honesty and alignment. Example behaviors of interest include misreporting results, falsely claiming success, evaluation awareness, and memetic effects within AI swarms.

As an early member of a highly collaborative team, you will learn and grow quickly, and work with our governance and infrastructure teams to scale your impact and technical reach.

Core responsibilities:

  • Develop novel, valid automated evaluations of AI agents' honesty and alignment.
  • Write code to implement and run automated evaluations, such as environment simulators or LLM-as-a-judge pipelines.
  • Design methods to improve the ecological of automated evaluations, and especially to measure which effects are increasing or decreasing as models become more capable.
  • Collaborate with our governance team to deliver high-impact evaluations for public policy.
  • Collaborate with scientists and research engineers to productionize best practices in AI behavior evaluation.

Minimum qualifications:

  • Expertise on quantitative generative AI evaluation and measurement. Good intuition about how to systematize and operationalize complex concepts and to work backwards from possible failure modes of agents.
  • Relevant experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
  • Proficiency in Python to implement analysis and evaluation tooling.
  • Meticulous, good experimental design, epistemic self-awareness and transparency.
  • Ability to iterate quickly and balance between scrappiness and thoroughness based on the impact needs of a project.
  • Strong communication skills, low ego, openness to giving and receiving feedback.

Preferred qualifications (not required):

  • Experience running automated evaluations at scale or in a production context.
  • Experience conducting controlled human subjects experiments to validate automated evaluation methods.
  • Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous stakeholder needs into concrete deliverables.
  • Experience and comfort using AI coding agents at work.

We are hiring at all levels of experience and would encourage those enthusiastic about the role who do not meet all of the qualifications to apply. We are located in San Francisco and excited to work together in-person. We are open to sponsoring international visas.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
373,573 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Senior GenAI Engineer 7 hours ago
In office • 6+ years exp • Bachelor's Degree
Python
C#
C#
.NET
Databases
Databricks
FAISS
Google BigQuery
Milvus
OpenSearch
Pinecone
Snowflake
Weaviate
AI/ML
AI Agents
AutoGen
AWS Bedrock
Copilot
CrewAI
Function Calling
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
LLM Guardrails
OpenAI
Prompt Engineering
RAG
Semantic Kernel
Vertex AI
DevOps
AWS
Azure
CI/CD
GCP
Git
Vector
Marketing
Salesforce
Apply
$31k – $79k per year (Estimated) • In office • 6+ years exp • Master's Degree • Bengaluru
Java
Python
Scala
AI/ML
AI Agents
Claude
Claude Code
Computer Vision
Cursor
Model Context Protocol
Multimodal AI
NLP
Spark
DevOps
Amazon S3
Analytics
A/B Testing
Apply
$122k – $166k per year • In office • Full-Time • 8+ years exp • Master's Degree • United States
Python
SQL
AI/ML
Accelerate
AI Agents
Amazon SageMaker
AWS Bedrock
NumPy
Pandas
PyTorch
RAG
Scikit-learn
TensorFlow
DevOps
AWS
AWS Lambda
Analytics
ETL/ELT
Apply
$103k – $207k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Wilmington • Philadelphia • Tysons • Pittsburgh
Python
SQL
AI/ML
Hadoop
Analytics
Tableau
Apply
$131k – $286k per year (Estimated) • Equity • Remote • Full-Time • New York
SQL
TypeScript
JavaScript
Databases
OpenSearch
PostgreSQL
AI/ML
Claude
Claude Code
LLM
LLM Evaluation
Frontend
React.js
DevOps
AWS
Apply
$250k – $450k per year • In office • Full-Time • San Francisco
Python
AI/ML
LLM
Apply
$250k – $500k per year • In office • Full-Time • San Francisco
AI/ML
Fine-tuning
Apply
VP of Engineering 4 months ago
$500k – $600k per year • In office • Full-Time • San Francisco
AI/ML
Claude
Apply
AI Behavior Engineer 5 months ago
$310k – $500k per year • In office • Full-Time • San Francisco
AI/ML
AI Agents
Apply
AI Systems Engineer 6 months ago
$350k – $600k per year • In office • Full-Time • San Francisco
Python
AI/ML
LLM
vLLM
AI Agents
Apply
$133k – $338k per year • Remote/Hybrid • Full-Time • 8+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
Knowledge Graph
DevOps
AWS
Azure
GCP
Platform Engineering
SLI/SLO/SLA
Apply
$208k – $332k per year • Remote/Hybrid • Full-Time • 9+ years exp • Bachelor's Degree • San Francisco
Apply
$197k – $314k per year • Remote/Hybrid • Full-Time • 8+ years exp • PhD • San Francisco
Apex
Java
Apex
MuleSoft
Java
Spring Boot
AI/ML
AI Agents
Model Context Protocol
RAG
Agentforce
DevOps
Docker
Kubernetes
Marketing
Salesforce
Apply
$220k – $413k per year (Estimated) • Remote/Hybrid • 12+ years exp • San Francisco
Java
Python
Ruby
Databases
DynamoDB
ElasticSearch
RocksDB
AI/ML
AI Agents
LLM Guardrails
Apply
$149k – $260k per year • In office • Full-Time • 6+ years exp • PhD • San Francisco
Go
JavaScript
TypeScript
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Prompt Engineering
Agentforce
OpenAI Codex
Frontend
React.js
DevOps
Kubernetes
GitHub
Marketing
Salesforce
Apply
See all jobs
This is one of many
373,573 more open roles from verified company boards, updated every day.