368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$250k – $450k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Infrastructure for understanding AI Infrastructure for understanding AI. Transluce is a non-profit research lab building the public tech stack for scalable oversight of AI.

Salary range: $250,000 - $450,000/year + benefits

Description: Transluce is a fast-moving nonprofit research lab building the public tech stack for AI evaluation and oversight. We are pioneering research into the behaviors of AI chatbots and their effect on user wellbeing, and we’re improving outcomes for millions of sensitive AI interactions with vulnerable users.

About the role: As an AI Behavior Researcher, you will lead projects to design and develop automated evaluations of frontier AI systems that are technically sophisticated, scientifically valid, and concretely impactful. This includes expanding on our existing evaluation pipelines to conduct novel analyses of AI behaviors that affect the autonomy and wellbeing of specific user groups (e.g., children or users located in countries beyond the United States).

As an early member of a highly collaborative team, you will learn and grow quickly, and work directly with frontier labs to improve AI evaluations design, with regulators to improve independent oversight of AI, and with domain experts and affected populations to enhance the realism and relevance of our evaluations.

Core responsibilities:

  • Develop novel, valid automated evaluations of AI’s impacts on users, including their mental health and decision making.
  • Write code to implement and run automated evaluations, such as user simulators or LLM-as-a-judge pipelines.
  • Design methods to improve the ecological validity and realism of automated evaluations for specific populations, such as customizing existing user simulation methods to capture the vocabulary used by children.
  • Write and revise judge rubrics to evaluate model behaviors related to user wellbeing and decision making, systematizing abstract, socially situated concepts into clear measurement criteria.
  • Collaborate with scientists and research engineers to productionize best practices in AI behavioral evaluation.

Minimum qualifications:

  • Expertise on quantitative generative AI evaluation and measurement. Good intuition about how to systematize and operationalize complex social concepts.
  • Relevant experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
  • Proficiency in Python to implement analysis and evaluation tooling.
  • Meticulous, good experimental design, epistemic self-awareness and transparency.
  • Ability to balance between the needs of AI researchers and domain experts, as well as between researchers and senior decision makers.
  • Strong communication skills, low ego, openness to giving and receiving feedback.

Preferred qualifications (not required):

  • Experience running automated evaluations at scale or in a production context.
  • Experience conducting controlled human subjects experiments to validate automated evaluation methods.
  • Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous stakeholder needs into concrete deliverables.
  • Experience or training in human-centered design or HCI research methods, including working with domain experts or impacted communities.
  • Experience or demonstrated interest in studying AI’s psychological or social impacts, such as for crisis support, manipulation or sycophancy, political persuasion, or displacing human relationships.
  • Experience designing multilingual generative AI evaluations.
  • Experience and comfort using AI coding agents at work.

We are hiring at all levels of experience and would encourage those enthusiastic about the role who do not meet all of the qualifications to apply. We are located in San Francisco and excited to work together in-person. We are open to sponsoring international visas.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$21k per year • In office • Contractor • Yekaterinburg
Python
SQL
Python
FastAPI
Flask
AI/ML
Claude
Claude Code
Embeddings
Function Calling
LLM
RAG
OpenAI
OpenAI Codex
Structured Outputs
DevOps
Docker
Git
Apply
$140k – $225k per year • Remote • Full-Time • 5+ years exp • Seattle
Python
Lua
Python
FastAPI
Celery
Pydantic
SQLAlchemy
Databases
pgvector
PostgreSQL
Redis
AI/ML
LLM
RAG
Hybrid Search
Reranking
Anthropic
LLM Evaluation
LLM Guardrails
OpenAI
Model Context Protocol
DevOps
OpenTelemetry
Apply
$98k – $132k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
R
SQL
Databases
Databricks
Snowflake
AI/ML
Embeddings
LangChain
LangGraph
LLM
NumPy
Pandas
Scikit-learn
Spark
TensorFlow
Transformers
Hugging Face
RAG
Analytics
Matplotlib
Seaborn
Apply
$98k – $132k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
R
SQL
Databases
Databricks
Snowflake
AI/ML
Embeddings
LangChain
LangGraph
LLM
NumPy
Pandas
Scikit-learn
Spark
TensorFlow
Transformers
Hugging Face
RAG
Analytics
Matplotlib
Seaborn
Apply
$98k – $132k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
R
SQL
Databases
Databricks
Snowflake
AI/ML
Embeddings
LangChain
LangGraph
LLM
NumPy
Pandas
Scikit-learn
Spark
TensorFlow
Transformers
Hugging Face
RAG
Analytics
Matplotlib
Seaborn
Apply
$250k – $500k per year • In office • Full-Time • San Francisco
AI/ML
Fine-tuning
Apply
VP of Engineering 3 months ago
$500k – $600k per year • In office • Full-Time • San Francisco
AI/ML
Claude
Apply
AI Behavior Engineer 4 months ago
$310k – $500k per year • In office • Full-Time • San Francisco
AI/ML
AI Agents
Apply
AI Systems Engineer 5 months ago
$350k – $600k per year • In office • Full-Time • San Francisco
Python
AI/ML
LLM
vLLM
AI Agents
Apply
$250k – $500k per year • In office • Full-Time • San Francisco
AI/ML
AI Agents
LLM
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
$173k – $260k per year • In office • Full-Time • PhD • San Francisco
JavaScript
Node JS
Python
Python
Celery
Django
Flask
Databases
RabbitMQ
Redis
AI/ML
Agentforce
AI Agents
DevOps
Akamai
AWS
CI/CD
Cloudflare
CloudFormation
Helm
Jenkins
Kubernetes
Spinnaker
Terraform
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.