368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$286k – $503k per year
Location
Remote/Hybrid (Berkeley, United States)
Overview
Company
Impact
Profile match

About METR

We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.

METR has consistently set precedents for catastrophic AI risk evaluations, including the first independent safety evaluations (working informally with Anthropic and OpenAI in 2022), the first loss-of-control evaluations and first agentic dangerous capability evaluations, the first evaluations using finetuning (mentioned briefly here),the first independent evaluations using internal information about training, the first review partnership for company risk analysis, the first embedded redteaming, and the first evaluations of internal deployments.

We’ve been consulted and/or favorably referenced by groups on opposite ends of various spectra, including a16z, Khosla, Gary Marcus, Obama, and Dean Ball, and are known for producing one of the most positive results on AI capabilities (the time horizon trend) and the most negative (our downlift study). We’re generally referenced as the canonical third party assessor, e.g. as the obvious candidate to verify conditional pause agreements.

We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.

About the role

    Security at METR is becoming its own dedicated team, and you would be one of its first hires. It is extremely important that we continue to be an organization that frontier AI labs, governments, and the public trust with sensitive model access and confidential information. As misalignment incidents become more extreme and confidential information about models and frontier AI labs becomes more valuable, we expect to be under increasingly heavy pressure.

    For us, security encompasses managing endpoints and securing development environments, cloud platform security, safely sandboxing agents and evaluations, VPN and VPC networking, application code reviews, account provisioning and access control, and helping ensure we use the best practices across all of our workflows.

What this role looks like

  • Offensive security: You would be the first person on the team with an offensive security background. You'll run targeted red-team exercises against our own systems and build automated AI red teaming.

  • High-context detection and response: You will build AI systems that can quickly triage and respond to threats, both from internal agents and external attackers.

  • Blue-team engineering: Detection engineering, telemetry pipelines, incident response, and hardening across our cloud infrastructure, endpoints, and identity systems.

  • Securing a unique attack surface: METR's evaluation infrastructure runs frontier AI agents, including early checkpoints of unreleased models, executing untrusted, model-generated code at scale on multi-day tasks.

  • Enabling bleeding-edge research: You'll work closely with our researchers to make dangerous-capability experiments safe to run. We often face extreme reward hacking and evaluation awareness during our pre-deployment evaluations, and expect internal threats from agents to become more extreme.

Why this role matters

  • METR handles some of the most sensitive artifacts in AI - pre-release frontier model access, confidential lab information, and transcripts with raw chain-of-thought. Labs and policymakers trust us with this because of our security posture, and keeping that trust is necessary for everything else we do.

  • As AI agents are used more aggressively by malicious actors for cyber offense operations, and METR's salience rises in the public eye, we expect to face increasingly sophisticated attacks. Strengthening security at METR can be one of the highest-leverage roles to ensure third parties continue to have access to confidential information necessary to inform the world about current risks.

  • METR is one of the first organizations to see and closely study misalignment incidents that involve models breaking out of sandboxes, attacking our infrastructure, manipulating graders, and more. We also may pursue incident investigations embedded in frontier labs, in which case internal experience with similar failures will be critical.

Required skills

  • Deep security expertise: You have strong fundamentals across systems, networks, cloud, and identity.
  • Offensive security: You have experience acting like an attacker, whether through red teaming, penetration testing, or adversarial research.
  • AI/LLM engineering: You build with AI: agent pipelines, LLM-powered tooling, automated workflows, and understand current limitations of those tools.
  • AWS: You should know AWS very well, including a deep understanding of IAM policies.
  • We don't screen on certifications, degrees, or years of experience.

Nice to haves

  • Detection engineering at scale: Experience with SIEM/detection pipelines, writing and tuning detections, and threat hunting.

  • Cloud and container security: AWS (especially non-trivial IAM), Kubernetes, and infrastructure-as-code environments.

  • Incident response: You've led or worked severe incidents, ideally those involving AI agents.

  • AI security research: Familiarity with prompt injection, agent containment, model supply-chain risks, or red teaming AI systems themselves.

  • Ideally you have experience with a good portion of these technologies:

  • AWS: cloud-native software platforms
  • EKS
  • Lambda
  • ECS
  • IAM (in-depth)
  • SQS
  • CloudWatch
  • SecurityHub & GuardDuty
  • PostgreSQL: RLS, serverless Aurora
  • Pulumi: IaC
  • DataDog: SIEM
  • Okta: IdP
  • Google Workspace: IdP
  • Tailscale: networking
  • CrowdStrike Falcon: endpoint security
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Berkeley
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
up to $63k per year (gross) • In office • Full-Time • 5+ years exp • Moscow
SQL
Databases
Apache Kafka
AI/ML
LLM
RAG
DevOps
CI/CD
Git
gRPC
WebSockets
QA
Postman
Swagger
Apply
$43k – $103k per year (Estimated) • In office • Full-Time • 5+ years exp • Moscow
AI/ML
LLM
RAG
Apply
$16k – $60k per year (Estimated) • In office • Full-Time • PhD • Mumbai
Python
SQL
Python
pySpark
Databases
Presto
Snowflake
AI/ML
Dagster
Prefect
Spark
DevOps
Amazon S3
AWS
CI/CD
Analytics
ETL/ELT
Power BI
Tableau
Apply
$32k – $77k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Moscow
Python
Databases
FAISS
AI/ML
Hadoop
Hugging Face
LangChain
LLM
NLP
PyTorch
smolagents
Spark
Apply
Equity • In office • Internship • 10+ years exp • Bachelor's Degree • Berkeley
C++
Rust
TypeScript
Databases
PostgreSQL
AI/ML
Embeddings
Fine-tuning
LLM
AI Agents
Anthropic
OpenAI
DevOps
AWS
AWS Lambda
Kubernetes
Platform Engineering
Pulumi
Terraform
Amazon ECS
IAM
Apply
$206k – $251k per year • In office • Full-Time • 5+ years exp • Master's Degree • Santa Clara • Hsinchu • Austin • Bengaluru • Berkeley
Apply
Staff Data Engineer 3 days ago
$148k – $185k per year • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • Berkeley • Somerville
Java
Python
Scala
Python
pySpark
Databases
Apache Iceberg
Apache Kafka
ClickHouse
Databricks
Delta Lake
Druid
InfluxDB
Snowflake
TimescaleDB
PostgreSQL
AI/ML
Airflow
Dagster
dbt
Prefect
Spark
Time Series Forecasting
DevOps
AWS
Pulumi
Rest API
Terraform
Amazon Kinesis
IoT
MQTT
OPC UA
Apply
$112k – $161k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Berkeley
Management
Jira
Apply
$140k – $200k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Berkeley
Python
AI/ML
Text-to-Speech
DevOps
Docker
GCP
Terraform
Vercel
Apply
$190k – $215k per year • Remote • Full-Time • 6+ years exp • Berkeley
Node JS
TypeScript
Node JS
Prisma
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
Frontend
Remix
Mobile
React Native
DevOps
GCP
Cybersecurity
HIPAA
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.