435,297open jobs
15,143companies
67,158added this week
Browse all
Salary
$182k – $238k per year
Location
In office (London)
Seniority
Junior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Apollo Research is an artificial intelligence safety organisation focused on deceptive model behaviour. Its evaluations test whether frontier models scheme or hide their intentions. The laboratory publishes interpretability research and advises policymakers on model risk.

THE OPPORTUNITY

We are currently buildingWatcher, a monitoring tool for coding agents. Our monitoring research agenda attempts to translate compute into safety at scale. Red-teaming previously sat inside the RS (Control) role as a partial responsibility. As it's grown from a single pilot into a recurring need, it now needs a dedicated owner.

As the AI Red Team Engineer, you will help build the practice of red-teaming AI monitors (both Watcher's own defenses and frontier labs' monitoring systems (see our pilot campaign red-teaming Anthropic's auto mode). You will hunt for attack surfaces monitors that haven't been tested against yet and turn what you find into fixes.You'll work closely with Marius (CEO & currently leads the monitoring efforts), control researchers and product engineers.

You will like this opportunity if you think like an attacker and want your adversarial findings to directly strengthen AI monitoring systems. You will join a small team and will have significant ability to shape the team & tech, and have the ability to earn responsibility quickly.

KEY RESPONSIBILITIES

  • Design and run red-teaming campaigns combining severity-graded failure-mode injection into real trajectories, static monitoring benchmarks (e.g. MonitoringBench), and dynamic off-policy control red-teaming.

  • Identify novel attack surfaces monitors haven't been tested against.

  • Build and maintain automated red-teaming pipelines that attack monitors at scale, rather than relying on one-off manual probing.

  • Design iterative adversarial red-team/blue-team games, working with RS (Control) on the blue-team side to keep escalating attack difficulty as monitors improve.

  • Track the research literature and real-world incidents on agent failure modes, monitor evasion, and adversarial robustness.

  • Turn each campaign into specific, actionable recommendations for monitor developers e.g. on time of intervention, user authorisation, trust boundaries, and monitor visibility

  • Write up campaign findings as external publications and internal reports.

  • Feed findings back into Watcher's monitor development and Apollo's blue-teaming research, so lessons learned externally harden our own product.

JOB REQUIREMENTS

    Must-haves

  • 2+ years of experience in offensive security, adversarial ML, or red-teaming of AI systems.

  • Strong experience with AI coding agents, extensive use of, comparison of, or direct development of frontier coding agents.

  • Experience designing and executing structured adversarial testing campaigns

  • Strong Python programming skills

  • Strong written communication as you will produce publications and campaign write-ups that need to be clear and credible externally.

  • Demonstrated ability to work independently on open-ended adversarial problems.

  • Good-to-haves

  • Familiarity with AI safety concepts, particularly agent-related risks

  • Experience with LLM-as-a-judge setups or AI monitoring more broadly

  • Background in penetration testing, CTFs, or computer security more broadly.

  • We want to emphasize that people who feel they don't fulfill all of these characteristics but think they would be a good fit for the position nonetheless are strongly encouraged to apply. We believe that excellent candidates can come from a variety of backgrounds and are excited to give you opportunities to shine.

REPRESENTATIVE PROJECTS

  • Run a red-teaming campaign against a frontier lab's monitor: work through real coding sessions with injected attacks across a range of failure modes, identify improvement areas, and deliver actionable recommendations the partner lab can implement.

  • Build an automated red-teaming pipeline that attacks Watcher's monitors at scale, rather than relying on one-off manual red-teaming.

  • Investigate a new attack surface: an emerging agent capability or novel evasion vector not yet in Apollo's failure mode catalog and produce a write-up and recommendations for the monitoring team.

BENEFITS

  • This role offers market competitive salary, equity, and competitive benefits.

  • Salary: San Francisco: $182,000 - $238,000; London: £122,000 - £160,000. We will be looking to meaningfully raise salaries soon.

  • Flexible work hours and schedule

  • Unlimited vacation

  • Unlimited sick leave

  • Up to 6 months of paid parental leave

  • Comprehensive health, dental and vision insurance

  • Retirement savings with competitive employer matching (e.g. 401(k) for US employees)

  • Lunch, dinner, and snacks are provided for all employees on workdays

  • Paid work trips, including staff retreats, business trips, and relevant conferences

  • A yearly $1,000 (USD) professional development budget

  • Relocation support and visa fees (if applicable)

LOGISTICS

  • Time Allocation: Full-time

  • Location: This is an in-person role working out of our London or San Francisco office. We offer flexible working hours and wfh arrangements.

  • Visa sponsorship: We sponsor visas in both the UK and US. Sponsorship isn't guaranteed for every role or candidate, but if we make you an offer, we'll work with you to find the right visa route.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,297 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$82k – $109k per year • In office • Full-Time • 6+ years exp • PhD • Rome
Python
JavaScript
TypeScript
Apex
Databases
Snowflake
Databricks
AI/ML
Cursor
LangChain
Claude
LlamaIndex
Prompt Engineering
AI Agents
Agentforce
LLM Guardrails
Multi-Agent Systems
DevOps
CI/CD
Apply
FinOps Engineer 1 day ago
Remote/Hybrid • 5+ years exp
Python
SQL
Databases
Databricks
AI/ML
Copilot
Claude
Claude Code
LLM
LLM Guardrails
DevOps
Terraform
GCP
Azure
AWS
Kubernetes
Grafana
Platform Engineering
Bicep
Azure AKS
FinOps
GitHub
Analytics
Power BI
Apply
$82k – $152k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • London
Python
Analytics
Microsoft Excel
Apply
$101k – $267k per year (Estimated) • In office • Internship • Bachelor's Degree • Singapore
Python
AI/ML
Copilot
LangChain
Scikit-learn
AI Agents
TensorFlow
Pandas
NumPy
PyTorch
Anomaly Detection
Copilot Studio
DevOps
Azure
Management
Confluence
Jira
SharePoint
Apply
$27k – $67k per year (Estimated) • In office • Full-Time • 8+ years exp • Hyderabad
Python
Verilog
C++
Perl
MATLAB
Chips/EDA
Cadence Virtuoso
Siemens ModelSim
Apply
$227k – $296k per year • In office • Full-Time • 5+ years exp • London
AI/ML
LLM
OpenAI
Anthropic
DevOps
CI/CD
Apply
$204k – $385k per year • In office • Full-Time • 5+ years exp • London
AI/ML
AI Agents
LLM
OpenAI
Anthropic
Red Teaming
Cybersecurity
MITRE ATT&CK
Threat Modeling
Apply
$204k – $385k per year • In office • Full-Time • 2+ years exp • London
Python
SQL
AI/ML
Fine-tuning
AI Agents
LLM
Apply
Security Engineer 11 days ago
$214k – $280k per year • In office • Full-Time • 5+ years exp • London
AI/ML
AI Agents
LLM
OpenAI
Anthropic
Cybersecurity
Zero Trust
Threat Modeling
Apply
$214k – $280k per year • In office • Full-Time • 5+ years exp • London
AI/ML
AI Agents
LLM
OpenAI
Anthropic
Red Teaming
DevOps
CI/CD
Apply
$105k – $175k per year (Estimated) • In office • Full-Time • London
Apply
$66k – $177k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
Marketing
Salesforce
Apply
Editor / Speechwriter 5 hours ago
$64k – $169k per year (Estimated) • In office • Full-Time • London
Apply
$73k – $122k per year (Estimated) • In office • Full-Time • 5+ years exp • London
Analytics
Power BI
Management
Outlook
Apply
$81k – $152k per year (Estimated) • Remote/Hybrid • Full-Time • London
Analytics
Power BI
Microsoft Excel
Apply
See all jobs
This is one of many
435,297 more open roles from verified company boards, updated every day.