435,295open jobs
15,159companies
67,452added this week
Browse all
Salary
$222k – $290k per year
Location
In office (London)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Apollo Research is an artificial intelligence safety organisation focused on deceptive model behaviour. Its evaluations test whether frontier models scheme or hide their intentions. The laboratory publishes interpretability research and advises policymakers on model risk.

THE OPPORTUNITY

We are buildingWatcher, a coding agent security product. Watcher is deployed in production and monitors billions of agent tokens per month across engineering teams at agent-building scale-ups and enterprises. We are looking for a product engineer to build and scale Watcher end to end.

This role includes backend and frontend components. You will own features across the full stack, from agent hooks and ingestion pipelines to monitoring dashboards and enterprise integrations. You can expect to own features from customer conversation to production deployment.

This is truly a “start-up role”: you will own big chunks of the product, make decisions fast, and ship at high volume. You will join a small team with significant ability to shape the product and tech, and you can earn more responsibility quickly. This is an individual contributor role but could lead to management responsibilities eventually, if desired.

KEY RESPONSIBILITIES

    Ship product end to end (~45%)

    Feature development

  • Design, build, and ship Watcher features across the stack: agent-side hooks and CLI tooling, ingestion and grading pipelines, real-time monitoring UI (Watcher Live), and the organization-wide Analyzer dashboard.

  • Own features end to end, from ambiguous requirements to production: talk to users, scope the work, make the technical decisions, put that into production, and iterate based on usage.

  • Build Watcher integrations: We currently support Claude Code and Codex, but we’re planning on increasing our coverage to other coding agents, LLM gateways, and anywhere that could benefit from agent monitoring.

  • Build the policy and configuration layer that lets security teams set org-wide guardrails for engineers to operate safely within.

  • Productionize research

  • Turn monitoring research (new monitors, grading strategies, backtesting results) into production-ready product features. We are "our own customer": You will use the product you build every day.

  • Infrastructure and scale (~35%)

  • Design and operate backend systems that process large volumes of agent logs in real time.

  • Own reliability: robust error handling, graceful degradation, and observability (logging, metrics, tracing, alerting). We should catch issues before customers do.

  • Architect data models and storage that handle both high-throughput writes and complex analytical queries over historical trajectories. Retention policies have to fit highly sensitive customer data.

  • Make sure Watcher is safe to deploy: Work with our AI Security & Control Engineer on tenant isolation, encryption, and access controls. Watcher touches some of the most sensitive data our customers have.

  • Customer and integration engineering (~20%)

  • Build and maintain integrations with the enterprise security stack, e.g. streaming monitoring alerts into SIEM systems and existing security operations workflows.

  • Support flexible deployment models: cloud-hosted, on-prem, and local backends, so sensitive agent logs never have to leave customer control.

  • Talk to customers, debug issues in their environments, and feed what you learn back into the roadmap.

REPRESENTATIVE PROJECTS

  • Secure a new frontier coding agent: Take an agent we don't yet support (e.g. Cursor) from zero to fully monitored. Build the hook/integration layer, normalize its trajectory format into our data model, validate monitor accuracy on real traffic, and ship it to customers.

  • Build the real-time blocking path: Design the low-latency pipeline that lets monitors block dangerous actions (e.g. git push --force, secret exfiltration) before they execute, while keeping p95 latency low enough that developers don't feel it. Handle traffic spikes and partial failures gracefully.

  • Ship an Analyzer capability end to end: Design and build an organization-wide view (e.g. failure trends over time, cross-session pattern detection) from data model through API to UI, based on what security teams actually need.

  • Deploy Watcher inside locked-down enterprises: Build and document the self-hosted/on-prem deployment story so that enterprises with strict data requirements can run Watcher inside their own infrastructure.

JOB REQUIREMENTS

    Must-haves

  • 4+ years building production software. You have shipped and operated real systems with real users, and you know what it takes to keep them reliable.

  • Full-stack capability. You are strong on the backend (APIs, data pipelines, databases, cloud infrastructure) and productive on the frontend (modern web UI). You don't need to be world-class at both, but you must be able to own a feature across the whole stack. Our stack is primarily Python and TypeScript.

  • High agency and ownership. You are willing to own big chunks of the product, make decisions fast with incomplete information, and be accountable for the outcome. You don't wait to be told what to do and have an accurate sense of the roadmap for the product.

  • High output volume. You have used Claude Code, Codex, Cursor, or similar tools heavily to accelerate or have built agents or agent tooling yourself.

  • Startup pace. You are excited about a fast-moving environment, comfortable with ambiguity and changing priorities, and willing to grind when it matters.

  • Product sense. You can talk to users, translate vague requirements into concrete designs, and make good calls about what to build and what to cut.

  • Strong nice-to-haves

  • Previous work on developer tools, monitoring/observability systems, or security products. Watcher sits at the intersection of all three.

  • Real-time or large-scale data processing experience (streaming pipelines, message queues, high-throughput log systems).

  • Enterprise or on-prem deployment experience. Shipping software into environments you don't control is its own skill.

  • Experience building LLM-powered applications, e.g. LLM-as-judge setups, evaluation pipelines, or familiarity with frameworks like Inspect.

  • Early-stage startup experience. You have been employee 1-15 somewhere and know what it feels like to build a product with no playbook.

  • Explicitly not required

  • Formal AI safety background. We need excellent product engineers who can learn the AI safety context, not AI safety researchers who need to learn engineering.

  • Management experience. This is an IC role, at least initially.

  • Deep experience in every technology we use. We care about demonstrated ability to learn and ship, not checkbox familiarity with our exact stack.

BENEFITS

  • This role offers market competitive salary, equity, and competitive benefits.

  • Salary: San Francisco: $222,000 - $290,000 London: £149,000 - £195,000. We will be looking to meaningfully raise salaries soon.

  • Our engineers effectively have an unlimited token budget. If a better result costs more compute, use it.

  • Flexible work hours and schedule

  • Unlimited vacation

  • Unlimited sick leave

  • Up to 6 months of paid parental leave

  • Comprehensive health, dental and vision insurance

  • Retirement savings with competitive employer matching (e.g. 401(k) for US employees)

  • Lunch, dinner, and snacks are provided for all employees on workdays

  • Paid work trips, including staff retreats, business trips, and relevant conferences

  • A yearly $1,000 (USD) professional development budget

  • Relocation support and visa fees (if applicable)

LOGISTICS

  • Time Allocation: Full-time

  • Location: This is an in-person role working out of our London or San Francisco office. We offer flexible working hours and some wfh arrangements.

  • Visa sponsorship: We sponsor visas in both the UK and US. Sponsorship isn't guaranteed for every role or candidate, but if we make you an offer, we'll work with you to find the right visa route.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,295 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$94k – $184k per year (Estimated) • In office • New York
Python
JavaScript
TypeScript
C#
AI/ML
Copilot
Claude
Claude Code
Frontend
React.js
DevOps
Rest API
Terraform
CI/CD
AWS
Kubernetes
Shift-Left
Self-Healing
Amazon EKS
GitHub
Cybersecurity
Shift-Left Security
QA
Selenium
JMeter
Playwright
Gatling
k6
Apply
Remote/Hybrid • 4+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
Senior AI Engineer 1 day ago
Remote/Hybrid • 9+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
$133k – $222k per year • Remote • 3+ years exp • Bachelor's Degree
AI/ML
Copilot
Claude
Apply
$126k – $265k per year (Estimated) • Remote • 8+ years exp
Python
SQL
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
$227k – $296k per year • In office • Full-Time • 5+ years exp • London
AI/ML
LLM
OpenAI
Anthropic
DevOps
CI/CD
Apply
$204k – $385k per year • In office • Full-Time • 5+ years exp • London
AI/ML
AI Agents
LLM
OpenAI
Anthropic
Red Teaming
Cybersecurity
MITRE ATT&CK
Threat Modeling
Apply
$182k – $238k per year • In office • Full-Time • 2+ years exp • London
Python
AI/ML
LLM
Anthropic
Red Teaming
DevOps
Vector
Apply
$204k – $385k per year • In office • Full-Time • 2+ years exp • London
Python
SQL
AI/ML
Fine-tuning
AI Agents
LLM
Apply
Security Engineer 11 days ago
$214k – $280k per year • In office • Full-Time • 5+ years exp • London
AI/ML
AI Agents
LLM
OpenAI
Anthropic
Cybersecurity
Zero Trust
Threat Modeling
Apply
$105k – $175k per year (Estimated) • In office • Full-Time • London
Apply
$66k – $177k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
Marketing
Salesforce
Apply
Editor / Speechwriter 5 hours ago
$64k – $169k per year (Estimated) • In office • Full-Time • London
Apply
$73k – $122k per year (Estimated) • In office • Full-Time • 5+ years exp • London
Analytics
Power BI
Management
Outlook
Apply
$81k – $152k per year (Estimated) • Remote/Hybrid • Full-Time • London
Analytics
Power BI
Microsoft Excel
Apply
See all jobs
This is one of many
435,295 more open roles from verified company boards, updated every day.