687,002open jobs
39,973companies
97,851added this week
Browse all
Salary
$104k – $258k per year (Estimated)
Location
Remote/Hybrid (Ramat Gan, Israel)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Alice is the most trusted partner for GenAI safety and security. Built on a decade of threat intelligence research, Alice helps organizations detect, prevent, and respond to AI risks across the full model lifecycle.

Description

We are seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. The ideal candidate will establish rigorous benchmarking methodologies, evaluate large language models (LLMs) against complex threats like Indirect Prompt Injections (IPI), and construct post-training evaluation pipelines that accurately measure realistic frontier-level security capabilities.

Key Responsibilities

  • RL Post-Training & Benchmarking: Execute post-training runs (e.g. GRPO) using mainstream open-weight generalist models against security-focused RL environments, targeting threat vectors like Indirect Prompt Injection (IPI). Reward Diagnostics & Trace Analysis - Analyze live loss curves and rollout traces to identify reward hacking, lazy policy convergence, and flawed or over/under-specified verifiers.
  • Task & Environment Auditing: Review tasks and multi-turn environments (including tool use, web navigation, and computer use) for realism, threat model accuracy, data distribution, and dataset balance.
  • Performance Reporting (Gym Cards): Generate comprehensive evaluation cards detailing hill-climbing performance uplift across checkpoints, failure modes, tokens/turns per rollout, and task-level success rates.
  • Integration & Orchestration: Integrate dockerized environments (e.g., Harbor format) into internal training frameworks, optimizing reset/statefulness semantics, concurrency, and throughput ceilings.

Requirements

Required Qualifications

  • Technical Background: M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning.
  • RL & Post-Training Expertise: Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures.
  • AI Security Expertise: Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks.
  • Infrastructure Skills: Proficiency in PyTorch, Docker containerization, and distributed training architectures.
  • Diagnostic Skills: Ability to analyze agent rollout traces, craft deterministic rubrics/verifiers, and debug complex reward shaping flaws.

Preferred Qualifications

  • Prior experience working with standard RL gym formats, such as Harbor.
  • Experience evaluating complex agentic workflows in tool-use or web-browser environments.
  • Familiarity with evaluating open-weight models similar to Llama or Mistral against adversarial workloads.

About Alice

Alice is a trust, safety, and security company built for the AI era. We safeguard the communicative technologies people use to create, collaborate, and interact- whether with each other or with machines.

In a world where AI has fundamentally changed the nature of risk, Alice provides end-to-end coverage across the entire AI lifecycle. We support frontier model labs, enterprises, and UGC platforms with a comprehensive suite of solutions: from model hardening evaluations and pre-deployment red-teaming to runtime guardrails and ongoing drift detection.

Alice is widely considered a global leader in online safety and AI security. We have some of the most forward-thinking and passionate minds in the world working to safeguard over 3 billion users across the largest AI and tech platforms.

If you're creative and driven to secure the future of AI, we want to hear from you!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
687,002 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Ramat Gan
In office • 5+ years exp
AI/ML
Cursor
Claude
FastAI
AI Agents
PyTorch
Google AI Studio
Lovable
Frontend
Sass
Design
Figma
Apply
$255k – $270k per year • Remote/Hybrid • Bachelor's Degree • San Francisco
AI/ML
Claude
Multimodal AI
LLM
Anthropic
Interpretability
DevOps
Incident Management
Cybersecurity
ISO 27001
SOC 2
Apply
$17k – $42k per year (Estimated) • In office • Full-Time • Bachelor's Degree
SQL
Databases
ClickHouse
Apache Kafka
DevOps
OpenShift
VMWare
Docker
Kubernetes
Hyper-V
Apply
$233k – $336k per year • In office • Full-Time • Bachelor's Degree • Cambridge
AI/ML
LLM
Apply
$110k – $150k per year • Remote • Full-Time • Bachelor's Degree • Philadelphia
Python
JavaScript
TypeScript
PowerShell
C#
Node JS
C#
.NET
Frontend
Angular
pnpm
DevOps
Terraform
GitHub Actions
Docker Swarm
Azure
CI/CD
ArgoCD
Docker
Kubernetes
Platform Engineering
Octopus Deploy
GitHub
Cybersecurity
HIPAA
Management
Agile
Apply
$119k – $242k per year (Estimated) • Remote • Full-Time • 3+ years exp • Master's Degree • New York
AI/ML
vLLM
Function Calling
AI Agents
DPO
SFT
GRPO
Post-training
LLM Evaluation
LLM Guardrails
Tool Use
Apply
$119k – $242k per year (Estimated) • Remote • Full-Time • 3+ years exp • Master's Degree • San Francisco
AI/ML
vLLM
Function Calling
AI Agents
DPO
SFT
GRPO
Post-training
LLM Evaluation
LLM Guardrails
Tool Use
Apply
$119k – $242k per year (Estimated) • Remote • Full-Time • 3+ years exp • Master's Degree
AI/ML
vLLM
Function Calling
AI Agents
DPO
SFT
GRPO
Post-training
LLM Evaluation
LLM Guardrails
Tool Use
Apply
$85k – $168k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • London
Cybersecurity
VirusTotal
Recorded Future
Apply
$56k – $134k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Ramat Gan
Apply
$56k – $134k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Ramat Gan
Apply
$72k – $173k per year (Estimated) • In office • Full-Time • 2+ years exp • Ramat Gan
JavaScript
C++
Cybersecurity
Wireshark
Ghidra
IDA Pro
Frida
Apply
Malware Researcher 2 days ago
$88k – $227k per year (Estimated) • Remote/Hybrid • 3+ years exp • Ramat Gan
Python
JavaScript
C++
Dart
Frontend
React.js
Mobile
Flutter
React Native
Cybersecurity
Ghidra
IDA Pro
Frida
Apply
Remote/Hybrid • PhD • Ramat Gan
Python
Rust
Python
pySpark
Rust
PyO3
Databases
Databricks
Apache Kafka
AI/ML
Cursor
Grok
Spark
Claude Code
MLFlow
AI Agents
AWS Bedrock
Gemini
LLM
Triton
OpenAI
Anthropic
OpenAI Codex
LLM Guardrails
DevOps
Terraform
CI/CD
AWS
Kubernetes
SLI/SLO/SLA
Apply
$112k – $254k per year (Estimated) • In office • Full-Time • 10+ years exp • Ramat Gan • Istanbul • Budapest • Lisbon
Cybersecurity
MITRE ATT&CK
Apply
See all jobs
This is one of many
687,002 more open roles from verified company boards, updated every day.