414,464open jobs
14,229companies
59,920added this week
Browse all
Salary
$90k – $110k per year
Location
Remote (United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Policy Evaluator, Violence & Threats based in United States.

This role sits at the intersection of AI safety, content policy, and expert human judgment, helping improve how advanced AI models handle violent and threatening content.

You’ll evaluate user requests, model responses, and conversation context to distinguish legitimate fictional, educational, historical, or defensive content from material that could enable real-world harm.

The work focuses heavily on nuanced edge cases where intent, context, and a single detail can materially change the appropriate policy decision.

You’ll contribute not only to evaluations, but also to adversarial testing, policy refinement, calibration, and the identification of emerging safety gaps.

This is a non-engineering role for someone with deep experience in areas such as violent fiction, military or emergency response, crisis intervention, threat assessment, trust and safety, or related fields.

You’ll work in a structured, feedback-rich environment where clear reasoning, consistency, and the ability to separate personal beliefs from policy standards are essential.

Because the role involves regular exposure to difficult material, candidates must be prepared to engage with sensitive content carefully, professionally, and sustainably.

Accountabilities

    • Evaluate user requests, AI responses, and conversation histories involving violence, weapons, threats, self-harm, dark fiction, and related sensitive topics.
    • Distinguish fictional, educational, historical, journalistic, defensive, or expressive content from requests that meaningfully facilitate real-world harm.
    • Assess whether an AI response provides actionable real-world capability, regardless of how the original request is framed.
    • Differentiate ordinary anger, frustration, venting, or dark humor from credible threats and potential crisis indicators.
    • Apply relevant customer policies consistently while considering intent, context, precedent, and established team guidance rather than relying on rote annotation.
    • Select defensible classifications for ambiguous cases and produce concise, well-supported rationales referencing policy language and relevant conversation details.
    • Develop and refine adversarial or borderline prompts that test how AI systems handle difficult policy boundaries.
    • Identify policy gaps, contradictions, recurring ambiguities, and emerging edge cases, escalating findings to project leads and policy teams.
    • Participate actively in calibration and adjudication discussions, respectfully challenging interpretations and updating judgments when stronger reasoning emerges.
    • Maintain high accuracy, consistency, and attention to detail across repetitive, feedback-heavy evaluation workflows.
    • Contribute to improving AI safety standards by turning nuanced human judgment into clear, auditable evaluation guidance.
    • Requirements

      • Demonstrated depth of experience in at least one relevant domain, such as violent fiction, game design or game mastering, film and television, military or law enforcement, emergency medicine, crisis counseling, threat assessment, trust and safety, content moderation, journalism, or law.
      • Strong ability to distinguish fictional or contextualized depictions of violence from content that facilitates or signals real-world harm.
      • Demonstrated understanding of how intent, context, language, and subtle changes in a request can materially affect a safety assessment.
      • Ability to make nuanced judgment calls while separating personal beliefs from the policy standard being applied.
      • Strong written communication skills, with the ability to explain complex decisions clearly enough for another evaluator to audit the reasoning.
      • Ability to remain open-minded, challenge assumptions constructively, and revise conclusions when stronger evidence or reasoning emerges.
      • Strong attention to detail and consistency when working through repeated evaluations involving difficult or sensitive material.
      • Familiarity with AI tools and a strong interest in understanding where language models may over-refuse, under-refuse, or misunderstand user intent.
      • Prior experience with AI evaluation, red teaming, data annotation, RLHF, trust and safety, content moderation, or language-model assessment is helpful but not required.
      • Experience with calibration sessions, inter-rater agreement, adjudication workflows, or structured policy evaluation is a plus.
      • Additional valuable experience may include published or produced work involving violence, military or security experience, crisis intervention, threat assessment, forensic or clinical psychology, violence prevention, weapons disciplines, or professional experience evaluating ChatGPT, Claude, Gemini, or similar AI systems.
      • A degree, security clearance, or technical/software engineering background is not required.
      • Ability to work remotely in the United States on a Monday-Friday schedule from 8:00 AM to 5:00 PM PT.
      • Benefits

        • Compensation: $45-$55 per hour.
        • Employment classification: W-2.
        • Work arrangement: Fully remote within the United States.
        • Schedule: Monday through Friday, 8:00 AM-5:00 PM Pacific Time.
        • Assignment: Ongoing opportunity with a planned start date of September 21, 2026.
        • Benefits eligibility: Eligible for available employee benefits.
        • Structured evaluation frameworks, professional guidelines, content rotation, and exposure limits designed to support sustainable work with sensitive material.
        • Access to mental health support given the nature of the content reviewed.
        • Opportunity to contribute directly to the safety, reliability, and responsible development of advanced AI systems.
        • Exposure to complex policy questions and emerging challenges at the intersection of AI, violence, threats, and content safety.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
414,464 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$100k – $150k per year • Remote • 6+ years exp • Master's Degree
Python
AI/ML
RLHF
Reinforcement Learning
AI Agents
Reward Modeling
Robotics
Reinforcement Learning
Apply
$100k – $175k per year • Remote • 6+ years exp • Master's Degree
Python
AI/ML
Fine-tuning
RLHF
Reinforcement Learning
Multimodal AI
Knowledge Distillation
PyTorch
LLM
Synthetic Data
DPO
FSDP
Model Distillation
Apply
$96k – $120k per year • Remote • 8+ years exp • Master's Degree
Python
AI/ML
RLHF
Reinforcement Learning
AI Agents
Reward Modeling
Robotics
Reinforcement Learning
Apply
$130k – $180k per year • Remote • 10+ years exp • Master's Degree
Python
AI/ML
LangGraph
LangChain
DeepSpeed
LlamaIndex
LoRA
Fine-tuning
RLHF
Multimodal AI
AI Agents
NLP
PEFT
QLoRA
Transformers
PyTorch
LLM
RAG
Ray
Hallucination
Synthetic Data
DPO
SFT
PPO
FSDP
Knowledge Graph
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Vector
Apply
$72k – $100k per year • Remote • 8+ years exp • Master's Degree
Python
AI/ML
Fine-tuning
RLHF
Reinforcement Learning
Multimodal AI
Knowledge Distillation
PyTorch
LLM
Synthetic Data
DPO
FSDP
Model Distillation
Apply
$30k – $72k per year (Estimated) • Remote • Full-Time • 10+ years exp
Apply
$59k – $141k per year (Estimated) • Remote • Full-Time • 10+ years exp
Apply
$35k – $89k per year (Estimated) • Remote • Contractor • 5+ years exp • Bachelor's Degree
Python
AI/ML
MLFlow
Evidently AI
TensorFlow
PyTorch
Amazon SageMaker
Recommender Systems
DevOps
Prometheus
GitLab CI
CI/CD
AWS
Grafana
GitLab
Apply
$53k – $131k per year (Estimated) • Remote • Contractor • 5+ years exp • Bachelor's Degree
Python
AI/ML
MLFlow
Evidently AI
TensorFlow
PyTorch
Amazon SageMaker
Recommender Systems
DevOps
Prometheus
GitLab CI
CI/CD
AWS
Grafana
GitLab
Apply
$11k – $83k per year (Estimated) • Remote/Hybrid • Contractor
Apply
See all jobs
This is one of many
414,464 more open roles from verified company boards, updated every day.