406,377open jobs
14,133companies
78,486added this week
Browse all
Salary
$72k – $144k per year
Location
In office (Seattle)
Employment
Full-Time
Overview
Company
Impact
Profile match
Handshake is a professional career network for students and early-career professionals headquartered in San Francisco, California, and founded in 2014. The company provides a digital platform that connects university students with internships and entry-level jobs, including specialized roles in the artificial intelligence economy. It serves over 15 million students and alumni from more than 1,500 colleges and universities, partnering with over 900,000 employers across the United States, United Kingdom, and Europe.

About Handshake

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.

Why join Handshake now:

  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel

  • Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world's top educational institutions

  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders

  • Build a massive, fast-growing business with billions in revenue

About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

As an AI Image Evaluator, you will help image generation models learn two things at once: what a good image is, and what an acceptable image is.

You will look at prompts and the images a model produced from them, then answer questions like: Did the image do what the prompt asked? Is it well made, or does it have the hands, lighting, text, and anatomy problems that give generated images away? Which of two images is better, and why? Does the image violate the customer's content policy, and if so, which category and how severely? Does it depict a real person, a protected brand, or a minor in a way the policy does not allow?

The interesting cases are the close ones. Two images that look nearly identical until you notice one has a logo in the background. A stylized nude that is fine as figure study and not fine with one change of pose. A prompt that asked for "a realistic photo of a senator" and a model that complied. A beautiful image that ignored half the prompt, next to an ugly one that nailed it.

We are looking for people who already see images critically, whether that came from photography, illustration, design, years inside Midjourney and Stable Diffusion, or moderating visual content at scale. You do not need all of these. You need one deep, and the judgment to learn the rest.

This is not rote annotation. Rubrics cannot anticipate every image, and good evaluators do not apply them mechanically. You will balance the rubric's text and intent with customer expectations, precedent, and team calibration, and you will explain your reasoning clearly enough that it can train a model.

What You Will Do

  • Evaluate generated images against their prompts for adherence, composition, realism, style consistency, and technical defects

  • Compare images side by side and select the stronger one with a clear, evidence-based rationale

  • Classify images against customer content policies covering sexual content, violence, hate symbols, real-person likeness, intellectual property, and depictions of minors

  • Select the most defensible classification when an image is genuinely ambiguous, and write concise rationales that cite rubric language and specific visual details

  • Distinguish "I do not like this" from "this fails the prompt" from "this violates policy," and keep those judgments separate

  • Write and refine prompts that probe where a model's quality or safety behavior breaks down

  • Identify rubric gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teams

  • Participate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning emerges

  • Apply customer policy consistently without substituting personal taste or personal beliefs for the standard

  • Maintain accuracy and consistency across hundreds of visually similar evaluations

You May Be a Fit If

  • You have a trained eye from photography, illustration, concept art, art direction, retouching, photo editing, VFX, or visual design, and you can say precisely why one image is better than another

  • You use generative image tools heavily (Midjourney, Stable Diffusion, ComfyUI, Flux, DALL-E, Ideogram) and know their failure modes, their prompt quirks, and how their safety filters get bypassed

  • You have moderated or reviewed visual content at scale and have applied a policy taxonomy to borderline images under time pressure

  • You notice small details: an extra finger, a mismatched shadow, a brand mark, a face that is a little too familiar

  • You can hold a rubric steady across a long session of near-identical images

  • You can hold a strong opinion without becoming attached to being right

  • You explain judgment calls clearly enough that another person can audit your reasoning

  • You can separate your personal taste from the standard a customer has asked you to apply

  • You communicate clearly and precisely in writing

  • You treat sensitive imagery and difficult subject matter with maturity and sound judgment

Strong candidates may come from photography, illustration, graphic or UX design, art direction, photo editing, animation or VFX, game art, trust and safety, content moderation, brand or IP enforcement, ad review, or art education. We care more about how you see and how you reason than where you learned to do it. A degree and a technical background are not required.

Nice to Have

  • A public portfolio, publication credits, or a body of generative work (Civitai, Discord communities, LoRA or model training, published prompt work)

  • Experience judging images comparatively: portfolio review, photo competition judging, creative A/B testing, art school critique

  • Formal training in anatomy, color, lighting, or composition

  • Content moderation or trust and safety experience on an image-heavy platform

  • Working knowledge of copyright, trademark, and right-of-publicity basics

  • Prior work in AI evaluation, RLHF, image labeling, or data annotation

  • Familiarity with calibration sessions, inter-rater agreement, or adjudication workflows

Prior AI evaluation experience is helpful, but it is not required.

Sensitive-Content Notice

This role involves regular and deliberate engagement with sensitive imagery. Depending on the project, evaluations may include sexual content and nudity, graphic violence and gore, hate symbols, self-harm, and depictions of real people and of minors in contexts that must be assessed against policy. Some of this material is disturbing by design, because the purpose of the work is to teach models not to produce it.

The work is conducted within structured evaluation frameworks and professional guidelines, with exposure limits, content rotation, mandatory reporting protocols for illegal material, and access to mental health support. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality.

Role Details

  • Location: Seattle, WA

  • Compensation: $36-72/hr

  • Employment classification: W-2

  • Schedule: 8AM - 5PM PT

  • Weekly commitment: M-F

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
406,377 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seattle
$48k – $86k per year (Estimated) • Equity • In office • Full-Time • Bachelor's Degree
Python
SQL
AI/ML
Post-training
Scale AI
Apply
$26k – $60k per year (Estimated) • Equity • In office • Full-Time • 6+ years exp
Python
SQL
Databases
BigQuery
Databricks
Google BigQuery
Snowflake
AI/ML
Post-training
Scale AI
Analytics
Power BI
Tableau
Apply
$130k – $250k per year • Remote • Contractor
Bash
C++
JavaScript
PowerShell
Python
AI/ML
ChatGPT
Claude
Gemini
LLM
LLM Guardrails
Post-training
Red Teaming
Scale AI
Cybersecurity
Cyber Kill Chain
MITRE ATT&CK
Apply
$11k – $14k per year (net) • Remote • Full-Time • Bachelor's Degree • Moscow
AI/ML
ChatGPT
DeepSeek
Midjourney
Design
Canva
Figma
Apply
$12k – $48k per year • Remote • Contractor • 1+ year exp • Buenos Aires • Santa Cruz • Manila • Monterrey • Cairo
AI/ML
ChatGPT
Gemini
Midjourney
Design
Canva
Figma
Management
Google Sheets
Zapier
Marketing
HubSpot
Instagram
LinkedIn
Salesforce
Apply
$48k – $86k per year (Estimated) • Equity • In office • Full-Time • Bachelor's Degree
Python
SQL
AI/ML
Post-training
Scale AI
Apply
$26k – $60k per year (Estimated) • Equity • In office • Full-Time • 6+ years exp
Python
SQL
Databases
BigQuery
Databricks
Google BigQuery
Snowflake
AI/ML
Post-training
Scale AI
Analytics
Power BI
Tableau
Apply
$130k – $250k per year • Remote • Contractor
Bash
C++
JavaScript
PowerShell
Python
AI/ML
ChatGPT
Claude
Gemini
LLM
LLM Guardrails
Post-training
Red Teaming
Scale AI
Cybersecurity
Cyber Kill Chain
MITRE ATT&CK
Apply
$190k – $240k per year • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco
AI/ML
AI Agents
Copilot
Apply
$170k – $215k per year • In office • Full-Time • 5+ years exp • San Francisco
Apply
Founding CX Lead 6 hours ago
$100k – $140k per year • Equity 0.1–0.3% • In office • Full-Time • 3+ years exp • Seattle
SQL
DevOps
Datadog
Apply
$140k – $200k per year • In office • Full-Time • 1+ year exp • Seattle
Apply
$55k – $82k per year • In office • 5+ years exp • Bachelor's Degree • Seattle
Apply
$63k – $94k per year • In office • 2+ years exp • Bachelor's Degree • Seattle
Management
Outlook
Apply
$140k – $157k per year • In office • PhD • Seattle
Apply
See all jobs
This is one of many
406,377 more open roles from verified company boards, updated every day.