368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$179k – $363k per year (Estimated)
Location
In office (New York, San Francisco, London)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Reflection AI is a company founded in 2024 by former Google DeepMind researchers who worked on AlphaGo and large language models. It builds autonomous coding agents and has committed to releasing frontier open-weight models as an American counterweight to Chinese open model releases. The company raised a very large round in 2025 to fund training at frontier scale.

Our Mission

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

Role Overview

The Data Flywheel team closes the gap between benchmark performance and useful performance in the real world. We identify and build the signals, data, and feedback loops that turn model usage into rigorous evaluations, targeted training data, and measurable improvements in future generations of models.

This is a hands-on technical role at the intersection of research and deployment. You'll take ambiguous model behaviors from first observation through measurement, intervention, and validated improvement, working across evaluation, human and synthetic data, infrastructure, post-training, and live deployments. You'll collaborate closely with researchers and engineers across the company, as well as customers, partners, vendors, and the open-source community.

What You’ll Do

  • Identify high-value data sources and partnership opportunities, deeply understand the underlying use cases, and translate them into representative evaluations

  • Bring new data sources online, from initial partner conversations and data scoping through quality validation and integration into production evaluation and training pipelines, where appropriate

  • Design and build evaluations, graders, and feedback loops that make priority real-world model behaviors measurable

  • Analyze model performance and failure modes, then translate those insights into targeted datasets, reward signals, and training interventions

  • Develop human and synthetic data strategies for capabilities where existing data is insufficient, including designing and running evaluation and data-collection programs with vendors

  • Build the infrastructure and pipelines needed to ingest, inspect, version, and evaluate data reliably at scale

  • Collaborate closely with pre-training, post-training, applied, and partnership teams to turn new signals into measurable model improvements

What We’re Looking For

  • Degree (BS, MS, or PhD) in Computer Science, Machine Learning, or related discipline, or equivalent practical experience

  • Deep technical understanding of LLM training and evaluation, with hands-on experience in areas such as evaluation design, data curation, reinforcement learning, or reward design

  • Strong software engineering skills and experience building automated data/evaluation pipelines or large-scale ML systems

  • A track record of owning high-impact projects end to end, navigating ambiguity, and adapting quickly as priorities change

  • A highly collaborative, action-oriented approach and excitement about defining how a new frontier lab measures and accelerates model progress

  • High agency and thrive in a fast-paced startup environment; bias for impact over process

  • Enjoy collaborating across research, engineering, operations, and product disciplines

What We Offer:

We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models.

We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported.

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.

  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.

  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.

  • Meals: Lunch and dinner are provided in the office daily.

  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.

  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.

  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.

  • Team building: We have regular off-sites, happy hours, and team celebrations.

Export Control Notice: This position may require access to technology or source code subject to the U.S. Export Administration Regulations. Any offer of employment for this role may be conditioned on the Company's ability to provide the candidate with access to such technology or source code in compliance with applicable U.S. export control laws, which may require the Company to seek government authorization.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
up to $63k per year (gross) • In office • Full-Time • 5+ years exp • Moscow
SQL
Databases
Apache Kafka
AI/ML
LLM
RAG
DevOps
CI/CD
Git
gRPC
WebSockets
QA
Postman
Swagger
Apply
$43k – $103k per year (Estimated) • In office • Full-Time • 5+ years exp • Moscow
AI/ML
LLM
RAG
Apply
$32k – $77k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Moscow
Python
Databases
FAISS
AI/ML
Hadoop
Hugging Face
LangChain
LLM
NLP
PyTorch
smolagents
Spark
Apply
$62k – $142k per year (Estimated) • In office • Full-Time • 6+ years exp • PhD • Madrid
Apex
JavaScript
Python
TypeScript
Databases
Databricks
Google BigQuery
Snowflake
AI/ML
Agentforce
AI Agents
Claude
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
Marketing
Salesforce
Apply
$133k – $266k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • San Francisco • London • New York
Databases
Apache Kafka
Delta Lake
Google BigQuery
Snowflake
AI/ML
Dagster
Flink
Great Expectations
DevOps
SLI/SLO/SLA
Apply
$134k – $268k per year (Estimated) • Equity • In office • Full-Time • San Francisco • London • New York
AI/ML
NCCL
DevOps
Kubernetes
Apply
$76k – $165k per year (Estimated) • Equity • In office • Full-Time • 5+ years exp • San Francisco • New York
Cybersecurity
Okta
Apply
$202k – $349k per year (Estimated) • Equity • In office • Full-Time • San Francisco • London • New York
AI/ML
LLM
Reinforcement Learning
Post-training
Pre-training
Apply
$174k – $327k per year (Estimated) • Equity • In office • Full-Time • New York • San Francisco • London
Python
AI/ML
LLM Guardrails
DevOps
Platform Engineering
Cybersecurity
Defense in Depth
Least Privilege
Apply
$220k – $350k per year • Remote/Hybrid • Full-Time • 15+ years exp • New York • Princeton
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Azure
Azure DevOps
CI/CD
GitHub
Platform Engineering
Design
Figma
Management
Jira
QA
Playwright
Apply
$80k – $115k per year • In office • Full-Time • PhD • New York
Apply
$60k – $116k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • New York
Apply
UX Research Lead 1 hour ago
$192k – $288k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • New York
AI/ML
Hallucination
Design
Axure RP
Figma
Sketch
InVision
Apply
$120k – $240k per year (Estimated) • Equity • Remote/Hybrid • 5+ years exp • New York
C#
C++
Go
JavaScript
TypeScript
Frontend
React.js
Redux
DevOps
AWS
Azure
GCP
IAM
Cybersecurity
FedRAMP
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.