368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$99k – $206k per year (Estimated)
Location
In office (New York)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Founded in 1898, Sunset Magazine has long covered all aspects of life in the Western United States, focusing in particular on travel, food & drink, home design, and gardening. Based in the Los Angeles area, Sunset is owned by the private equity fi...

About Sunset

At its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses.

In 2025, we had a unique insight: the data every company generates each day through collaboration, communication, and building is some of the most valuable training data in the world. Public and synthetic data can only get frontier models so far, so the next generation of model progress depends on real, proprietary data grounded in how actual businesses operate. We are a primary source of it, partnering directly with the frontier AI labs building what comes next.

Why Join Sunset Now

  • We have scaled from $0 to a multi-eight-figure run rate in a matter of months

  • We have raised from top-tier investors, including Floodgate, Afore, Ludlow, and Hustle Fund

  • We are small enough that you will carry outsized responsibility and grow as quickly as the company does

  • You will partner with and build for some of the fastest and most important companies in the world

  • You will help build a massive, category-defining business from the ground floor

The Role

Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. That creates a difficult measurement problem. A system can improve aggregate F1 while missing a high-risk slice, remove more sensitive information while also destroying useful context, or pass one stage while defects escape somewhere else in the pipeline.

As Sunset's first Data Scientist focused on evaluation, you will establish how we know whether that data is actually getting better. You will build the datasets, experiments, quality measures, and feedback loops that expose hidden failures, accelerate model and pipeline improvement, and give the team confidence in what it delivers.

This is a hands-on, zero-to-one role at the intersection of data science, AI, and a real production system. You will write Python and SQL, construct evaluation corpora, study failure patterns, design comparisons, calibrate human and model-based judgments, and turn the result into a clear decision. The questions are scientifically difficult, but the output must be practical enough to change what the team builds and ships.

You will work closely with Machine Learning, Product Engineering, Data Engineering, Security, Quality, domain experts, and the team making delivery decisions. Machine Learning Engineers own changing model behavior. You own the credibility of the evidence used to decide whether a model, pipeline, or delivery change actually made the data safer or more useful.

Questions You Might Answer

  • Did a higher NER or entity-resolution score actually reduce sensitive misses across the messages, documents, tables, and providers that matter?

  • Is a new model finding more sensitive information, or simply removing more of the useful structure our customers need?

  • Can we trust a golden dataset, a human review process, or an LLM judge enough to use it for a release decision?

  • Which customer, modality, entity, language, or format slices are hidden by a strong aggregate result?

  • Where did a quality loss enter between source data, processing, de-identification, review, and delivery?

  • What is the smallest credible experiment that would tell us whether to ship, revise, or stop a change?

What You'll Do

  • Define what high-quality and safe-to-deliver data mean across de-identification, structure preservation, semantic coherence, and customer utility

  • Design representative samples and build golden, adversarial, replay, and production-like corpora with explicit provenance, labeling policy, agreement, adjudication, and versioning

  • Turn ambiguous concepts such as “useful,” “clean,” or “safe” into measurable claims with known uncertainty and clear decision consequences

  • Evaluate detectors, models, prompts, judges, thresholds, review workflows, and pipeline changes using comparisons that can support a real decision

  • Break aggregate results into the modalities, providers, entity classes, customer contexts, languages, formats, and risk tiers that reveal consequential failures

  • Connect local measures to escaped sensitive information, avoidable over-redaction, preserved data utility, review burden, rework, and delivery acceptance

  • Build reproducible analysis, evaluation pipelines, and high-fidelity environments using Python, SQL, synthetic data, historical replay, seeded failures, and programmatic verifiers

  • Establish holdout and evaluation practices that keep the evidence trustworthy while model and product teams iterate quickly

  • Use modern AI tools deeply for analysis, corpus development, coding, review, and hypothesis generation while independently verifying their output

What Success Looks Like

  • The team has a decision-grade baseline for a priority Clean Data quality claim and trusts it enough to use in model, pipeline, release, and delivery decisions

  • Improvements are judged by the slices and failure costs that matter, not only by an aggregate benchmark

  • The company can distinguish a true gain from label noise, sample bias, leakage, evaluator error, or a shifted workload

  • Changes that improve one stage cannot hide escaped defects, over-redaction, utility loss, or review burden somewhere else

  • At least one consequential decision changes because the evidence reveals a risk, tradeoff, or opportunity that was previously unclear

  • Evaluation becomes faster and more repeatable without sacrificing independence or rigor

  • Quality claims communicate uncertainty honestly and remain understandable to engineers, customers, and risk owners

You Might Thrive Here If

  • You have at least three years of professional experience in applied science, data science, machine learning, quantitative research, or a closely related role

  • You have designed evaluations or experiments that changed a product, model, release, or operational decision

  • You understand sampling, uncertainty, precision, recall, F1, calibration, agreement, class imbalance, distribution shift, and imperfect labels

  • You can investigate messy, multi-stage data systems and determine where an apparent gain or loss actually came from

  • You are comfortable writing Python and SQL and building reproducible technical artifacts rather than handing requirements to an engineering team

  • You can protect the independence of an evaluation while collaborating closely with the people whose work it evaluates

  • You have startup experience and enjoy broad ownership, changing context, and building the measurement foundation while decisions are already moving quickly

  • You use AI tools fluently but do not confuse an articulate model output with valid evidence

  • You communicate uncertainty and difficult findings directly, without hiding behind false precision

This Role May Not Be for You If

  • You want to optimize models as your primary job rather than determine whether changes actually improve delivered data

  • You prefer descriptive dashboards that stop short of changing a decision

  • You treat labels, benchmarks, or model-based judges as ground truth without investigating how they fail

  • You need a perfectly defined dataset and research plan before you can make progress

  • You are uncomfortable disagreeing with a technically strong team when the evidence does not support its conclusion

  • You do not want AI tools to be part of your daily scientific and technical workflow

Bonus

  • Experience evaluating NER, entity resolution, information extraction, document understanding, multimodal, retrieval, or LLM systems

  • Experience with privacy, de-identification, data quality, model risk, safety, or other high-trust decision systems

  • Experience designing human-review, adjudication, weak-supervision, or active-learning systems

  • Experience building adversarial corpora, replay systems, simulation environments, programmatic verifiers, or model-judge evaluations

  • Experience connecting offline measures to escaped defects, customer outcomes, review effort, or preserved data utility

  • Experience measuring quality across multi-stage batch or data pipelines

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$19k – $28k per year (net) • Remote/Hybrid • Contractor • 2+ years exp • Samara
JavaScript
Python
SQL
Databases
Apache Kafka
MS SQL
PostgreSQL
AI/ML
Claude
Claude Code
Copilot
Cursor
Model Context Protocol
DevOps
Git
GitLab
Kibana
Design
Figma
Management
Confluence
Draw.io
Jira
Miro
QA
Postman
Swagger
Apply
Full Stack Engineer 11 hours ago
$47k – $103k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad • Bengaluru
Java
SQL
TypeScript
JavaScript
Frontend
Angular
DevOps
AWS
Azure
Apply
Lead Java Developer 11 hours ago
$31k – $81k per year (Estimated) • In office • Full-Time • 5+ years exp • Kochi
Java
SQL
Kotlin
Java
SLF4J
Spring Boot
Kotlin
Mockito
DevOps
CI/CD
Cybersecurity
SonarQube
Apply
Full Stack Engineer 11 hours ago
$47k – $102k per year (Estimated) • In office • Full-Time • 12+ years exp • Pune • Chennai • Hyderabad
JavaScript
SQL
DevOps
Azure
Apply
$20k – $43k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Moscow
SQL
Analytics
Power BI
Apply
Security Lead 18 days ago
$135k – $293k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
Synthetic Data
Model Context Protocol
Cybersecurity
SOC 2
Least Privilege
Apply
Platform Engineer 18 days ago
$117k – $263k per year (Estimated) • In office • Full-Time • New York
AI/ML
Synthetic Data
DevOps
AWS
CI/CD
Platform Engineering
Progressive Delivery
Terraform
Cybersecurity
Least Privilege
SOC 2
Apply
$146k – $274k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
Python
AI/ML
Data Augmentation
Embeddings
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
NER
ONNX
Quantization
Synthetic Data
TensorRT
Human-in-the-Loop
AI Agents
Apply
Engineering Manager 18 days ago
$179k – $343k per year (Estimated) • In office • Full-Time • 2+ years exp • New York
AI/ML
Synthetic Data
Apply
AI Product Engineer 18 days ago
$136k – $273k per year (Estimated) • In office • Full-Time • New York
Node JS
Python
TypeScript
JavaScript
AI/ML
LangChain
LangGraph
LLM
Prompt Engineering
Synthetic Data
AI Agents
Function Calling
LLM Guardrails
Structured Outputs
Frontend
React.js
Apply
$197k – $374k per year (Estimated) • In office • Full-Time • 8+ years exp • PhD • New York
AI/ML
AI Agents
Claude
LangChain
OpenAI
Vertex AI
Management
n8n
Zapier
Apply
$96k – $134k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • New York
JavaScript
Swift
TypeScript
Java
Java
Spring Framework
Databases
Apache Kafka
PostgreSQL
AI/ML
AI Agents
Claude
Copilot
Fine-tuning
Flink
LangChain
LangGraph
Llama
LlamaIndex
Prompt Engineering
PyTorch
RAG
TensorFlow
Transformers
Devin
Hugging Face
OpenAI
Frontend
Angular
React.js
Mobile
MVC
DevOps
AWS
CI/CD
Docker
Kubernetes
OpenShift
Splunk
Vector
GitHub
Analytics
Tableau
Apply
$150k – $180k per year • In office • Full-Time • PhD • New York
Python
AI/ML
Anthropic
Anthropic SDK
Computer Vision
Fine-tuning
LangChain
LlamaIndex
LLM
OpenAI
OpenAI SDK
RAG
DevOps
AWS
Azure
GCP
Apply
Senior AI Architect 2 hours ago
$131k – $136k per year • In office • Full-Time • 4+ years exp • Master's Degree • New York
Python
Databases
Databricks
AI/ML
Anthropic
Computer Vision
EU AI Act
LLMOps
OpenAI
DevOps
AWS
Azure
GCP
Terraform
Cybersecurity
GDPR
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.