368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$146k – $274k per year (Estimated)
Location
In office (New York)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Founded in 1898, Sunset Magazine has long covered all aspects of life in the Western United States, focusing in particular on travel, food & drink, home design, and gardening. Based in the Los Angeles area, Sunset is owned by the private equity fi...

About Sunset

At its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses.

In 2025, we had a unique insight: the data every company generates each day through collaboration, communication, and building is some of the most valuable training data in the world. Public and synthetic data can only get frontier models so far, so the next generation of model progress depends on real, proprietary data grounded in how actual businesses operate. We are a primary source of it, partnering directly with the frontier AI labs building what comes next.

Why Join Sunset Now

  • We have scaled from $0 to a multi-eight-figure run rate in a matter of months

  • We have raised from top-tier investors, including Floodgate, Afore, Ludlow, and Hustle Fund

  • We are small enough that you will carry outsized responsibility and grow as quickly as the company does

  • You will partner with and build for some of the fastest and most important companies in the world

  • You will help build a massive, category-defining business from the ground floor

The Role

Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. The data does not arrive in one clean modality. It spans messages, documents, tables, files, images, metadata, and provider-specific structures, with important context distributed across all of them.

You will improve how well our system understands and protects that data. Your initial scope will be a prioritized subset of named-entity recognition, entity and identity resolution, structured extraction, classification, semantic review, or other model-backed parts of the de-identification pipeline. We do not expect one person to be an expert in every modality. The goal is measurable improvement in the areas you own: better precision, recall, F1, high-risk coverage, and preserved data utility across the failure modes that matter.

This is an applied, production-facing ML role. You will study errors, form hypotheses, build datasets and experiments, improve or replace models, and ship the result into a live pipeline. Evaluation, reproducibility, observability, and safe releases matter because they let us identify, ship, and verify meaningful model improvements in production.

What You'll Do

  • Own and improve NER, entity resolution, structured or tabular detection, document understanding, semantic review, or related de-identification systems

  • Transform model failures and capability ceilings into a prioritized improvement roadmap

  • Design active-learning loops that combine model sweeps, LLM-assisted review, clustering, and uncertainty signals to identify the examples most worth hand-labeling

  • Build representative datasets and benchmarks, and use decision-relevant metrics to reveal strengths, weaknesses, uncertainty, and failure costs

  • Choose and combine deterministic rules, classical ML, fine-tuning, embeddings, multimodal models, and LLM-based approaches based on the problem and evidence

  • Design experiments, tune thresholds, analyze precision-recall and utility tradeoffs, and explain which changes are real, uncertain, or limited to particular conditions

  • Productionize improvements with reproducible artifacts, evaluation evidence, runtime instrumentation, and safe rollout

  • Optimize inference cost, latency, and throughput without hiding regressions in quality or high-risk recall

  • Build high-fidelity evaluation environments with seeded failure modes and programmatic verifiers that expose subtle regressions

  • Build reliable model- or agent-based harnesses with bounded behavior and explicit output verification when the problem calls for them

  • Partner with Applied Science on measurement and calibration, Data and Product Engineering on pipeline and review systems, and Security and Quality on acceptable risk

  • Use AI engineering tools deeply to accelerate research, implementation, error analysis, and evaluation while verifying their output

What Success Looks Like

  • Model improvements generalize beyond the examples used to develop them and hold up in replay, shadow, and production evidence

  • Priority modalities and entity classes show credible improvements in precision, recall, F1, or other decision-relevant quality measures

  • High-risk misses decline without unacceptable over-redaction or loss of useful structure

  • New formats and modalities can be covered without relying on brittle one-off fixes

  • Improvements reduce meaningful delivery risk, review or rework burden, or loss of data utility rather than moving only an isolated benchmark

  • The team can explain why a model changed, where it improved or regressed across consequential failure modes and data segments, and whether the change should ship

  • The path from error discovery to a trustworthy production improvement becomes faster and more repeatable

  • Quality gains remain inside acceptable inference-cost, latency, and operational constraints

You Might Thrive Here If

  • You have 3+ years of professional machine learning or software engineering experience, including improving models in production

  • You have startup experience, enjoy broad ownership, and thrive when requirements are evolving or incomplete

  • You use modern AI tools fluently and verify their output

  • You have personally moved model quality through error analysis, data work, experimentation, implementation, deployment, and iteration

  • You have a strong grasp of precision, recall, F1, calibration, thresholding, class imbalance, imperfect labels, distribution shift, and representative evaluation

  • You are an applied engineer first: a strong Python and software engineer who can work inside data pipelines and production systems, not only notebooks

  • You have a bias toward action while maintaining scientific and engineering rigor

  • You are curious and stay current with relevant state-of-the-art methods

  • You choose techniques based on the shape of the problem and can combine deterministic, statistical, neural, and LLM-based approaches

  • You communicate uncertainty and tradeoffs clearly to scientists, engineers, and people making delivery or risk decisions

This Role May Not Be for You If

  • You want to focus on research novelty without owning measurable production improvement

  • You prefer optimizing one aggregate benchmark without investigating consequential failure modes, data segments, and failure costs

  • You want data preparation, evaluation, deployment, and production diagnosis to belong entirely to other teams

  • You reach for a larger model before understanding the errors, constraints, and simpler alternatives

  • You do not want AI tools to be part of your daily engineering and research workflow

Bonus

  • Experience with NER, entity resolution, information extraction, document understanding, multimodal systems, or privacy-preserving ML

  • Experience with hyperparameter tuning, data augmentation, model merging, ensembles, knowledge distillation, or multimodal model training

  • Experience fine-tuning or adapting transformer, GLiNER, embedding, vision-language, or small specialized models

  • Experience with active learning, uncertainty sampling, weak supervision, human-in-the-loop review, or LLM-assisted evaluation pipelines

  • Experience building goldens, adversarial corpora, replay systems, model bakeoffs, agentic harnesses, or programmatic evaluation environments

  • Experience with difficult ML or labeling problems

  • Experience with ONNX Runtime, TensorRT, model pruning, quantization, or other CPU/GPU inference optimization

  • Experience with sensitive enterprise data or other high-trust production systems

  • Experience with synthetic data generation and managing the synth-to-real gap

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
Staff Data Engineer 5 hours ago
$160k – $200k per year • In office • Full-Time • 6+ years exp • Chicago
Python
SQL
Databases
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Arize Phoenix
AutoGen
AWS Bedrock AgentCore
CrewAI
dbt
Fine-tuning
Function Calling
LangChain
LangGraph
LangSmith
LLM
LLM Evaluation
LLM Guardrails
Model Context Protocol
Prefect
Prompt Engineering
RAG
Semantic Kernel
Semantic Search
Semantic Search
Weights & Biases
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
Kubernetes
Vector
Analytics
ETL/ELT
Apply
$118k – $168k per year • In office • Full-Time • Ottawa • Halifax
Databases
Snowflake
AI/ML
AI Agents
AutoGen
CrewAI
Human-in-the-Loop
LangChain
LLM
DevOps
AWS
Azure
Apply
$73k – $183k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Zug
Python
Python
FastAPI
Databases
PostgreSQL
Redis
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
DevOps
Amazon EKS
AWS
AWS CDK
CI/CD
Datadog
Kubernetes
OpenTelemetry
Platform Engineering
Apply
$136k – $253k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Frisco • New York • Toronto • Ann Arbor
Python
SQL
Java
Java
Flyway
AI/ML
AWS Bedrock
Claude
LLM
Anthropic
AWS Bedrock AgentCore
LLM Guardrails
LLMOps
DevOps
Amazon EKS
AWS
CI/CD
Datadog
Docker
Kubernetes
Apply
$24k – $63k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • India
Crystal
Groovy
JavaScript
Perl
Python
Ruby
SQL
TypeScript
Java
Java
Apache Tomcat
Gradle
Hibernate
Maven
Spring Boot
Spring MVC
Databases
Apache Kafka
Db2
Oracle
PostgreSQL
RabbitMQ
AI/ML
Fine-tuning
Frontend
Angular
JQuery
DevOps
Apache HTTP Server
AWS
Azure
CI/CD
Docker
GCP
Jenkins
Kubernetes
Rest API
Cybersecurity
Checkmarx
SonarQube
Apply
Security Lead 18 days ago
$135k – $293k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
Synthetic Data
Model Context Protocol
Cybersecurity
SOC 2
Least Privilege
Apply
Platform Engineer 18 days ago
$117k – $263k per year (Estimated) • In office • Full-Time • New York
AI/ML
Synthetic Data
DevOps
AWS
CI/CD
Platform Engineering
Progressive Delivery
Terraform
Cybersecurity
Least Privilege
SOC 2
Apply
Engineering Manager 18 days ago
$179k – $343k per year (Estimated) • In office • Full-Time • 2+ years exp • New York
AI/ML
Synthetic Data
Apply
Data Scientist 18 days ago
$99k – $206k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
Python
SQL
AI/ML
LLM
Multimodal AI
NER
Synthetic Data
Apply
AI Product Engineer 18 days ago
$136k – $273k per year (Estimated) • In office • Full-Time • New York
Node JS
Python
TypeScript
JavaScript
AI/ML
LangChain
LangGraph
LLM
Prompt Engineering
Synthetic Data
AI Agents
Function Calling
LLM Guardrails
Structured Outputs
Frontend
React.js
Apply
$220k – $350k per year • Remote/Hybrid • Full-Time • 15+ years exp • New York • Princeton
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Azure
Azure DevOps
CI/CD
GitHub
Platform Engineering
Design
Figma
Management
Jira
QA
Playwright
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • New York
AI/ML
AI Agents
Apply
$80k – $115k per year • In office • Full-Time • PhD • New York
Apply
$60k – $116k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • New York
Apply
UX Research Lead 2 hours ago
$192k – $288k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • New York
AI/ML
Hallucination
Design
Axure RP
Figma
Sketch
InVision
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.