368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$172k – $322k per year (Estimated)
Location
Remote (United States)
Seniority
Staff · 10+ years exp
Overview
Company
Impact
Profile match
Airbnb is an American technology company that operates a global online marketplace for short-term and long-term homestays, vacation rentals, and local experiences. Founded in 2008 and headquartered in San Francisco, the platform connects property hosts with travelers across more than 220 countries and regions. Operating on an asset-light, peer-to-peer business model, the company facilitates search, secure payment processing, identity verification, and customer support for millions of guest arrivals worldwide.

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way.

The Community You Will Join: 

AI and ML are at the heart of the Airbnb product. From Trust to Payments, and from Customer Service to Marketing, we rely on ML to ensure that guests and hosts have the best possible experience with Airbnb.

The Core ML team is responsible for driving CSxAI (Customer Support x Artificial Intelligence) initiatives by adopting Generative AI technologies to enable an intelligent, scalable, and exceptional service experience. The team develops and enhances AI models, ML services, and tools including LLM fine-tuning and optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning, and guardrails for a wide range of applications at Airbnb.

The richness of Airbnb's data, the complexity of its marketplace, and the variety innate in our product mean that we need to operate at the state of the art of AI practice. We are committed to long-term innovation to solve complex problems, and to do that we need experienced ML

The Difference You Will Make:

In this Senior Staff role, you will set technical direction and lead execution for ML evaluation and the end-to-end data flywheel powering CSxAI products (e.g., assistive agents, issue resolution, and tooling). Your work will define how we measure quality, how we turn feedback into learning signals, and how we continuously improve models and products safely and efficiently. You will partner closely with product, engineering, design, operations to build evaluation systems that are trusted, scalable, and actionable - connecting offline metrics to online outcomes.

A Typical Day: 

  • Define evaluation strategy and success metrics for GenAI systems, aligning offline evaluation with online business and customer experience outcomes.
  • Build and scale evaluation frameworks (golden sets, synthetic data, automated regressions, rubric-based grading, LLM-as-judge where appropriate) with strong controls for bias, drift, and reliability.
  • Design the data flywheel: instrumentation, feedback collection, data quality checks, labeling strategy, dataset versioning, and governance to support continuous improvement.
  • Lead cross-functional quality initiatives across product, ops, and engineering, driving clarity on what “good” looks like and how teams act on evaluation results.
  • Develop and productionize pipelines for dataset creation, model monitoring, evaluation-at-scale, and continuous testing (pre-deploy and post-deploy).
  • Drive technical decisions and architecture for evaluation and data infrastructure, balancing speed, rigor, cost, and safety.

Minimum Qualifications:

  • Educational Background: PhD in Computer Science, Mathematics, Statistics, or related technical field (or equivalent practical experience).
  • Industry Experience: 10+ years building, testing, and shipping ML/AI systems end-to-end; including 2+ years of experience with GenAI/LLM systems in production.
  • Leadership Experience: 5+ years leading large, ambiguous technical initiatives as a senior IC, influencing roadmap and engineering/science direction across teams.
  • Technical Proficiency:
    • Deep expertise in evaluation methodology (offline/online alignment, metric design, human-in-the-loop evaluation, A/B testing, power analysis, regression testing).
    • Hands-on experience with GenAI systems, including orchestration, retrieval, tool calling, memory, etc.
    • Experience building data pipelines and quality systems (labeling workflows, dataset curation, versioning, monitoring, and governance).
    • Solid ML fundamentals and best practices (model selection, training/serving, monitoring, reliability, and model lifecycle management).

Preferred Qualifications:

  • Customer Support Systems: Experience applying ML/AI to customer support workflows (e.g., agent assist, classification/routing, resolution recommendation, QA).
  • Infrastructure & Quality at Scale: Experience building robust evaluation platforms for agent behavior validation, safety/guardrails, and continuous improvement.
  • Agile Practice for Applied AI: Proven ability to take evaluation and data flywheel work from incubation to production, iterating quickly while maintaining scientific rigor.

Your Location:

This position is US - Remote Eligible. The role may include occasional work at an Airbnb office or attendance at offsites, as agreed to with your manager. While the position is Remote Eligible, you must live in a state where Airbnb, Inc. has a registered entity. Click here  for the up-to-date list of excluded states. This list is continuously evolving, so please check back with us if the state you live in is on the exclusion list. If your position is employed by another Airbnb entity, your recruiter will inform you what states you are eligible to work from.

How We'll Take Care of You:

Our job titles may span more than one career level. The actual base pay is dependent upon many factors, such as: training, transferable skills, work experience, business needs and market demands. The base pay range is subject to change and may be modified in the future. This role may also be eligible for bonus, equity, benefits, and Employee Travel Credits.  

Pay Range

$244,000—$305,000 USD

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$47k – $106k per year (Estimated) • In office • Moscow
AI/ML
Fine-tuning
LangChain
LLM
NLP
RAG
Transformers
Knowledge Graph
AI Agents
Apply
$37k – $111k per year (Estimated) • Equity • Remote/Hybrid • Internship • 1+ year exp • Bachelor's Degree • Madrid
Python
SQL
Databases
Databricks
AI/ML
AI Agents
Fine-tuning
Function Calling
LangChain
LlamaIndex
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Context Engineering
Edge AI
OpenAI
Robotics
Digital Twin
Apply
$50k – $146k per year (Estimated) • Equity • Remote/Hybrid • Internship • 1+ year exp • Bachelor's Degree • Munich
Python
SQL
Databases
Databricks
AI/ML
AI Agents
Fine-tuning
Function Calling
LangChain
LlamaIndex
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Context Engineering
Edge AI
OpenAI
Robotics
Digital Twin
Apply
$16k – $43k per year (Estimated) • In office • Master's Degree • Saint Petersburg
AI/ML
ChatGPT
Claude
Claude Code
Cursor
LLM
RAG
Lovable
Replit
DevOps
SLI/SLO/SLA
Management
Jira
n8n
Zapier
Marketing
Zendesk
Apply
$39k – $78k per year (Estimated) • Remote/Hybrid • Moscow
AI/ML
GigaChat
LLM
YandexGPT
Apply
$172k – $286k per year (Estimated) • Remote • 9+ years exp
Databases
Apache Kafka
Druid
MySQL
Redis
AI/ML
Flink
DevOps
ZooKeeper
Apply
$111k – $190k per year (Estimated) • Remote • 8+ years exp • Master's Degree
Python
SQL
AI/ML
ChatGPT
Claude
Analytics
A/B Testing
Apply
$148k – $266k per year (Estimated) • Remote/Hybrid • 8+ years exp • San Francisco
Management
Airtable
Asana
Jira
Apply
$141k – $253k per year (Estimated) • Remote • 5+ years exp • PhD
Java
Python
Databases
Apache Kafka
AI/ML
Fine-tuning
Hallucination
Kubeflow
LLM
LoRA
Prompt Engineering
PyTorch
Ray
RLHF
Spark
TensorFlow
PEFT
LLM Guardrails
DevOps
Vector
Apply
$175k – $326k per year (Estimated) • Remote • 10+ years exp • PhD
Python
AI/ML
Fine-tuning
LLM
Multimodal AI
PyTorch
Quantization
RAG
Reinforcement Learning
Edge AI
LLM Guardrails
Post-training
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.