998,188open jobs
59,534companies
165,604added this week
Browse all
Salary
$300k – $350k per year
Location
In office (Mountain View)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026.

Overview
Company
Impact
Profile match
Bespoke Labs optimizes AI agents using cutting-edge research on data curation and evolutionary algorithms. Trusted by Fortune 500 enterprises and frontier labs.

About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.

Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.

About the Role

This is a delivery-heavy role focused on standing up our enterprise post-training capability. Demand is inbound, and our goal is to ship custom, high-performing models to 2-3 paying enterprise customers by year-end. Over the first 6-12 months, you will do the execution work required to solve real-world enterprise problems while helping build the scalable product underneath.

You will not be running abstract experiments or training models solely for benchmarks. You will sit directly at the intersection of enterprise demand and applied post-training-curating data, building rigorous eval suites, fine-tuning models, and proving their value to enterprise stakeholders. We will measure you on the production impact, robustness, and delivery of models shipped to real users, not on published papers.

The thing we care about most is whether you have done this before. If you have post-trained an LLM, shipped it to production users, managed regression risks, and owned the evals from end-to-end, we want to talk.

What You'll Do

  • Ship enterprise-grade models. Post-train, fine-tune, and align open-weight and proprietary base models for complex business domains, ensuring they perform reliably in production.

  • Build bespoke evaluation suites. Define what "quality" means for subjective domain-specific tasks, create custom benchmarks, and calibrate LLM judges against human domain experts.

  • Curate and filter high-impact datasets. Build and run production data flywheels combining real production traces, human labeling, and synthetic augmentation with strict filtering standards.

  • Manage and mitigate regression risk. Rigorously track downstream performance to ensure fine-tuning for new behaviors doesn't silently degrade core capabilities or reasoning.

  • Engage directly with stakeholders. Sit in front of product and enterprise customers to understand their requirements, translate vague domain preferences into technical eval metrics, and explain model behavior clearly.

  • Deploy for cost and privacy. Fine-tune open-weight architectures (e.g., Llama, Qwen, Mistral, DeepSeek) to hit strict enterprise latency, cost, and privacy targets.

  • Direct frontier tools and workflows. Leverage state-of-the-art post-training techniques, preference tuning, and data curation tooling to maximize output quality and delivery speed.

What We're Looking For

  • A record of shipped models. You have post-trained at least one LLM that was deployed to real users in production, and you can show how you measured its success.

  • End-to-end eval ownership. Demonstrated experience building benchmarks, creating eval datasets, and getting stakeholders to agree on clear metrics for complex or subjective tasks.

  • Deep understanding of regression risks. You can instinctively explain how training a model on new behaviors impacts existing capabilities and how to prevent it.

  • Customer-facing or product empathy. Experience collaborating directly with non-ML stakeholders, enterprise customers, or product managers to turn requirements into model behavior.

  • Strong software and ML fundamentals. Fluency in modern post-training frameworks, data processing pipelines, and code bases designed for production deployment.

  • Ownership mindset. You take complete responsibility for the full post-training lifecycle-from raw data to model deployment and failure analysis-without requiring close supervision.

You May Be a Good Fit If You Also

  • Have post-trained conversational or task-oriented assistants (e.g., support agents, multi-turn chat, tool-using agents)

  • Have built LLM judges or reward models and calibrated them against human raters

  • Have operated a production data flywheel: traces → labeling → synthetic augmentation → retrain

  • Have extensive hands-on experience with open-weight models (Llama, Qwen, Mistral, DeepSeek) for cost, latency, or privacy optimization

  • Come from forward-deployed engineer (FDE), founder, or early-stage startup backgrounds

What We Offer

  • Location: Mountain View, CA (Preferred) or San Francisco, CA (Onsite); Remote considered

  • Base Salary: $300,000 - $350,000 USD / year

  • Additional Comp: 25% performance-based bonus + equity

Benefits & Perks:

  • Health, dental, and vision coverage

  • 401(k)

  • Daily onsite lunch provided

  • Visa sponsorship and relocation support available

  • Direct impact on how the industry trains and evaluates agents

We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
998,188 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Mountain View
$181k – $323k per year • Equity • In office • Full-Time • 4+ years exp • PhD • San Francisco
AI/ML
Reinforcement Learning
Multimodal AI
Computer Vision
VLM
Synthetic Data
World Models
Apply
LLM Platform Engineer 15 days ago
$172k – $329k per year • Equity • In office • Full-Time • 4+ years exp • San Francisco
Python
AI/ML
LoRA
Fine-tuning
Embeddings
Multimodal AI
Knowledge Distillation
Computer Vision
AI Agents
PEFT
Transformers
LLM
RAG
Reranking
Hybrid Search
Text-to-Speech
Model Distillation
DevOps
CI/CD
AWS
Apply
≈ $70k – $162k per year (Estimated) • In office • Full-Time • 1+ year exp • PhD • Prairie View
Python
AI/ML
Scikit-learn
Computer Vision
TensorFlow
Keras
PyTorch
Ray
Machine Learning
Analytics
Seaborn
Matplotlib
Plotly
Apply
$150k – $200k per year • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree • United States
Python
AI/ML
Fine-tuning
JAX
Multimodal AI
Computer Vision
AI Agents
TensorFlow
PyTorch
Machine Learning
DevOps
GCP
Apply
≈ $139k – $263k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • United States
Python
SQL
C++
AI/ML
Machine Learning
Apply
$75k – $125k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Chicago
Python
SQL
Databases
Weaviate
Pinecone
Apache Kafka
AI/ML
LangGraph
AutoGen
LangChain
Spark
Fine-tuning
AI Agents
Langfuse
Ragas
Semantic Kernel
Falcon
AWS Bedrock
CrewAI
LLM
Hybrid Search
LLMOps
Human-in-the-Loop
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Analytics
Tableau
Power BI
ETL/ELT
Looker
Apply
$158k – $193k per year • Hybrid • Full-Time • 5+ years exp • Master's Degree • United States
Python
SQL
MATLAB
SAS
SPSS
Databases
Snowflake
Databricks
AI/ML
Copilot
Unstructured.io
Claude Code
Fine-tuning
Prompt Engineering
LLM
RAG
Human-in-the-Loop
Machine Learning
DevOps
GCP
Azure
AWS
Cybersecurity
HIPAA
Apply
$99k – $225k per year • In office • TS/SCI • Full-Time • 2+ years exp • Bachelor's Degree • Dayton
Python
Java
Rust
Scala
AI/ML
CUDA Toolkit
Reinforcement Learning
Scikit-learn
Computer Vision
NLP
TensorFlow
Pandas
NumPy
PyTorch
LLM
RAG
Hallucination
CUDA
Machine Learning
DevOps
GCP
Podman
Azure
CI/CD
AWS
Docker
Kubernetes
GitLab
Apply
Data Engineer 1 day ago
$99k – $225k per year • In office • TS/SCI • Full-Time • 6+ years exp • High School Diploma • Fort Belvoir
Python
PowerShell
Databases
ElasticSearch
AI/ML
Fine-tuning
Anomaly Detection
DevOps
CI/CD
AIOps
Amazon ECS
Cybersecurity
Zero Trust
SIEM
Apply
≈ $20k – $45k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bengaluru
Databases
ElasticSearch
AI/ML
Fine-tuning
DevOps
SLURM
Jenkins
Linux
Windows
TCP/IP
Apply
$35k – $55k per year • Remote (LATAM, United States) • Full-Time
AI/ML
Reinforcement Learning
Tool Use
Apply
$250k – $300k per year • In office • Full-Time • Mountain View
AI/ML
Reinforcement Learning
AI Agents
Post-training
Tool Use
DevOps
CI/CD
Apply
Founding Recruiter 8 months ago
$150k – $225k per year • In office • Full-Time • 3+ years exp • Mountain View
AI/ML
Reinforcement Learning
AI Agents
Scale AI
Tool Use
Apply
≈ $213k – $441k per year (Estimated) • In office • 6+ years exp • PhD • Mountain View
AI/ML
RLHF
Gemini
SFT
Post-training
Apply
≈ $174k – $329k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Mountain View
Python
C++
AI/ML
Reinforcement Learning
Apply
≈ $171k – $324k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Mountain View
AI/ML
Fine-tuning
Multimodal AI
Computer Vision
Gemini
Machine Learning
Apply
$234k – $350k per year (gross) • Equity • Hybrid • Full-Time • New York • Mountain View
AI/ML
AI Agents
EU AI Act
Cybersecurity
GDPR
Game Dev
Unity
Apply
See all jobs
This is one of many
998,188 more open roles from verified company boards, updated every day.