1,099,169open jobs
64,006companies
186,251added this week
Browse all
Salary
$180k – $280k per year
Location
In office (Beverly Hills)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Jul 29, 2026. StudyFetch scores B on the Alion truth index.

Overview
Company
Impact
Profile match
StudyFetch turns course materials into flashcards, quizzes and tutoring sessions. Its assistant answers questions grounded in the uploaded content. The company serves students and educators.

About Studyfetch

StudyFetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered education. We’re growing fast with backing from top-tier investors and a mission that’s redefining the future of education and ethical learning.

Why this role exists

We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most.

Nobody has settled how to measure whether an AI tutor teaches. Public benchmarks tell you a model can answer a question. They don't tell you whether a fourteen-year-old understood the explanation, whether the model handed over the answer when it should have asked a follow-up, or whether the student could still do the problem a week later. We train and tune our own models for learning outcomes rather than leaderboard scores, and that only works if someone can define what better means and prove when we've hit it.

That's this role. You'll build the evaluation and measurement layer that sits under model training, product decisions, and learning science across both StudyFetch and Honen. This is a founding-team role. You'll work directly with the people making the decisions, and the standard you set is the one every model and feature gets held to.

What we believe

Every learner deserves the chance to succeed. StudyFetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce.

Accessible to everyone. Learning should reach every person, whatever their background, role, or resources.

Meet people where they are. Every course adapts to a person's pace, their level, and the way they learn best.

Learning never stops. From a first job to a new career, people keep growing at every stage of life.

We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material.

What you'll own

Evaluation for the models we train. We fine-tune our own model family for tutoring, and you decide how we know whether a new checkpoint is better than the last one. That covers accuracy and reasoning, and it also covers child safety, resistance to sycophancy, and whether the model teaches Socratically instead of answering outright. You'll design the evals, run them against every candidate model, and hold the release bar.

An internal benchmark for multi-turn tutoring. Single-turn Q&A benchmarks miss almost everything we care about. You'll build and maintain our benchmark for real tutoring conversations, decide what it measures, and defend those choices to researchers outside the company. Expect to publish parts of it.

The link between product data and model training. The Learn Engine records what worked for past students at the personal, course, topic, and global level. You'll turn that into training signal and into evidence: which interventions moved mastery, which ones only moved engagement, and which model behaviors correlate with a student actually learning.

Data quality for expert-verified content. We build assessment questions with subject-matter experts, starting in nursing licensure and expanding into medical and legal. You'll measure agreement between experts, catch where the official answer key is out of date, and design how those verifications feed back into training.

Analytics across both products. Retention, activation, feature adoption, conversion, and how each of those moves when a model changes. You'll build the dashboards and reporting that product and leadership actually use, and you'll say plainly when the data can't answer the question yet.

Instrumentation we don't have. You'll find the missing events, telemetry, and logging, then work with engineering to add them. Most of the interesting questions here are currently unanswerable because nobody logged the right thing.

The measurement bar for the team. How we run experiments, what counts as a result, when a change ships. The patterns you set are the ones the rest of the team follows.

What we're looking for

You're a strong fit if either of these is true:

  • A PhD in statistics, computer science, machine learning, economics, physics, computational social science, or another quantitative field, plus 5+ years applying it to real products or research, or
  • Fewer credentials on paper and a track record of owning evaluation or measurement for an AI product that shipped to real users. Show us the work.

Beyond that:

You've evaluated LLMs in production, not just read about it. You can walk us through an eval suite you built, what it caught, what it missed, and what you'd design differently now.

You're rigorous about causality. Experimental design, controls, sample size, and knowing when an observational result is all you're going to get. You can tell the difference between a real effect and a dashboard that moved.

You write and speak clearly. You'll present to engineers, to the founders, and to external research partners in the same week. You state uncertainty as a number when you can and in plain words when you can't.

You use AI every day and have informed opinions about it. Genuine curiosity is the one thing we can't teach.

The mission is why you're here. What sustains you through the hard weeks is the learner on the other end, the one who finally understands because of something you built.

The stack you'll work in

You don't need every item below, but you should be deep in most and able to ramp quickly on the rest:

  • Analysis: Python (Pandas, NumPy, SciPy, Jupyter), expert-level SQL, large-scale datasets
  • Statistics: experimental design, causal inference, Bayesian and frequentist methods, hypothesis testing
  • AI/LLM: eval frameworks, LLM-as-judge and its failure modes, RAG, embeddings, agent workflows, fine-tuning and post-training
  • Data: MongoDB and PostgreSQL, vector databases, warehouse and pipeline tooling
  • Infra: GCP, and comfort reasoning about inference cost, latency, throughput, and GPU utilization
  • Reporting: dashboards people return to, in whatever tool gets there fastest

Learning science, psychometrics, or item response theory is a real plus. So is having worked with children's data and the rules that come with it.

What to expect

This is an in-person role at a fast pace, with periods of intense work around major launches.

You'll have significant ownership and autonomy with limited oversight. The role suits researchers who do their best work with room to run.

You'll be the first person in this seat. Some weeks are model evaluation, some weeks are a retention question from the founders, and you'll have to decide which one matters more that week.

It's a strong fit for people who have shipped analysis that changed a decision. If your experience has been mostly reports that nobody acted on, this likely isn't the right match.

Compensation & benefits

$180,000-$280,000 base salary, plus equity

  • 100% employer-paid Medical, Dental, and Vision; 75% dependent coverage
  • 401(k) with employer matching
  • Daily team dinner provided in-office
  • A small, mission-driven team changing how the world learns
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,099,169 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Beverly Hills
Développeur IA senior 6 months ago
≈ $91k – $201k per year (Estimated) • In office • Montreal
Python
AI/ML
LangGraph
LangChain
Scikit-learn
TensorFlow
Pandas
NumPy
PyTorch
LLM
RAG
DevOps
Azure
AWS
Apply
$84k – $108k per year • Hybrid • 5+ years exp • Bachelor's Degree • Vancouver
AI/ML
Model Context Protocol
Embeddings
LLM
RAG
LLMOps
DevOps
Azure
CI/CD
Apply
≈ $103k – $219k per year (Estimated) • Remote (Canada) • Full-Time • 5+ years exp • Bachelor's Degree • Canada
C#
C#
.NET
Databases
PostgreSQL
AI/ML
Model Context Protocol
AI Agents
DevOps
Terraform
CI/CD
AWS
Platform Engineering
Apply
≈ $98k – $217k per year (Estimated) • In office • Full-Time • 5+ years exp • Markham
AI/ML
Embeddings
LLM
RAG
Apply
≈ $90k – $200k per year (Estimated) • In office • 7+ years exp • Toronto
Python
JavaScript
TypeScript
AI/ML
Copilot
Model Context Protocol
AI Agents
LLM
RAG
LLM Guardrails
DevOps
GitHub Actions
Azure
CI/CD
Docker
Kubernetes
Apply
≈ $101k – $226k per year (Estimated) • In office • Bachelor's Degree • Columbus
Python
SQL
C#
DevOps
TCP/IP
Apply
$125k – $150k per year • In office • 5+ years exp • Bachelor's Degree • Canada
JavaScript
TypeScript
SQL
C#
Node JS
C#
.NET
DevOps
GCP
Datadog
Kubernetes
GitLab
Management
Confluence
Jira
Agile
Scrum
Kanban
Apply
≈ $81k – $153k per year (Estimated) • In office • Full-Time • 3+ years exp • San Jose
Python
AI/ML
InfiniBand
Apply
≈ $96k – $191k per year (Estimated) • In office • 5+ years exp • Alpharetta
SQL
Databases
ClickHouse
AI/ML
Claude
AI Agents
Cybersecurity
Crowdstrike
Marketing
Salesforce
Apply
Metallurgist 1 day ago
≈ $89k – $171k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Columbus
SQL
Management
Outlook
Apply
$70k – $100k per year • In office • Part-Time • Master's Degree • Beverly Hills
Python
AI/ML
Fine-tuning
Reinforcement Learning
Multimodal AI
Computer Vision
NLP
Transformers
PyTorch
RAG
Synthetic Data
Hugging Face
Post-training
Machine Learning
DevOps
Git
Docker
Linux
Apply
≈ $124k – $257k per year (Estimated) • In office • Full-Time • 8+ years exp • Beverly Hills
AI/ML
LLM
LLM Guardrails
Agentic Workflows
DevOps
Terraform
GCP
Pulumi
Azure
AWS
Cloudflare
IAM
DNS
Cybersecurity
ISO 27001
SOC 2
NIST 800-53
FedRAMP
Zero Trust
Management
Google Workspace
Apply
$85k – $115k per year • In office • Full-Time • 3+ years exp • Beverly Hills
Management
Asana
Slack
Marketing
YouTube
Instagram
Apply
$135k – $165k per year • In office • Full-Time • 6+ years exp • Beverly Hills
Apply
$150k – $180k per year • In office • Full-Time • 5+ years exp • Beverly Hills
Design
Figma
Management
Agile
Apply
$56k – $68k per year • In office • Full-Time • 1+ year exp • PhD • Beverly Hills
Apply
$90k – $110k per year • In office • Full-Time • 4+ years exp • PhD • Beverly Hills
Apply
$70k – $80k per year • In office • Full-Time • 2+ years exp • PhD • Beverly Hills
Marketing
YouTube
Instagram
Apply
$140k – $155k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Beverly Hills
Apply
In office • Part-Time • High School Diploma • Beverly Hills
Apply
See all jobs
This is one of many
1,099,169 more open roles from verified company boards, updated every day.