1,456,966open jobs
88,711companies
228,521added this week
Browse all
Salary
≈ $172k – $324k per year (Estimated)
Location
In office (Cupertino)
Seniority
Middle · 3+ years exp
Visa
H-1B filings in 12 months: 6,350 · for this role: 239 · green card filings: 37

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Mar 25, 2026. Apple scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Apple is an American multinational technology company founded in 1976 by Steve Jobs, Steve Wozniak and Ronald Wayne, and headquartered in Cupertino, California. It designs and sells consumer hardware including the iPhone, Mac, iPad, Apple Watch, AirPods and Vision Pro, together with the operating systems and silicon that run them. A growing services division built around the App Store, iCloud, Apple Music, Apple TV+ and Apple Pay now contributes a large share of profit, making Apple one of the most valuable companies in the world.

We build frontier foundation models that power intelligent experiences at Apple. Our team works across the full training lifecycle: including pre-training foundation models, and developing mid-training approaches that bridge general capability and task-specific performance. What makes our work distinct is that we're engineering models specifically for Apple silicon and optimized for experiences that are private, personal, and deeply integrated into the OS. We're solving frontier problems in reward modeling to resist reward hacking, handling sparse and delayed rewards in agentic settings, and aligning models reliably across the spectrum from open-ended creative tasks to precise, action-taking workflows. If you're drawn to hard problems where the research and the product are inseparable, this is the team.

Description

This is a hands-on role focused on the models that power Apple products used daily by over a

billion people. You will design evaluation systems where the outcome is not just a score, but an

actionable signal - one that drives model improvement and predicts real user experience.

Working alongside model training and product teams, you will close the loop between evaluation

and improvement.

Our work spans three areas:

  • Frontier capability assessment: benchmarking against the state of the art in reasoning,

code, knowledge, and agentic workflows

  • Product-aligned evaluation: measuring model quality in ways that reflect real user

experience

  • Evaluation-to-training integration: feeding actionable insights back into the model

development cycle

You may focus on one area or work across multiple, depending on your background and

interests.

We build frontier foundation models that power intelligent experiences at Apple. Our team works across the full training lifecycle: including pre-training foundation models, and developing mid-training approaches that bridge general capability and task-specific performance. What makes our work distinct is that we're engineering models specifically for Apple silicon and optimized for experiences that are private, personal, and deeply integrated into the OS. We're solving frontier problems in reward modeling to resist reward hacking, handling sparse and delayed rewards in agentic settings, and aligning models reliably across the spectrum from open-ended creative tasks to precise, action-taking workflows. If you're drawn to hard problems where the research and the product are inseparable, this is the team.

Minimum Qualifications

3+ years of experience in AI model evaluation, NLP, or a related area (e.g., natural language generation, information retrieval, or conversational AI)

Strong fundamentals in machine learning, natural language processing, and statistical analysis

Proficiency in Python and experience with ML frameworks (PyTorch, JAX, or equivalent)

Demonstrated ability to translate research insights into practical implementations

Strong experimental design skills: ability to design rigorous comparisons and draw valid conclusions from results

Clear technical communication: ability to distill evaluation results into actionable recommendations for cross-functional partners

MS or PhD in Computer Science, Machine Learning, Natural Language Processing or a related technical field. Equivalent practical experience will be considered.

Preferred Qualifications

PhD in Computer Science, Machine Learning, NLP, or a related field

Direct experience evaluating large language models, e.g. benchmark design, model-based judging

Track record of collaborating with model training and data teams to turn evaluation findings into training improvements

Experience building reusable evaluation tooling or analysis frameworks adopted across teams

Familiarity with human evaluation methodology and experience partnering with annotation teams or vendors to assess model quality

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,456,966 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Cupertino
$95k – $135k per year • Hybrid • 3+ years exp • Bachelor's Degree • Atlanta
Python
SQL
PowerShell
Databases
MS SQL
Azure SQL Database
DevOps
Azure
Windows Server
Analytics
ETL/ELT
Apply
Sr Data Scientist 9 days ago
≈ $108k – $206k per year (Estimated) • In office • 4+ years exp • Master's Degree • Dallas
Python
Java
SQL
C++
Python
pySpark
C++
TensorFlow C++
AI/ML
Spark
Scikit-learn
TensorFlow
Pandas
NumPy
Keras
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$94k – $179k per year • Remote (United States) • 3+ years exp • Bachelor's Degree • United States
Python
SQL
Python
Flask
Django
Databases
Snowflake
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
dbt
Prompt Engineering
DevOps
Rest API
Apply
Data Scientist 1 month ago
≈ $114k – $205k per year (Estimated) • Remote (United States) • 5+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
AI/ML
Interpretability
DevOps
Azure
AWS
Apply
$101k – $204k per year • Remote (United States) • 2+ years exp • Bachelor's Degree • United States
Python
SQL
SAS
AI/ML
Metaflow
Machine Learning
Analytics
Microsoft Excel
Apply
In office • Internship • Bachelor's Degree • Longmont
Python
AI/ML
Copilot
Claude
ChatGPT
Prompt Engineering
AI Agents
Gemini
LLM
RAG
Agentic Workflows
Machine Learning
DevOps
CI/CD
Git
Apply
$70k – $85k per year • Remote (United States) • 2+ years exp • Bachelor's Degree • Boston
Python
R
R
Shiny
Analytics
Power BI
Management
Microsoft Office
Apply
≈ $92k – $174k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Chicago
Python
SQL
Databases
Snowflake
Databricks
Google BigQuery
BigQuery
AI/ML
Time Series Forecasting
Synthetic Data
DevOps
GCP
Analytics
ETL/ELT
QA
TestNG
Pytest
Apply
≈ $74k – $174k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Richardson
Python
Java
C++
Perl
LISP
Visual Basic
AI/ML
Machine Learning
DevOps
Linux
Unix
Apply
$203k – $305k per year • Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Northbrook
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
LLM
Feature Store
LLM Guardrails
Machine Learning
DevOps
Azure
Platform Engineering
Incident Management
Analytics
Power BI
Azure Data Factory
Apply
≈ $163k – $307k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Cupertino
Python
SQL
AI/ML
Spark
Time Series Forecasting
Machine Learning
DevOps
Git
Analytics
ETL/ELT
Apply
≈ $162k – $305k per year (Estimated) • In office • 3+ years exp • Master's Degree • Seattle
Python
SQL
Python
FastAPI
AI/ML
LLM
Streamlit
Time Series Forecasting
Human-in-the-Loop
Tool Use
Management
Agile
Apply
≈ $201k – $397k per year (Estimated) • In office • 5+ years exp • Master's Degree • Cupertino
Python
SQL
AI/ML
LLM
Hallucination
Synthetic Data
Analytics
A/B Testing
Apply
≈ $183k – $362k per year (Estimated) • In office • 5+ years exp • Master's Degree • Austin
Python
SQL
AI/ML
LLM
Hallucination
Synthetic Data
Analytics
A/B Testing
Apply
≈ $192k – $367k per year (Estimated) • In office • Bachelor's Degree • New York
Java
SQL
Scala
Java
Maven
Gradle
Databases
Cassandra
Apache Kafka
AI/ML
Hadoop
Flink
DevOps
Git
AWS
Docker
Kubernetes
AWS Lambda
Amazon S3
Apply
≈ $217k – $420k per year (Estimated) • In office • 3+ years exp • Master's Degree • Cupertino
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Reinforcement Learning
AI Agents
TensorFlow
PyTorch
Machine Learning
Robotics
Reinforcement Learning
Apply
≈ $227k – $397k per year (Estimated) • In office • 10+ years exp • Cupertino
Design
Sketch
Adobe After Effects
Apply
≈ $217k – $420k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Cupertino
AI/ML
Diffusion Models
Computer Vision
NLP
Transformers
Machine Learning
Apply
≈ $156k – $293k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Cupertino
Management
Agile
Apply
≈ $178k – $347k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Cupertino
Python
Swift
Swift
Swift Concurrency
AI/ML
Edge AI
Mobile
SwiftUI
Core ML
Swift Testing
Apply
See all jobs
This is one of many
1,456,966 more open roles from verified company boards, updated every day.