1,134,601open jobs
65,337companies
199,586added this week
Browse all
Salary
≈ $196k – $371k per year (Estimated)
Location
In office (San Diego)
Seniority
Staff · 8+ years exp

First seen by Alion on Oct 1, 2026.

Overview
Company
Impact
Profile match
Apple is an American multinational technology company founded in 1976 by Steve Jobs, Steve Wozniak and Ronald Wayne, and headquartered in Cupertino, California. It designs and sells consumer hardware including the iPhone, Mac, iPad, Apple Watch, AirPods and Vision Pro, together with the operating systems and silicon that run them. A growing services division built around the App Store, iCloud, Apple Music, Apple TV+ and Apple Pay now contributes a large share of profit, making Apple one of the most valuable companies in the world.
Would you like to contribute to Machine Learning and Generative AI technologies? Are you passionate about measuring what matters and ensuring AI systems work reliably for everyone? Do you believe that rigorous evaluation - including holding models accountable to fairness standards - is what separates great ML from good ML? We truly believe it is! We are defining what exceptional looks like for machine learning across Wallet, Payments, and Commerce. As a Machine Learning Engineer specializing in Evaluation, you will establish the evaluation criteria, metrics frameworks, and quality standards that determine when models are ready to reach hundreds of millions of users. Your judgment shapes model quality and earns the confidence to ship. You'll work at the intersection of rigorous ML science and high-impact product decisions, collaborating closely with ML Engineering, Product, Privacy, and Legal teams. This unique opportunity puts you at the center of model quality - designing adversarial test strategies, surfacing failure modes before they reach users, and owning the sign-off process that ensures Apple's financial features meet the highest bar for accuracy, robustness, and reliability.

Description

The ideal candidate is a rigorous, curious ML practitioner who believes that how you measure a model is just as important as how you train it. You think critically about what metrics actually capture, know how models break in the real world, and hold quality standards others find uncomfortably high - including on dimensions like fairness. You will own the full evaluation lifecycle for ML models across Wallet features - designing test frameworks, adversarial corpora, and benchmarks that reflect the diversity of Apple's global user base, then making the final quality call before any model ships. Your findings directly shape model development priorities and product decisions at scale.

Minimum Qualifications

M.S. in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related technical field strongly preferred.

Bachelor's degree with 7+ years hands-on experience in ML evaluation, model quality, or applied research will be considered

5+ years of hands-on ML experience, with deep expertise in model evaluation, offline metrics design, and behavioral testing

Strong track record designing evaluation frameworks for production ML systems - not just accuracy/F1, but precision-recall tradeoffs, calibration, fairness, and task-specific quality dimensions

Creative mindset with the ability to translate standard ML evaluation metrics (F1, AUC, etc.) into utility and user trust measures

Experience testing for distribution shift, out-of-distribution generalization, and temporal drift in real-world deployed models

Proven ability to construct adversarial test suites, aggressor scenarios, and edge-case corpora that surface model failure modes before they reach users

Experience with structured and semi-structured document understanding, OCR pipelines, or financial data extraction is a strong plus

Strong programming skills in Python; fluency with evaluation tooling, data pipelines, and experiment tracking (e.g., MLflow, W&B, or equivalent)

Excellent communication skills - ability to translate metric results into product-quality narratives for engineering and executive audiences

Experience owning model quality sign-off in a cross-functional launch process

Preferred Qualifications

PhD in Computer Science, Data Science, Statistics, AI/ML, or a related field.

Experience with Bayesian or causal graph-based approaches to data generation.

Experience with causal approaches to fairness evaluation - counterfactual fairness, causal Shapley values, or structural causal model-based bias auditing.

Experience evaluating models under privacy constraints or on-device inference settings is a plus.

Familiarity with confidence calibration techniques and uncertainty quantification a plus

Background in financial services, fintech, or consumer payment products

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,134,601 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Diego
In office • Shenzhen
AI/ML
Function Calling
A2A
Structured Outputs
Context Engineering
Agentic Workflows
Apply
≈ $122k – $230k per year (Estimated) • In office • Morrisville
Python
PowerShell
Bash
Perl
DevOps
Windows
Apply
$224k – $280k per year • Remote (United States) • Full-Time
Databases
Databricks
AI/ML
AI Agents
LLM
Agentic Workflows
Multi-Agent Systems
Machine Learning
Cybersecurity
SOC 2
HIPAA
FedRAMP
Threat Modeling
Apply
$180k – $210k per year • Hybrid • 8+ years exp • Bachelor's Degree • Chicago
Python
JavaScript
AI/ML
AI Agents
LLM
RAG
Hallucination
LLM Guardrails
DevOps
CI/CD
Windows
Apply
$183k – $205k per year • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Python
Ruby
Python
Pydantic
Ruby
Ruby on Rails
AI/ML
LangChain
Scikit-learn
Machine Learning
DevOps
Azure
AWS
Kubernetes
Apply
$144k – $216k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Dallas
Python
SQL
Databases
Snowflake
Databricks
AI/ML
Model Context Protocol
Prompt Engineering
RAG
A2A
Context Engineering
Multi-Agent Systems
Machine Learning
DevOps
Azure
Git
AWS
Analytics
Power BI
Management
Agile
Apply
$50k – $96k per year • In office • Internship • Master's Degree • Melbourne
Python
SQL
Python
pySpark
AI/ML
Spark
TensorFlow
Pandas
PyTorch
Machine Learning
DevOps
Docker
Apply
≈ $43k – $112k per year (Estimated) • Hybrid • Bachelor's Degree • Campinas
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
AI Agents
DevOps
Azure DevOps
Azure
CI/CD
Analytics
Power BI
Metabase
Azure Data Factory
Apply
$160k – $260k per year • Equity 0.5–2% • In office • Full-Time • San Francisco
AI/ML
Time Series Forecasting
Machine Learning
Apply
≈ $105k – $208k per year (Estimated) • Equity • Hybrid • Full-Time • United States
AI/ML
AI Agents
Machine Learning
Apply
≈ $233k – $440k per year (Estimated) • Equity • In office • 8+ years exp • Master's Degree • Cupertino
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Gemma
Fine-tuning
Embeddings
Quantization
JAX
Multimodal AI
Knowledge Distillation
Function Calling
AI Agents
Llama
Mistral
TensorFlow
PyTorch
LLM
RAG
BERT
Reranking
Semantic Search
Human-in-the-Loop
Semantic Search
Edge AI
Recommender Systems
Tool Use
Model Distillation
Machine Learning
Apply
≈ $228k – $432k per year (Estimated) • Equity • In office • 2+ years exp • Master's Degree • Cupertino
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Hadoop
Spark
TensorFlow
PyTorch
Machine Learning
Apply
≈ $198k – $376k per year (Estimated) • Equity • In office • 4+ years exp • Bachelor's Degree • Culver City
Python
AI/ML
LangGraph
AutoGen
LangChain
AI Agents
Arize Phoenix
LangSmith
TensorFlow
PyTorch
CrewAI
Braintrust
Agentic Workflows
Tool Use
Machine Learning
Apply
≈ $225k – $425k per year (Estimated) • Equity • In office • 2+ years exp • Bachelor's Degree • Cupertino
Python
AI/ML
Prompt Engineering
NLP
RAG
SFT
Machine Learning
Apply
≈ $263k – $542k per year (Estimated) • Equity • In office • 10+ years exp • Bachelor's Degree • Cupertino
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Gemma
Fine-tuning
Embeddings
Quantization
JAX
Multimodal AI
Knowledge Distillation
Function Calling
AI Agents
Llama
Mistral
TensorFlow
PyTorch
LLM
RAG
BERT
Reranking
Semantic Search
Human-in-the-Loop
Semantic Search
Edge AI
Recommender Systems
Tool Use
Model Distillation
Machine Learning
Apply
$182k – $204k per year • In office • Full-Time • 12+ years exp • San Diego
Apply
≈ $91k – $170k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Diego
Analytics
Microsoft Excel
Management
Outlook
Apply
Truck Driver 1 day ago
$52k per year • Equity • In office • High School Diploma • San Diego
Apply
$38k per year • Equity • In office • Contractor • 1+ year exp • High School Diploma • San Diego
Management
Microsoft Office
Apply
$104k – $155k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Diego • Los Angeles • Irvine
Apply
See all jobs
This is one of many
1,134,601 more open roles from verified company boards, updated every day.