534,639open jobs
19,018companies
74,061added this week
Browse all
Salary
$312k – $389k per year
Location
Remote/Hybrid (Sunnyvale, United States)
Seniority
Staff
Overview
Company
Impact
Profile match
Wayve is a British autonomous driving company founded in Cambridge in 2017 that trains end-to-end neural networks to drive rather than assembling hand-written rules around high-definition maps. Its AI Driver learns from camera video and can be deployed on standard vehicle sensor sets, which is what lets the company demonstrate driving in cities it has never mapped. Headquartered in London and backed by a one billion dollar SoftBank-led round plus investment from NVIDIA, Microsoft and Uber, it partners with carmakers to ship assisted driving software and is developing fully driverless deployments.

The role

As a Senior / Staff Machine Learning Engineer in Wayve's AV Core organisation, you will advance reinforcement learning methods for end-to-end driving models. You will identify where learning from reward or feedback can improve beyond behavior cloning, then take promising ideas from design through large-scale experiments, rigorous evaluation, and integration into our best driving models.

Driving Core team develops the learning methods that turn diverse driving data into robust closed-loop behavior. You will be a technical owner for reinforcement learning within the group, working closely with researchers and engineers across AV Core, Simulation, Evaluation, and Product Engineering. Success means producing measurable improvements in driving behavior.

Core Model Safety team develops the core model competencies that enable safe, driverless operation. You will lead the technical direction and delivery of a learned emergency trajectory model for low-frequency, high-consequence maneuvers such as evasive steering and emergency braking. You will take the programme from problem definition through modelling, evaluation, integration, and evidence for deployment.

Key responsibilities

  • Shape and execute the reinforcement learning roadmap for Driving Core / Core Model Safety, selecting problems and methods against clear behavioral gaps and measurable success criteria.
  • Develop and evaluate post-behavior-cloning optimization methods, including offline and off-policy reinforcement learning as well as other reward-guided approaches; design the regularization, data strategy, and diagnostics needed to make policies reliably better.
  • Help improve the reward models and related learning signals used to train and evaluate driving policies, working with partner teams to strengthen their quality, scalability, and downstream usefulness.
  • Build robust training and experimentation workflows using large-scale driving data; diagnose distribution shift, objective misspecification, optimization instability, and data or evaluation bias.
  • Define evidence across offline metrics, open-loop tests, closed-loop simulation, and on-road evaluation, and distinguish genuine policy improvement from benchmark overfitting.
  • Productionize successful methods in the shared ML stack, communicate decisions and results clearly, and raise the technical bar through design reviews, code reviews, and mentoring.

About you

In order to set you up for success as a Staff / Senior Machine Learning Engineer at Wayve, we’re looking for the following skills and experience.

Essential

  • A strong track record developing and experimentally validating reinforcement learning or closely related sequential decision-making methods on complex, high-dimensional problems.
  • Deep understanding of modern reinforcement learning fundamentals, including policy and value learning, off-policy learning, function approximation, distribution shift, and the failure modes of learned objectives.
  • Hands-on experience with behaviour cloning, reinforcement learning, or related methods.
  • Proficiency in Python and PyTorch, with strong software engineering practices and hands-on experience building reliable machine learning training and evaluation systems.
  • Excellent experimental judgement: able to turn an ambiguous behavioral problem into falsifiable hypotheses, useful metrics, disciplined ablations, and clear technical decisions.
  • Senior-level ownership and collaboration: able to lead a substantial technical area, work across research and engineering boundaries, and bring others along through clear written and verbal communication.

Desirable

  • Experience with offline reinforcement learning, imitation learning, reward modeling, preference learning, or post-training of large neural policies.
  • Experience in autonomous vehicles, robotics, control, or another domain where policies interact with safety-critical physical systems, including an understanding of motion planning, vehicle dynamics, control, or collision avoidance.
  • Experience with closed-loop simulation, off-policy evaluation, uncertainty or calibration, and evaluation under rare or shifted conditions.
  • Experience training multimodal, transformer-based, or generative policy models at scale.
  • Proficiency in C++, CUDA, distributed training, or performance optimization for production machine learning systems.

This is a full-time role based in our office in Sunnyvale. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. The reasonably estimated salary for this role ranges from $311,850 to $389,400, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
534,639 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sunnyvale
$86k – $185k per year • Remote/Hybrid • 2+ years exp • Bachelor's Degree • Toronto
Python
SQL
AI/ML
Reinforcement Learning
TensorFlow
PyTorch
DevOps
GCP
Azure
AWS
Apply
$110k – $148k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Orlando
Python
SQL
Databases
Snowflake
Analytics
Tableau
Power BI
Microsoft Excel
Management
SharePoint
Apply
Senior Data Scientist 9 hours ago
$142k – $190k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Glendale
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
AI/ML
Spark
DevOps
AWS
Analytics
Tableau
A/B Testing
Looker
Management
Agile
Scrum
Apply
Remote/Hybrid • Bachelor's Degree
Python
SQL
Python
pySpark
Databases
Databricks
AI/ML
Spark
dbt
DevOps
Terraform
Azure
CI/CD
GitHub
Analytics
Power BI
ETL/ELT
Apply
Data Scientist 9 hours ago
$69k – $128k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • United States
Python
SQL
Apply
$210k – $311k per year • Equity • Remote/Hybrid • Sunnyvale
Robotics
Teleoperation
Apply
$77k – $162k per year (Estimated) • Remote/Hybrid • Tokyo
Apply
$83k – $244k per year (Estimated) • In office • 3+ years exp • London
AI/ML
Model Context Protocol
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
Kubernetes
Apply
Remote/Hybrid • London
Cybersecurity
Okta
ISO 27001
SOC 2
Zero Trust
Apply
$119k – $293k per year (Estimated) • Remote/Hybrid • PhD • London
Python
AI/ML
Agentic Workflows
DevOps
Kubernetes
Apply
$146k – $267k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Sunnyvale
Kotlin
Kotlin
Kotlin Coroutines
Databases
Apache Kafka
Mobile
Jetpack Compose
MVVM
Room
Clean Architecture
ViewModel
DevOps
Azure
CI/CD
Apply
$129k – $273k per year (Estimated) • In office • Full-Time • 8+ years exp • Sunnyvale
Design
SolidWorks
Apply
$153k – $366k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Master's Degree • Sunnyvale • Toronto
AI/ML
AI Agents
Cerebras
OpenAI
Edge AI
DevOps
HPC
Cybersecurity
Wireshark
Apply
Category Manager 10 hours ago
$103k – $210k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Master's Degree • Sunnyvale
Analytics
Microsoft Excel
Management
Agile
Kanban
Apply
$104k – $155k per year • In office • Full-Time • 9+ years exp • Bachelor's Degree • Sunnyvale
Design
AutoCAD
Apply
See all jobs
This is one of many
534,639 more open roles from verified company boards, updated every day.