435,295open jobs
15,159companies
67,452added this week
Browse all
Salary
$159k – $285k per year (Estimated)
Location
Remote (United States)
Seniority
Principal
Employment
Full-Time
Overview
Company
Impact
Profile match
Cerence builds conversational assistants for vehicles. Its speech and language software ships in hundreds of millions of cars. The company was spun out of the automotive division of Nuance.

A Moving Experience.

What You Will Work On

  • Design andoperatedistributed training systems for large neural networks(autoregressive, diffusion,State Space Modelsetc.)across GPU clusters

  • Optimisemulti-node, multi-GPU execution tomaximizethroughput andutilization

  • Diagnose&resolve bottlenecks across compute, memory, and network

  • Improve training stability and fault tolerance at scale

  • Partner with research and applied ML teams to productionizelarge-modeltraining pipelines

Core Responsibilities

  • Distributed Training Infrastructure

  • Build andoptimizeGPU cluster orchestration using:

  • Slurm

  • Kubernetes

  • Ray

  • RunAI

  • Ensure efficient scheduling, isolation, and fairness across training workloads

  • Communication & Networking

  • Optimizeand debug distributed communication using:

  • NCCL

  • RDMA

  • InfiniBand

  • NVLink

  • Minimizenetworking bottlenecks that dominateend-to-endtraining time

  • Training Frameworks

  • Scale large-model training using:

  • PyTorchDistributed

  • Megatron-LM

  • DeepSpeed

  • Ownmulti-nodelaunch configurations, failure recovery, and performance tuning

  • Memory & Performance Optimization

  • Apply advanced memory optimization techniques:

  • Activation checkpointing

  • ZeRO(Stage 1-3) and offload strategies

  • Balance compute, memory, and communication to push model size and batch scale

What Success Looks Like

  • GPUutilizationconsistently stays high (>80-90%)

  • Training scales cleanly from single node to dozens or hundreds of GPUs

  • Communication overhead is minimized and predictable

  • Large training jobs run stably for days or weeks without failure

  • New models can be trained faster, larger, and more reliably than before

Required Experience & Skills

  • Strongly Required

  • Deephands-onexperience with distributed systems or ML systems

  • Experience runninglarge-scaleworkloads on GPU clusters

  • Production experience withPyTorchdistributed training

  • Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)

  • Low-levelunderstanding of GPU communication and networking

  • Critical Technical Skills

  • GPU orchestration:Slurm, Kubernetes, Ray,RunAI

  • Communication libraries: NCCL, RDMA, InfiniBand,NVLink

  • Training frameworks:PyTorchDistributed,Megatron-LM,DeepSpeed

  • Memoryoptimisation: activation checkpointing,ZeROoffload techniques

Common ProblemsYou’llBe Solving

  • Many teams fail at scale because:

  • GPUutilizationis low despite large clusters

  • Networking and communication dominate training time

  • Training jobs crash or become unstable at large scale

  • You will be explicitly focused oneliminatingthese failure modes.

Ideal Background

  • This role is a strong fit for individuals who have worked as:

  • ML Systems Engineer

  • Distributed Systems Engineer

  • AI Infrastructure Engineer

  • HPC Engineer transitioning into ML

  • Experience working with large language models or foundation models is a strong plus, but deep systemsexpertiseis valued over pure model architecture experience.

Why This Role Matters

Without robust distributed training infrastructure, progress on largemodelsstalls. This role directly enables:

  • Larger models

  • Faster iteration cycles

  • More reliable research-to-production pipelines

You will be building the foundation that makeslarge-scaleAI possible.

Cerence Inc. (Nasdaq: CRNC and www.cerence.com) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers - from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC - to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages.

As Cerence looks to the future and continues an ambitious growth agenda, we need someone to join the team and help build the future of voice and AI in cars. This is an exciting opportunity to join Cerence’s passionate, dedicated, global team and be a part of meaningful innovation in a rapidly growing industry.

EQUAL OPPORTUNITY EMPLOYER

Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement.

All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes:

- Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace.

- Following security procedures to report any suspicious activity.

- Having respect for corporate security procedures to allow those procedures to be effective.

- Adhering to company's compliance and regulations.

- Encouraging to follow a zero tolerance for workplace violence.

- Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).

- Demonstrative knowledge of information security through internal training programs.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,295 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
$73k – $178k per year (Estimated) • Remote/Hybrid • Full-Time • Belfast
Java
SQL
Databases
MongoDB
Apache Kafka
DevOps
CI/CD
Kubernetes
Amazon S3
Apply
$145k – $218k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Mississauga
Python
Java
AI/ML
LangChain
Claude
NLP
Llama
TensorFlow
PyTorch
RAG
Hugging Face
OCR
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Apply
Software Engineer 1 day ago
In office • 3+ years exp • Bachelor's Degree • Guadalajara
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
AI/ML
Copilot
Cursor
Claude
Claude Code
Frontend
Vue.js
React.js
DevOps
Rest API
Azure
CI/CD
AWS
Docker
Kubernetes
GitHub
Apply
$26k – $73k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Guadalajara
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
AI/ML
Copilot
Cursor
Claude
Claude Code
Frontend
Vue.js
React.js
DevOps
Rest API
Azure
CI/CD
AWS
Docker
Kubernetes
GitHub
Apply
FinOps Engineer 1 day ago
Remote/Hybrid • 5+ years exp
Python
SQL
Databases
Databricks
AI/ML
Copilot
Claude
Claude Code
LLM
LLM Guardrails
DevOps
Terraform
GCP
Azure
AWS
Kubernetes
Grafana
Platform Engineering
Bicep
Azure AKS
FinOps
GitHub
Analytics
Power BI
Apply
$168k – $253k per year • Equity • Remote • Full-Time • 7+ years exp • United States
Apply
$93k – $221k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • PhD • Aachen
Python
AI/ML
Fine-tuning
Speech Recognition
LLM
Text-to-Speech
Apply
$24k – $57k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune
Python
AI/ML
Text-to-Speech
Edge AI
DevOps
Azure
CI/CD
Jenkins
Docker
Kubernetes
GitLab
Apply
$148k – $237k per year • Equity • Remote • Full-Time • 10+ years exp • United States
AI/ML
Physical AI
Apply
$35k – $93k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree
Python
Java
Kotlin
C++
AI/ML
Speech Recognition
DevOps
Git
GitLab
Management
Confluence
Jira
Apply
$154k – $231k per year • Remote/Hybrid • Full-Time • 4+ years exp • United States
Apply
$56k – $83k per year • Remote • Full-Time • 6+ years exp • Bachelor's Degree • United States
Marketing
Salesforce
Apply
$122k – $184k per year • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Los Angeles • Irvine • Anaheim
Apply
$64k – $96k per year • Remote • Full-Time • 21+ year exp • High School Diploma • Overland Park
Apply
$58k – $88k per year • Remote • Full-Time • 21+ year exp • High School Diploma • United States
Apply
See all jobs
This is one of many
435,295 more open roles from verified company boards, updated every day.