386,863open jobs
10,118companies
50,530added this week
Browse all
Salary
$132k – $288k per year (Estimated)
Location
In office
Employment
Full-Time
Overview
Company
Impact
Profile match
ElevenLabs is a voice artificial intelligence company founded in 2022 by two friends from Poland who were frustrated by how badly foreign films were dubbed into their own language. It builds speech synthesis, voice cloning, dubbing and conversational voice agents, and its models are widely regarded as the most natural sounding in the field, which has made it the default choice for audiobooks, video localisation and voice interfaces. Headquartered between London and New York and valued in the billions, it has also had to build unusually strict consent and detection controls because voice cloning is so easily abused.

About ElevenLabs

ElevenLabs is an AI research and product company transforming how we interact with technology.

We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always.

We have expanded from voice into three main platforms:

  • ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.

  • ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.

  • ElevenAPI gives developers access to our leading AI audio foundational models.

Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you.

How we work

  • High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.

  • Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.

  • AI first: We use AI to move faster with higher-quality results. We do this across the whole company-from engineering to growth to operations.

  • Excellence everywhere: Everything we do should match the quality of our AI models.

  • Global team: We prioritize your talent, not your location.

What we offer

  • Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.

  • Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.

  • Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.

  • Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.

  • Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.

  • Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.

About ElevenLabs

ElevenLabs is an AI research and product company transforming how we interact with technology.

We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always.

We have expanded from voice into three main platforms:

- ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.

- ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.

- ElevenAPI gives developers access to our leading AI audio foundational models.

Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you.

How we work

  • High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.

  • Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.

  • AI first: We use AI to move faster with higher-quality results. We do this across the whole company-from engineering to growth to operations.

  • Excellence everywhere: Everything we do should match the quality of our AI models.

  • Global team: We prioritize your talent, not your location.

What we offer

  • Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.

  • Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.

  • Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.

  • Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.

  • Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.

  • Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.

About the role

Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense. When the clusters just work, research moves faster. Your impact is measured directly in training throughput and researcher velocity.

This is a builder-operator role with real breadth: one week you're writing automation that eliminates a whole class of manual work, the next you're benchmarking a new provider's InfiniBand fabric or on-site bringing new hardware online. You'll have unusual scope and autonomy - we're a lean team where decisions are made by the people closest to the problem.

What you’ll be doing

  • Operate and improve our GPU fleet end to end: provisioning, scheduling, monitoring, upgrades, capacity planning

  • Build automation that keeps the fleet healthy without human intervention - node health checks, automated draining and remediation, burn-in pipelines for new capacity

  • Own the stack beneath the training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, high-speed networking (InfiniBand/RoCE)

  • Run and tune job scheduling (Slurm or similar) so researchers get compute fairly and fast

  • Build and maintain high-performance storage for datasets and checkpoints

  • Hunt down performance problems: stragglers, degraded links, thermal issues, flaky GPUs - and fix the class of problem, not just the instance

  • Evaluate rented GPU capacity: benchmark it, validate it, hold providers to their SLAs

  • Hands-on hardware work when it's needed: racking, cabling, diagnostics, coordinating with datacenter staff and vendors

  • Keep clusters secure by default: access control, network isolation, secrets

Requirements

  • Have run large-scale Linux server or GPU environments in production and enjoy both building and operating

  • Know the NVIDIA stack well - drivers, CUDA, NCCL, DCGM - or have deep systems experience and learn hardware stacks fast

  • Are comfortable with bare-metal environments, server hardware, and high-speed networking

  • Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform

  • Are happy digging into noisy data (metrics, logs, PromQL) to find what's actually wrong

  • Like owning real scope end to end and being the person others rely on

  • Don't consider any task above or beneath you - datacenter trips included

Nice to have

  • Experience supporting ML training workloads from the infra side (distributed training failure modes, checkpointing patterns)

  • Experience evaluating and working with GPU cloud providers

  • Parallel filesystems (WEKA, VAST, etc) or large-scale object storage

  • BMC/IPMI/Redfish automation, PXE provisioning at scale

  • Power and cooling awareness for dense GPU deployments

How we work

Small team, high trust, minimal bureaucracy. We automate aggressively so on-call is sane, and we fix root causes so the same page never fires twice. You'll work directly with the researchers whose jobs run on your clusters - short feedback loops, no ticket queues between you and impact.

Location

This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.

# HPC Infrastructure Engineer - GPU Clusters

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,863 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$191k – $397k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv • Yokneam
AI/ML
CUDA
CUDA Toolkit
NCCL
DevOps
Kubernetes
Apply
$95k – $200k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Palo Alto
Go
Python
AI/ML
CUDA
CUDA Toolkit
Fine-tuning
Mistral SDK
NCCL
PyTorch
DevOps
Karpenter
Kubernetes
Platform Engineering
Cybersecurity
Kyverno
Apply
$99k – $132k per year • In office • Full-Time • Heilbronn
Bash
PowerShell
Python
Apply
$85k – $183k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Fort Belvoir
DevOps
Ansible
Cybersecurity
Palo Alto NGFW
Apply
$40k – $73k per year • In office • Full-Time • 1+ year exp • High School Diploma • United States
DevOps
Ansible
AWS
Apply
$88k – $172k per year (Estimated) • Remote/Hybrid • Full-Time • London
Python
SQL
AI/ML
ElevenLabs
Human-in-the-Loop
Apply
$105k – $220k per year (Estimated) • Remote • Full-Time
AI/ML
CUDA
CUDA Toolkit
ElevenLabs
Knowledge Distillation
Quantization
SGLang
TensorRT
Triton
vLLM
DevOps
GitHub
Apply
$54k – $97k per year (Estimated) • Remote • Full-Time
Python
AI/ML
ElevenLabs
Apply
Remote • Full-Time
Python
AI/ML
ElevenLabs
Apply
$106k – $168k per year (Estimated) • In office • Full-Time
Python
AI/ML
ElevenLabs
Apply
See all jobs
This is one of many
386,863 more open roles from verified company boards, updated every day.