368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$118k – $282k per year (Estimated)
Location
In office (Los Angeles)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
HeyGen is a company founded in 2020 that generates presenter videos from text using synthetic avatars and cloned voices. Its translation feature re-voices existing footage in other languages while matching lip movement, which made it popular with marketing and training teams. The company reports fast revenue growth serving businesses that produce large volumes of localised video.

About HeyGen

At HeyGen, our mission is to make visual storytelling accessible to all. Over the last decade, visual content has become the preferred method of information creation, consumption, and retention. But the ability to create such content, in particular videos, continues to be costly and challenging to scale. Our ambition is to build technology that equips more people with the power to reach, captivate, and inspire audiences.

Learn more at www.heygen.com.  Visit our Mission and Culture doc here

We are seeking a seasonedTechnical Leader to build and scale the foundational compute infrastructure that powers our state-of-the-art AI models-from multimodal training data pipelines to high-throughput, low-latency video generation.

Responsibilities

You will be the core engineer responsible for building the robust, efficient, and scalable platform that enables our research and production teams to rapidly iterate on HeyGen's generative video models. Your contributions will directly impact model performance, developer productivity, and the final quality of every AI-generated video.

  • Optimize GPU Utilization: Design and implement mechanisms to aggressively optimize GPU and cluster utilization across thousands of devices for inference, training, data processing and large-scale deployment of our state-of-art video generation models.

  • Develop Large-Scale AI Job Framework: Build highly scalable, reliable frameworks for launching and managing massive, heterogeneous compute jobs, including multi-modal high-volume data ingestion/processing, distributed model training, and continuous evaluation/benchmarking.

  • Enhance Observability: Develop world-class observability, tracing, and visualization tools for our compute cluster to ensure reliability, diagnose performance bottlenecks (e.g., memory, bandwidth, communication).

  • Accelerate Pipelines: Collaborate closely with AI researchers and AI engineers to integrate innovative acceleration techniques (e.g., custom CUDA kernels, distributed training libraries) into production-ready, scalable training and inference pipelines.

  • Infrastructure Management: Champion the adoption and optimization of modern cloud and container technologies (Kubernetes, Ray) for elastic, cost-efficient scaling of our distributed systems.

Minimum Requirements

We are looking for a highly motivated engineer with deep experience operating and optimizing AI infrastructure at scale.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

  • 5+ years of full-time industry experience in large-scale MLOps, AI infrastructure, or HPC systems.

  • Experience with data frameworks and standards like Ray, Apache Spark, LanceDB

  • Strong proficiency in Python and a high-performance language such as C++  for developing core infrastructure components.

  • Deep understanding and hands-on experience with modern orchestration and distributed computing frameworks such as Kubernetes and Ray.

  • Experience with core ML frameworks such as PyTorch, TensorFlow, or JAX.

Preferred Qualifications

  • Master's or PhD in Computer Science or a related technical field.

  • Demonstrated Tech Lead experience, driving projects from conceptual design through to production deployment across cross-functional teams.

  • Prior experience building infrastructure specifically for Generative AI models (e.g., diffusion models, GANs, or large language models) where cost and latency are critical.

  • Proven background in building and operating large-scale data infrastructure (e.g., Ray, Apache Spark) to manage petabytes of multi-modal data (video, audio, text).

  • Expertise in GPU acceleration and deep familiarity with low-level compute programming, including CUDA, NCCL, or similar technologies for efficient inter-GPU communication.

What HeyGen Offers

  • Competitive salary and benefits package.
  • Dynamic and inclusive work environment.
  • Opportunities for professional growth and advancement.
  • Collaborative culture that values innovation and creativity.
  • Access to the latest technologies and tools.

HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Los Angeles
Data Scientist 1 day ago
$25k – $56k per year (Estimated) • Remote • Bachelor's Degree • Moscow
C++
Python
SQL
AI/ML
Computer Vision
CUDA
CUDA Toolkit
TensorRT
DevOps
Docker
Git
Kubernetes
Apply
$20k – $48k per year (Estimated) • In office • Moscow
C#
C++
Java
Python
SQL
DevOps
CI/CD
Management
Draw.io
Jira
Apply
$21k – $55k per year (Estimated) • In office • Full-Time • Bengaluru
C++
Go
Java
C++
Protobuf
Databases
Apache Ignite
ElasticSearch
PostgreSQL
RabbitMQ
Redis
DevOps
CI/CD
Docker
gRPC
Kibana
OpenTelemetry
Apply
$65k – $156k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tel Aviv
C++
DevOps
Platform Engineering
RTOS
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
C++
Java
Python
SQL
C#
TypeScript
JavaScript
Java
Maven
C#
.NET
Databases
Apache Kafka
MySQL
AI/ML
ChatGPT
Copilot
Frontend
Angular
DevOps
CI/CD
Docker
Jenkins
Kubernetes
Prometheus
Apply
$180k – $240k per year • In office • San Francisco
Go
Python
C++
Java
C++
CMake
Java
Gradle
DevOps
Bazel
Buildkite
CI/CD
CircleCI
GitHub Actions
GitLab CI
Jenkins
Kubernetes
GitHub
GitLab
Apply
Security Engineer 2 months ago
$131k – $289k per year (Estimated) • In office • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
DevOps
AWS
IAM
Cybersecurity
ISO 27001
SOC 2
Apply
$200k per year • In office • 7+ years exp • Los Angeles
AI/ML
AI Agents
Claude
Claude Code
CrewAI
Cursor
Google ADK
LangGraph
LangChain
Model Context Protocol
OpenAI
OpenAI Agents SDK
DevOps
CI/CD
GitHub
Apply
$83k – $235k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Los Angeles
C++
Python
C++
PyTorch C++
TensorFlow C++
Databases
LanceDB
AI/ML
Accelerate
CUDA
CUDA Toolkit
Diffusion Models
JAX
Multimodal AI
PyTorch
Ray
Spark
TensorFlow
NCCL
Scale AI
DevOps
Kubernetes
HPC
Apply
$114k – $256k per year (Estimated) • In office • Bachelor's Degree • Los Angeles
Python
AI/ML
Multimodal AI
PyTorch
TensorFlow
Edge AI
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$94k – $294k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Portland • Milwaukee • Dallas • Columbus
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$150k – $185k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Los Angeles
AI/ML
Human-in-the-Loop
Apply
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.