368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$235k – $353k per year
Location
Remote/Hybrid (Redwood City, United States)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Luma AI is an artificial intelligence company that specializes in building advanced 3D visual technology and generative AI tools. Founded in 2021, the company is best known for its realistic text-to-3D models, NeRF (Neural Radiance Fields) capture app, and its high-quality video generation model, Dream Machine. By enabling creators and developers to easily capture, generate, and edit photorealistic 3D assets and video, it aims to make next-generation digital media creation accessible to everyone.

You'll own the reliability of Luma's 10k+ GPU fleet: the scheduling, efficiency, and resilience that research and products depend on. As a Staff AI Infrastructure Engineer, you'll be a technical authority who turns deep systems knowledge into repeatable, company-wide reliability, and a leader other strong engineers want to work with.

This is close-to-the-metal work - kernels, containers, schedulers, networking, storage, GPU behavior - under demand hard enough that yesterday's solutions break regularly. It's also a technical-leadership role: you'll set the bar and grow the team. If most of your experience has been inside highly abstracted internal platforms where others owned the underlying machinery, this likely isn't a match.

What You'll Own

  • Architect and operate large, heterogeneous GPU environments under extreme demand, improving utilization and performance where small gains change company outcomes.

  • Resolve failures spanning hardware, OS, runtimes, and orchestration, and eliminate whole classes of instability.

  • Define how infrastructure and workloads evolve as cluster size and concurrency grow - scheduling, placement, resource management.

  • Work directly with research to build the systems new model capabilities require, and scale inference without sacrificing reliability or latency.

  • Hire and develop exceptional systems and reliability engineers, and set the bar for depth, judgment, and production ownership.

  • Shape product and research architecture early through strong partnerships.

First 90 Days

One way the first 90 could unfold.

  • Days 1-30 - Immerse & Diagnose: Learn the fleet, its failure modes, and the biggest reliability and utilization gaps.

  • Days 30-60 - Ship & Validate: Eliminate a recurring class of instability or land a utilization or performance win that moves company outcomes.

  • Days 60-90 - Scale & Systemize: Set the reliability direction, redesign ahead of where today's abstractions will fail, and begin building the team.

What You Bring

  • Deep expertise in Linux and distributed systems.

  • Experience operating GPU or accelerator clusters in real production environments.

  • Strong fluency in Kubernetes and modern open-source infrastructure.

  • Comfort debugging across hardware, kernel, runtime, and orchestration, and understanding how systems behave under contention and at scale.

  • You write code and build automation, and think in bottlenecks, failure modes, and trade-offs.

  • Judgment engineers trust, especially when things break.

Nice to Have

  • You raise reliability standards company-wide and influence product and research architecture early.

  • You build partnerships rather than ticket queues, and attract and level up strong engineers.

  • Curiosity for how models use infrastructure, because improving systems expands what becomes possible.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence - the next step beyond language models comes from vision. Luma is an equal opportunity employer.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Redwood City
Director, Engineering 11 hours ago
$42k – $91k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Bengaluru
Kotlin
Node JS
Swift
TypeScript
JavaScript
Node JS
Nest.JS
Databases
Apache Kafka
Frontend
Angular
React.js
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub Actions
Kubernetes
GitHub
QA
Appium
BrowserStack
Playwright
Selenium
Apply
$34k – $83k per year (Estimated) • Remote/Hybrid • Full-Time • Bengaluru
Node JS
JavaScript
Databases
Apache Kafka
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Datadog
FinOps
GitHub Actions
Grafana
Istio
Kubernetes
Platform Engineering
Prometheus
SLI/SLO/SLA
Terraform
GitHub
IAM
Cybersecurity
GDPR
Microsoft Defender
Microsoft Defender for Cloud
Okta
PCI DSS
Management
ServiceNow
Apply
$23k – $65k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • India
JavaScript
SQL
TypeScript
Java
Java
Spring Boot
Spring Cloud
Spring Framework
Frontend
Angular
GraphQL
DevOps
CI/CD
CircleCI
Docker
Jenkins
Kubernetes
OpenShift
Cybersecurity
SonarQube
Analytics
Tableau
Apply
Java Developer 11 hours ago
$19k – $51k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • India
Java
Kotlin
Java
Spring Boot
Spring Cloud
Kotlin
Mockito
Databases
Apache Kafka
Cassandra
PostgreSQL
RabbitMQ
Mobile
JUnit
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Cybersecurity
SonarQube
Apply
$32k – $64k per year (Estimated) • Remote • Contractor • Saint Petersburg
Python
DevOps
CI/CD
GitLab CI
Kubernetes
SLI/SLO/SLA
GitLab
Apply
$93k – $208k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Berlin
Python
TypeScript
Apply
$170k – $360k per year • Remote/Hybrid • Full-Time • Redwood City
Python
AI/ML
Multimodal AI
Ray
Spark
Apply
$79k – $188k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • London
Python
TypeScript
Apply
$43k – $152k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Paris
Python
TypeScript
Apply
$225k – $300k per year • Remote/Hybrid • Full-Time • 7+ years exp • Redwood City
AI/ML
Claude
Claude Code
Cursor
Design
Figma
Apply
$91k – $182k per year (Estimated) • In office • 3+ years exp • Redwood City
Python
DevOps
AWS
CI/CD
GCP
Terraform
Apply
$96k – $120k per year • In office • Internship • Bachelor's Degree • Redwood City
JavaScript
AI/ML
AI Agents
Apply
$96k – $120k per year • In office • Internship • Master's Degree • Redwood City
JavaScript
Python
AI/ML
AI Agents
Time Series Forecasting
DevOps
GitHub
Apply
$85k – $170k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Redwood City
Java
Scala
Databases
MySQL
Apply
$83k – $176k per year (Estimated) • In office • 12+ years exp • Bachelor's Degree • Redwood City
AI/ML
OpenAI Codex
Management
Confluence
Jira
Smartsheet
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.