706,097open jobs
41,961companies
98,408added this week
Browse all
Salary
$270k – $330k per year
Location
In office (New York)
Seniority
Staff · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luig i), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll own the full lifecycle of a machine, from accepting and benchmarking new hardware from a growing set of providers, to network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation of unhealthy hosts. You'll manage a team of 3-8 engineers while staying hands-on across the stack which involves BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services, and you'll shape our long-term path further down the stack.

Requirements:

  • 7+ years of experience writing high-quality production code

  • 3+ years of direct people management experience, ideally leading a team of engineers through project planning, growth, and performance conversations

  • Experience operating large fleets of physical hardware at scale (bare metal provisioning, BMC/IPMI, PXE and network boot, firmware) or building the control planes that manage them (the more challenges you've worked through, the better)

  • Strong cloud skills

  • Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers, etc.)

  • Experience working with hardware and colocation providers, including hardware acceptance testing and benchmarking

  • Track record of setting technical direction and driving architectural decisions across a team

  • Willingness to step into the thick of it with our on-call rotation and respond to production incidents

Nice-to-Haves:

  • Experience with GPUs and the NVIDIA software stack (drivers, health monitoring, RDMA/NVLink) in production

  • Prior experience with Go

Key Things the Team Is Working On:

  • Automatic remediation of unhealthy machines (power cycling, reimaging, GPU recovery) to maximize uptime and minimize operator toil.

  • Automatic integration of new CPU, GPU, and storage servers into the fleet while managing hardware and network heterogeneity.

  • Network health monitoring and reliability across many datacenters, and standardization of bare metal network configuration.

  • Automatic hardware acceptance testing and benchmarking (CPU, disk, GPU, interconnect, network).

  • Custom network bootloader, machine image pipeline, and kernel and firmware management across the fleet.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
706,097 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$128k – $214k per year • In office • TS/SCI • 20+ years exp • Bachelor's Degree
DevOps
Linux
Windows
Apply
$135k – $189k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
Python
Verilog
C++
SystemVerilog
VHDL
MATLAB
DevOps
Git
Linux
TCP/IP
Cybersecurity
Wireshark
Chips/EDA
JTAG Debuggers
Management
Agile
Apply
$87k – $182k per year • In office • Top Secret • 5+ years exp • Bachelor's Degree • Reston
Databases
Oracle
ActiveMQ
DevOps
OpenShift
Kubernetes
GitLab
Linux
Windows
Cybersecurity
NIST 800-53
Apply
$53k – $91k per year • In office • Full-Time • Bachelor's Degree • Washington
Python
C++
Bash
MATLAB
DevOps
Linux
Analytics
Matplotlib
Apply
$80k – $128k per year • In office • Top Secret • 2+ years exp • Bachelor's Degree
Python
Java
C++
Java
Maven
Gradle
DevOps
CI/CD
Jenkins
Git
Docker
Kubernetes
GitLab
Linux
Windows
Management
Confluence
Jira
Agile
Apply
$270k – $330k per year • In office • Full-Time • 7+ years exp • San Francisco • New York
Rust
DevOps
Amazon S3
Linux
Analytics
Seaborn
Matplotlib
Apply
$250k – $300k per year • In office • Full-Time • 5+ years exp • San Francisco • New York
Python
AI/ML
NVLink
DevOps
Linux
BGP
Analytics
Seaborn
Matplotlib
Apply
$250k – $300k per year • In office • Full-Time • 5+ years exp • San Francisco • New York
Rust
DevOps
Amazon S3
Linux
Analytics
Seaborn
Matplotlib
Apply
$140k – $276k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
DevOps
CI/CD
Analytics
Seaborn
Matplotlib
Apply
Account Manager 6 days ago
$300k per year • Remote • Full-Time • San Francisco • New York
Analytics
Seaborn
Matplotlib
Apply
$74k – $129k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Boston • New York
Apply
$101k – $115k per year • In office • Full-Time • 1+ year exp • Bachelor's Degree • McLean • New York • Chicago
Management
Agile
Apply
$101k – $115k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • McLean • Charlotte • Richmond • New York • Plano
DevOps
GCP
Azure
AWS
SRE
Chaos Engineering
Management
Agile
Apply
$119k – $136k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • McLean • Charlotte • Richmond • New York • Plano
DevOps
GCP
Azure
AWS
SRE
Chaos Engineering
Management
Agile
Apply
$245k – $279k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco • McLean • Cambridge • San Jose • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
AI Agents
PyTorch
LLM
CUDA
Hugging Face
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
See all jobs
This is one of many
706,097 more open roles from verified company boards, updated every day.