706,063open jobs
41,958companies
98,117added this week
Browse all
Salary
$270k – $330k per year
Location
In office (San Francisco, New York)
Seniority
Staff · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luig i), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll set technical direction for the primitives that other teams (filesystems, training, sandboxes) build on, balancing durability, latency, throughput, and cost. You'll own the roadmap from today's hardest problems (garbage collection at petabyte scale, active-active replication, rate limiting that protects the upstream without wasting utilization) to the architectural bets that decide what blobnet becomes: storage colocated with the GPUs, tiered writes, and capacity planning against provider limits. You'll manage a team of 3-8 engineers while staying hands-on across the stack, from local disk and page cache to distributed blob storage and garbage collection, and you'll guide the observability, automation, and on-call practices that keep the system healthy as it grows by orders of magnitude.

Requirements:

  • 7+ years of experience writing high-quality production code

  • 3+ years of direct people management experience, ideally leading a team of engineers through project planning, growth, and performance conversations

  • Experience building high-performance distributed storage or caching systems at a large scale (the more challenges you've worked through, the better)

  • Strong cloud skills, including deep familiarity with object storage (S3 or similar), CDNs, and their consistency, throughput, and cost characteristics

  • Strong knowledge of low-level operating system foundations (Linux kernel, file systems, page cache, containers, etc.)

  • Experience with replication, content addressing, and consistency models in multi-region or multi-cloud systems

  • Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration), including owning cost and capacity planning

  • Track record of setting technical direction and driving architectural decisions across a team, and of building the primitives other teams depend on

  • Willingness to step into the thick of it with our on-call rotation and respond to production incidents

Nice-to-Haves:

  • Experience with data engineering at petabyte-scale.

  • Prior experience with Rust

Key Things the Team Is Working On:

  • P2P sharing of data across workers within a single datacenter to dramatically reduce ingress

  • Replicating data across multiple blob storage providers

  • Automating garbage collection across hundreds of petabytes of data

  • Deploying colocated storage clusters to datacenters to accelerate high-throughput customer workloads

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
706,063 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$142k – $224k per year • In office • Full-Time • 8+ years exp • High School Diploma • Omaha • Des Moines
DevOps
Amazon S3
Apply
$178k – $331k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Boston
SQL
Databases
Databricks
Amazon Redshift
AI/ML
Streamlit
Amazon SageMaker
DevOps
Git
AWS
GitHub
Amazon S3
Analytics
Tableau
Power BI
Apply
$148k – $236k per year • In office • Full-Time • 8+ years exp • Master's Degree • Santa Clara
Python
C++
AI/ML
CUDA Toolkit
TensorRT
CUDA
Edge AI
DevOps
Git
Linux
Robotics
ROS
Nav2
ROS2
MoveIt
SLAM
Path Planning
Motion Planning
Perception
Apply
In office • Full-Time • 1+ year exp • Bachelor's Degree • Tel Aviv
Python
Bash
DevOps
CI/CD
Jenkins
Docker
Kubernetes
GitLab
Linux
Management
Agile
Scrum
QA
Pytest
Apply
$38k – $82k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Shanghai
Python
C++
AI/ML
Synthetic Data
Edge AI
Physical AI
DevOps
Linux
Robotics
ROS
Isaac Sim
Isaac Lab
Sim-to-Real
Apply
$270k – $330k per year • In office • Full-Time • 7+ years exp • New York
AI/ML
NVLink
DevOps
Linux
Analytics
Seaborn
Matplotlib
Apply
$250k – $300k per year • In office • Full-Time • 5+ years exp • San Francisco • New York
Python
AI/ML
NVLink
DevOps
Linux
BGP
Analytics
Seaborn
Matplotlib
Apply
$250k – $300k per year • In office • Full-Time • 5+ years exp • San Francisco • New York
Rust
DevOps
Amazon S3
Linux
Analytics
Seaborn
Matplotlib
Apply
$140k – $276k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
DevOps
CI/CD
Analytics
Seaborn
Matplotlib
Apply
Account Manager 6 days ago
$300k per year • Remote • Full-Time • San Francisco • New York
Analytics
Seaborn
Matplotlib
Apply
$245k – $279k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco • McLean • Cambridge • San Jose • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
AI Agents
PyTorch
LLM
CUDA
Hugging Face
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$245k – $279k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • McLean • San Francisco • Richmond • Chicago • New York
Python
Go
JavaScript
Rust
TypeScript
C#
Scala
AI/ML
AI Agents
Machine Learning
DevOps
GCP
Azure
AWS
HPC
Apply
Cloud Engineer 5 5 hours ago
$230k – $262k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • McLean • San Francisco • San Jose • Richmond • New York
Python
JavaScript
Java
Node JS
Scala
DevOps
Terraform
GCP
CloudFormation
Crossplane
Pulumi
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$197k – $225k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Jose • San Francisco • McLean • Cambridge • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
Fine-tuning
AI Agents
PyTorch
LLM
CUDA
Hugging Face
TPU
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$219k – $250k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Jose • San Francisco • McLean • Cambridge • New York
AI/ML
Fine-tuning
RLHF
Quantization
NLP
Transfer Learning
VLM
PyTorch
LLM
Tokenization
Self-Supervised Learning
Hugging Face
SFT
Pre-training
Machine Learning
DevOps
AWS
Apply
See all jobs
This is one of many
706,063 more open roles from verified company boards, updated every day.