706,797open jobs
41,963companies
99,144added this week
Browse all
Salary
$250k – $300k per year
Location
In office (San Francisco, New York)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luig i), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on top of it, whether that is a kernel panic, a broadcast storm during boot, or getting our container runtime to run on new architectures and platforms.

Requirements:

  • 5+ years of experience writing high-quality production code

  • Experience operating fleets of physical hardware (bare metal provisioning, BMC/IPMI, PXE or network boot, firmware) or building the control planes that manage them (the more challenges you've worked through, the better)

  • Strong cloud skills

  • Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers, etc.)

  • Effective at debugging across layers, from BGP flapping, Linux RPS, and vBIOS bugs to a Python control-plane service

  • Willingness to step into the thick of it with our on-call rotation and respond to production incidents

Nice-to-Haves:

  • Experience with GPUs and the NVIDIA software stack in production (drivers, health monitoring, XIDs, RDMA/NVLink)

  • Prior experience with Go

Key Things the Team Is Working On:

  • Automatic remediation of unhealthy machines (power cycling, reimaging, GPU recovery) to maximize uptime and minimize operator toil.

  • Automatic integration of new CPU, GPU, and storage servers into the fleet while managing hardware and network heterogeneity.

  • Network health monitoring and reliability across many datacenters, and standardization of bare metal network configuration.

  • Automatic hardware acceptance testing and benchmarking (CPU, disk, GPU, interconnect, network).

  • Custom network bootloader, machine image pipeline, and kernel and firmware management across the fleet.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
706,797 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$81k – $194k per year (Estimated) • Remote • Contractor • Bachelor's Degree
Python
Go
JavaScript
Java
Rust
TypeScript
C++
Node JS
Apply
$60k – $151k per year (Estimated) • In office • Full-Time • 5+ years exp • New Zealand
DevOps
Ansible
Red Hat
VMWare
Dynatrace
Linux
Management
Confluence
Jira
ServiceNow
SharePoint
Apply
$76k – $164k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sydney
Python
C
C++
C
Embedded C
DevOps
CI/CD
Jenkins
Bitbucket
GitHub
Management
Confluence
Jira
Agile
Apply
$58k per year • In office • Internship • San Francisco
Python
Apply
In office • Internship
Python
Analytics
Microsoft Excel
Apply
$270k – $330k per year • In office • Full-Time • 7+ years exp • San Francisco • New York
Rust
DevOps
Amazon S3
Linux
Analytics
Seaborn
Matplotlib
Apply
$270k – $330k per year • In office • Full-Time • 7+ years exp • New York
AI/ML
NVLink
DevOps
Linux
Analytics
Seaborn
Matplotlib
Apply
$250k – $300k per year • In office • Full-Time • 5+ years exp • San Francisco • New York
Rust
DevOps
Amazon S3
Linux
Analytics
Seaborn
Matplotlib
Apply
$140k – $276k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
DevOps
CI/CD
Analytics
Seaborn
Matplotlib
Apply
Account Manager 6 days ago
$300k per year • Remote • Full-Time • San Francisco • New York
Analytics
Seaborn
Matplotlib
Apply
$245k – $279k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco • McLean • Cambridge • San Jose • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
AI Agents
PyTorch
LLM
CUDA
Hugging Face
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$245k – $279k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • McLean • San Francisco • Richmond • Chicago • New York
Python
Go
JavaScript
Rust
TypeScript
C#
Scala
AI/ML
AI Agents
Machine Learning
DevOps
GCP
Azure
AWS
HPC
Apply
Cloud Engineer 5 8 hours ago
$230k – $262k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • McLean • San Francisco • San Jose • Richmond • New York
Python
JavaScript
Java
Node JS
Scala
DevOps
Terraform
GCP
CloudFormation
Crossplane
Pulumi
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$197k – $225k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Jose • San Francisco • McLean • Cambridge • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
Fine-tuning
AI Agents
PyTorch
LLM
CUDA
Hugging Face
TPU
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$219k – $250k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Jose • San Francisco • McLean • Cambridge • New York
AI/ML
Fine-tuning
RLHF
Quantization
NLP
Transfer Learning
VLM
PyTorch
LLM
Tokenization
Self-Supervised Learning
Hugging Face
SFT
Pre-training
Machine Learning
DevOps
AWS
Apply
See all jobs
This is one of many
706,797 more open roles from verified company boards, updated every day.