708,596open jobs
42,039companies
101,035added this week
Browse all
Salary
$300k – $350k per year
Location
In office (San Francisco)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role

Modal's LLM inference platform delivers frontier performance for open-source models with best-in-class elasticity and developer experience, made in part possible by our custom runtime with GPU memory snapshots and multi-cloud substrate.

We're looking for a leader to own the direction and execution of this platform to continue to establish us as the clear market leader, working closely with customers like Cognition, Doordash, Ramp, and many more. You'll be leading a group of highly talented engineers working on our market-leading LLM inference offering, spanning the serving stack, routing infrastructure, internal agentic optimization platform, and the user-facing product surface area.

This is a hands-on leadership role - expect to split your time between technical contribution, product shaping and people management depending on what the team needs. You'll set direction, remove blockers, and build a strong engineering culture as your team tackles hard problems in distributed computing, frontier inference serving, and performance optimization.

Responsibilities

Team

  • Recruit, hire, and grow a high-performing team of engineers; run regular 1:1s focused on coaching, feedback, and career growth.

  • Set clear performance expectations, hold a high bar, and build an environment where engineers do their best work.

  • Foster a culture of ownership, accountability, customer obsession, and continuous improvement.

Technical and Product Direction

  • Drive technical and product decisions through design reviews, code reviews, and architectural discussions.

  • Lead our efforts working with customers with novel or frontier workloads so they can be successful running on Modal.

  • Translate learnings from frontier customers into a roadmap for our internal optimization platform and user-facing product, so that gains from optimization and research efforts are accessible to all of our users.

  • Establish standards for reliability and product excellence; ensure the team owns projects end-to-end, from spec through production.

Cross-Functional Leadership

  • Partner with business operations and compute strategy on continuing to develop our strategy for compute purchases across a variety of accelerators.

  • Collaborate with GTM on product launches, positioning, and improving win rates for inference opportunities.

  • Help guide the roadmap for teams building the infrastructure platform underlying inference, as well as adjacent product teams.

Requirements

  • 10+ years of industry experience, including 3+ years in a leadership role

  • Track record building high-performance systems at scale

  • Strong background in cloud infrastructure

  • Deep knowledge of low-level OS foundations (Linux kernel, file systems, containers, etc.)

  • Nice to have: Experience working with LLM inference in production and familiarity with underlying concepts like engines, kernels, routing, KV cache management and speculative decoding.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
708,596 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Lead AI Engineer 1 day ago
$40k – $96k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
Java
Node JS
AI/ML
Copilot
Claude
Model Context Protocol
Prompt Engineering
AI Agents
NLP
AWS Bedrock
LLM
Amazon SageMaker
AWS Bedrock AgentCore
DevOps
OpenShift
CloudFormation
CI/CD
Jenkins
AWS
Docker
Kubernetes
Platform Engineering
AWS Lambda
Amazon EC2
GitLab
Amazon S3
Apply
$250k – $325k per year • In office • Full-Time • New York
Python
JavaScript
TypeScript
AI/ML
Function Calling
AI Agents
LLM
LLM Evaluation
Tool Use
Frontend
React.js
Apply
$32k – $88k per year (Estimated) • In office • Moscow
AI/ML
Embeddings
Computer Vision
NLP
Speech Recognition
LLM
RAG
Whisper
Hugging Face
Machine Learning
Analytics
ETL/ELT
Apply
Editor 1 day ago
$65k – $70k per year • In office • 3+ years exp • Bachelor's Degree
AI/ML
AI Agents
Management
Google Sheets
Google Docs
Agile
Apply
$37k – $97k per year (Estimated) • In office • Ho Chi Minh City
Python
Go
JavaScript
TypeScript
SQL
Databases
MySQL
PostgreSQL
Redis
pgvector
RabbitMQ
ElasticSearch
Apache Kafka
OpenSearch
AI/ML
Claude
Claude Code
Embeddings
Function Calling
Gemini
LLM
RAG
OpenAI
Structured Outputs
LLM Guardrails
Frontend
Vue.js
Pinia
Vite
Vue Router
Mobile
Dependency Injection
DevOps
Rest API
Helm
GitHub Actions
OpenTelemetry
Prometheus
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
QA
Playwright
Apply
$140k – $276k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
DevOps
CI/CD
Analytics
Seaborn
Matplotlib
Apply
Account Manager 7 days ago
$300k per year • Remote • Full-Time • San Francisco • New York
Analytics
Seaborn
Matplotlib
Apply
$140k – $165k per year • In office • Full-Time • New York
Analytics
Seaborn
Matplotlib
Apply
People Lead, UK 13 days ago
$79k – $170k per year (Estimated) • In office • Full-Time • 5+ years exp • London
Analytics
Seaborn
Matplotlib
Apply
$150k – $270k per year • In office • Full-Time • New York
SQL
AI/ML
LLM
DevOps
Kubernetes
Linux
Cybersecurity
SIEM
Analytics
Seaborn
Matplotlib
Apply
$245k – $279k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco • McLean • Cambridge • San Jose • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
AI Agents
PyTorch
LLM
CUDA
Hugging Face
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$245k – $279k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • McLean • San Francisco • Richmond • Chicago • New York
Python
Go
JavaScript
Rust
TypeScript
C#
Scala
AI/ML
AI Agents
Machine Learning
DevOps
GCP
Azure
AWS
HPC
Apply
Cloud Engineer 5 11 hours ago
$230k – $262k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • McLean • San Francisco • San Jose • Richmond • New York
Python
JavaScript
Java
Node JS
Scala
DevOps
Terraform
GCP
CloudFormation
Crossplane
Pulumi
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$197k – $225k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Jose • San Francisco • McLean • Cambridge • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
Fine-tuning
AI Agents
PyTorch
LLM
CUDA
Hugging Face
TPU
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$219k – $250k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Jose • San Francisco • McLean • Cambridge • New York
AI/ML
Fine-tuning
RLHF
Quantization
NLP
Transfer Learning
VLM
PyTorch
LLM
Tokenization
Self-Supervised Learning
Hugging Face
SFT
Pre-training
Machine Learning
DevOps
AWS
Apply
See all jobs
This is one of many
708,596 more open roles from verified company boards, updated every day.