368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$314k – $465k per year
Location
Remote/Hybrid (San Francisco, San Jose, Bellevue, United States)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Lambda is a specialized AI infrastructure provider that offers high-performance GPU cloud compute, clusters, and hardware tailored for deep learning and machine learning workloads. The company enables AI developers and research teams to train, fine-tune, and deploy large language models efficiently through scalable cloud instances and dedicated on-premise GPU servers.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco/Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

In the world of distributed AI, raw GPU and CPU horsepower is just a part of the story. High-performance networking and storage are the critical components that enable and unite these systems, making groundbreaking AI training and inference possible.

The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage, networking, GPU and CPU hardware.

Our expertise lies at the intersection of:

  • High-Performance Distributed Storage Solutions and Protocols: We engineer the protocols and systems that serve massive datasets at the speeds demanded by modern clustered GPUs.

  • Dynamic Networking: We design advanced networks that provide multi-tenant security and intelligent routing without compromising performance, using the latest in AI networking hardware.

  • Compute Virtualization: We enable cutting-edge virtualization and clustering that allows AI researchers and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs.

About the Role:

We are seeking a seasoned Staff Storage Software Engineer with deep experience designing and deploying storage protocol solutions at scale across object, block, and file paradigms.

This is a unique opportunity to work at the intersection of large-scale distributed systems and the rapidly evolving field of artificial intelligence infrastructure. This is an opportunity to have a significant impact on the future of AI. You will be building the foundational infrastructure that powers some of the most advanced AI research and products in the world.

What You’ll Do

  • Technical Leadership: Set technical direction for storage software architecture across petabyte-scale deployments, authoring and reviewing design docs, mentoring senior engineers, and serving as the technical anchor for cross-functional initiatives spanning storage, networking, compute, and control plane teams. Represent the storage software team in architectural reviews, roadmap planning, and customer-facing technical discussions.

  • Execution: Design, develop, and maintain high-performance storage systems software across file (NFS, SMB, Lustre), block (NVMe-oF, iSCSI), and object (S3) protocols. Build distributed systems for orchestrating storage resources, integrate with NVMe/GPU-direct/DPU-accelerated hardware, and troubleshoot complex production issues across performance, protocol, and hardware failure domains. Own the full lifecycle from requirements and design through deployment, monitoring, and maintenance, including benchmarking, profiling, and capacity planning tooling.

  • Collaboration: Partner closely with storage software, networking, control plane, Kubernetes, observability, compute, and fleet engineering teams to deliver cross-functional infrastructure initiatives, define and track storage SLOs/SLIs, and ensure reliable deployment and maintenance of distributed storage infrastructure.

  • Innovate: Stay current with AI and HPC storage research, evaluate emerging protocols and hardware (from open-source filesystems to vendor-specific accelerated storage), and optimize solutions for AI workloads including checkpoint I/O, high-throughput dataset serving, and latency-sensitive inference pipelines.

You Have:

  • Experience: 10+ years in storage systems engineering, with 5+ years in a technical lead or Staff+ IC role. Proven track record designing and operating multi-petabyte storage infrastructure in production data center or cloud environments. Background in HPC, AI/ML infrastructure, or large-scale cloud storage.

  • Systems-Level Programming: Strong proficiency in C, C++, Rust, or Go. Ability to write high-performance, concurrent, production-grade systems code. Familiarity with DPDK/SPDK and kernel-bypass data paths is a plus; kernel-level storage driver or storage daemon experience is even better.

  • Storage Protocol & API Expertise: Deep hands-on experience with two or more protocols, object (S3), block (iSCSI, NVMe-oF), or file (NFS, SMB, Lustre, DAOS), including implementing or maintaining protocol servers/clients in production, not just consuming them.

  • Storage Performance Optimization: Experience profiling and tuning for throughput, latency, and IOPS under real workloads using tools like fio, elbencho.

  • Modern Storage Technologies: Working knowledge of NVMe, NVMe-oF, RDMA (RoCE or InfiniBand), and DPUs (e.g., NVIDIA BlueField).

  • Operational Acumen: Comfortable in physical data center environments, rack-scale infrastructure, storage hardware, failure domains. Experience designing for reliability, writing runbooks, and driving incident response. Familiar with storage observability tooling (Prometheus, Grafana, log aggregation, tracing).

Nice to Have

  • Experience with NVIDIA BlueField DPUs or SuperNICs for accelerated storage data paths, including GPUDirect Storage implementation.

  • Deep production experience with enterprise or HPC storage platforms: Vast Data, Weka, NetApp, or Lustre.

  • Experience deploying and operating Ceph like service at scale (100PB+) in an HPC or AI infrastructure environment.

  • Familiarity with emerging storage technologies such as CXL memory pooling, computational storage, or ZNS (Zoned Namespace) SSDs.

  • Experience contributing to or maintaining open-source storage projects (e.g., Ceph, DAOS, Lustre, MinIO).

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$71k – $170k per year (Estimated) • In office • Full-Time • Netanya
Python
TypeScript
AI/ML
Accelerate
Fine-tuning
LangChain
LLM
NLP
Prompt Engineering
PyTorch
RAG
Edge AI
Hugging Face
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$18k – $46k per year (Estimated) • In office • Full-Time • Moscow
DevOps
Ansible
ArgoCD
AWS
CI/CD
Docker
GitLab CI
Helm
Kubernetes
OpenTofu
Terraform
Terragrunt
Yandex Cloud
GitLab
Apply
$20k – $54k per year (Estimated) • In office • Full-Time • Pune
Java
SQL
DevOps
GCP
Incident Management
Kubernetes
IAM
QA
JMeter
Selenium
Apply
AI Engineer 1 day ago
$25k – $103k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
Embeddings
Gemini
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Semantic Search
Spark
TensorFlow
Hugging Face
LLM Guardrails
LLMOps
OpenAI
Semantic Search
DevOps
AWS
Azure
CI/CD
Apply
$24k – $55k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Moscow
DevOps
AWS
Azure
SLI/SLO/SLA
Cybersecurity
GDPR
Management
Jira
ServiceNow
Apply
$89k – $119k per year • Equity • In office • Full-Time • Atlanta
DevOps
AWS Lambda
AWS
Marketing
Zendesk
Apply
$380k – $445k per year • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • San Francisco
Go
Python
SQL
Databases
DynamoDB
Presto
DevOps
AWS Lambda
AWS
Amazon S3
Apply
$297k – $440k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • San Jose • Bellevue
AI/ML
InfiniBand
DevOps
AWS Lambda
Incident Management
AWS
HPC
Apply
$278k – $325k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • San Jose
Databases
Google BigQuery
DevOps
AWS Lambda
AWS
Cybersecurity
Okta
Management
Google Workspace
Jira
Marketing
Salesforce
Apply
$137k – $183k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Elk Grove Village
AI/ML
InfiniBand
DevOps
AWS Lambda
AWS
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 2 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.