368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$314k – $465k per year
Location
Remote/Hybrid (Bellevue, San Francisco, San Jose, United States)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Lambda is a specialized AI infrastructure provider that offers high-performance GPU cloud compute, clusters, and hardware tailored for deep learning and machine learning workloads. The company enables AI developers and research teams to train, fine-tune, and deploy large language models efficiently through scalable cloud instances and dedicated on-premise GPU servers.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

About the Role

As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.

The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.

What You'll Do

We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:

  • Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.

  • Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.

  • Guide design of compute platform multi-tenant security model

  • Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.

  • Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.

  • Work with customers on translating vague customer technical requirements into concrete engineering deliverables.

  • Set engineering standards and lead design reviews for mission-critical cloud software at scale.

Who You are

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.

  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.

  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.

  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.

  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).

  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.

Nice to Have

  • Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).

  • Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)

  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).

  • Experience with Cloud Service Provider Kubernetes offerings.

  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bellevue
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$120k – $220k per year • Equity 0.2–0.8% • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
C++
Go
JavaScript
Python
TypeScript
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
$89k – $119k per year • Equity • In office • Full-Time • Atlanta
DevOps
AWS Lambda
AWS
Marketing
Zendesk
Apply
$380k – $445k per year • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • San Francisco
Go
Python
SQL
Databases
DynamoDB
Presto
DevOps
AWS Lambda
AWS
Amazon S3
Apply
$297k – $440k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • San Jose • Bellevue
AI/ML
InfiniBand
DevOps
AWS Lambda
Incident Management
AWS
HPC
Apply
$278k – $325k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • San Jose
Databases
Google BigQuery
DevOps
AWS Lambda
AWS
Cybersecurity
Okta
Management
Google Workspace
Jira
Marketing
Salesforce
Apply
$137k – $183k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp
AI/ML
InfiniBand
DevOps
AWS Lambda
AWS
Apply
Engineering Manager 5 hours ago
$161k – $310k per year (Estimated) • In office • Bachelor's Degree • Bellevue
AI/ML
Edge AI
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
$236k – $339k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bellevue
Java
Python
Databases
Snowflake
AI/ML
AI Agents
Feature Store
Apply
$160k – $210k per year • Remote/Hybrid • Full-Time • 6+ years exp • Denver • Bellevue
C#
C++
Go
Python
SQL
Databases
Azure Cosmos DB
Azure SQL Database
DynamoDB
MySQL
AI/ML
Edge AI
DevOps
AWS
Azure
Azure AKS
CI/CD
GCP
Google GKE
Kubernetes
SLI/SLO/SLA
Analytics
Power BI
Management
UiPath
Apply
$159k – $285k per year (Estimated) • In office • Full-Time • 10+ years exp • Bellevue
AI/ML
AI Agents
Management
Smartsheet
Apply
$152k – $278k per year (Estimated) • In office • 8+ years exp • Bellevue
SQL
Cybersecurity
GDPR
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.