368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$240k – $312k per year
Location
Remote/Hybrid (San Francisco, San Jose, Bellevue, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Lambda is a specialized AI infrastructure provider that offers high-performance GPU cloud compute, clusters, and hardware tailored for deep learning and machine learning workloads. The company enables AI developers and research teams to train, fine-tune, and deploy large language models efficiently through scalable cloud instances and dedicated on-premise GPU servers.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.

What You'll Do

  • Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure

  • Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs

  • Develop tooling and automation to reduce operational toil and improve reliability

  • Collaborate with software, platform, and networking teams to improve service reliability and deployment workflows

  • Deploy and maintain network monitoring, observability, and management tools

  • Improve deployment safety through CI/CD pipelines, GitOps workflows, testing, and progressive rollouts

  • Drive operational excellence through observability, incident management, capacity planning, postmortems, and participation in the on-call rotation

You

  • Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role

  • Have experience operating and supporting large-scale distributed systems in production

  • Have experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations

  • Have experience participating in on-call rotations and incident response

  • Have strong troubleshooting skills across Linux systems, Kubernetes, distributed systems, and networking

  • Have experience with observability platforms, monitoring, alerting, and metrics

  • Are comfortable working on the Linux command line and have a solid understanding of the Linux networking stack

  • Have experience with multi-datacenter and hybrid cloud environments

  • Have experience automating infrastructure and operational workflows using Python, Ansible, or similar tools

  • Have experience designing and operating CI/CD and GitOps deployment workflows

Nice To Have

  • Experience building and operating Software Defined Networks (SDN), including OpenStack Neutron, OVN, and OVS

  • Experience operating production-scale SDNs in a cloud environment (e.g., infrastructure powering AWS VPC-like networking services)

  • Software development experience in Go and/or Python (C is a plus)

  • Experience automating infrastructure and network configuration using Kubernetes, Helm, Terraform, and Ansible

  • Deep understanding of the Linux networking stack and its interaction with network virtualization technologies, SR-IOV, and DPDK

  • Understanding of the SDN ecosystem and modern cloud networking architectures

  • Experience diagnosing complex production issues across infrastructure, networking, and application layers

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$68k – $104k per year (Estimated) • Equity • Remote • Full-Time • 3+ years exp • Warsaw
AI/ML
Anthropic
DevOps
AWS
Azure
CI/CD
GCP
Helm
Kubernetes
Platform Engineering
Terraform
Cybersecurity
SonarQube
Apply
$82k – $200k per year (Estimated) • Equity • Remote • Full-Time
C#
C#
ASP.NET Core
Blazor
Entity Framework Core
SignalR
xUnit
AI/ML
Anthropic
Mobile
Dependency Injection
DevOps
AWS
Azure
CI/CD
GCP
Platform Engineering
Apply
$63k – $97k per year (Estimated) • Equity • Remote • Full-Time • Warsaw
Java
Kotlin
Scala
Scala
Cats
Cats Effect
ZIO
Databases
Amazon Aurora
Apache Kafka
Databricks
DynamoDB
PostgreSQL
Snowflake
AI/ML
Anthropic
Frontend
GraphQL
DevOps
AWS
Azure
CI/CD
GCP
Platform Engineering
Terraform
Amazon Kinesis
Analytics
A/B Testing
Apply
$15k – $34k per year (Estimated) • In office • 3+ years exp • Yekaterinburg
DevOps
CI/CD
Docker
Apply
$137k – $246k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Austin
Java
DevOps
AWS
CI/CD
Docker
Git
GitHub Actions
GitLab CI
Jenkins
Kubernetes
GitHub
GitLab
Management
Confluence
Jira
Apply
$297k – $440k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • San Jose • Bellevue
AI/ML
InfiniBand
DevOps
AWS Lambda
Incident Management
AWS
HPC
Apply
$278k – $325k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • San Jose
Databases
Google BigQuery
DevOps
AWS Lambda
AWS
Cybersecurity
Okta
Management
Google Workspace
Jira
Marketing
Salesforce
Apply
$137k – $183k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp
AI/ML
InfiniBand
DevOps
AWS Lambda
AWS
Apply
$251k – $335k per year • Equity • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • San Jose
AI/ML
RAG
DevOps
AWS Lambda
Kubernetes
AWS
HPC
Apply
$296k – $395k per year • Equity • Remote/Hybrid • Full-Time • San Francisco • San Jose
DevOps
AWS
AWS Lambda
Azure
GCP
HPC
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$293k – $385k per year • In office • Full-Time • San Francisco
AI/ML
OpenAI
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.