368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$230k – $340k per year
Location
Remote/Hybrid (San Francisco, San Jose, Bellevue, United States)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Lambda is a specialized AI infrastructure provider that offers high-performance GPU cloud compute, clusters, and hardware tailored for deep learning and machine learning workloads. The company enables AI developers and research teams to train, fine-tune, and deploy large language models efficiently through scalable cloud instances and dedicated on-premise GPU servers.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.

What You’ll Do

  • Architect, deploy, and operate Kubernetes clusters across AWS and Lambda's bare-metal datacenters.

  • Build and maintain automation for cluster lifecycle management - provisioning, upgrades, and scaling.

  • Own the reliability, performance, and security of Kubernetes workloads in production.

  • Implement observability, logging, and alerting for clusters and critical workloads.

  • Partner with product teams to design scalable, cloud-native services and CI/CD pipelines.

  • Set the standards for resource management, networking, and RBAC across the platform.

  • Lead incident response, root-cause analysis, and post-mortems for platform issues.

  • Mentor engineers and raise the bar for platform engineering across the org.

You

  • 5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale.

  • Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting).

  • Strong with Helm, Kustomize, or similar, and GitOps-based delivery.

  • Proficient with infrastructure-as-code (Terraform, Pulumi, or equivalent).

  • Solid grounding in networking, service meshes, and container runtimes.

  • Hands-on with observability stacks (Prometheus, Grafana, OpenTelemetry).

  • Strong coding skills in Go or Python for automation and tooling.

  • Practical security experience: network policies, secrets management, and image scanning.

Nice to Have

  • Experience with multi-cluster, multi-cloud, or hybrid environments.

  • Knowledge of GPU scheduling, HPC workloads, or ML/AI infrastructure.

  • Experience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows).

  • Exposure to cost optimization and capacity planning for large clusters.

  • Contributions to CNCF or Kubernetes open-source projects.

  • CKA/CKS certification.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$18k – $46k per year (Estimated) • Remote • Full-Time • Tula
C#
JavaScript
Node JS
SQL
TypeScript
C#
.NET
Node JS
InversifyJS
Databases
DynamoDB
MySQL
AI/ML
Claude
Copilot
Cursor
OpenAI Codex
Frontend
Angular
React.js
Tailwind CSS
Mobile
Dependency Injection
DevOps
AWS
AWS Lambda
CI/CD
OpenTelemetry
Rest API
Terraform
Amazon CloudWatch
Amazon S3
API Gateway
GitHub
Cybersecurity
HIPAA
Apply
$19k – $53k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
Bash
JavaScript
Python
TypeScript
Frontend
Angular
React.js
DevOps
AWS
Azure
Datadog
Docker
GCP
Grafana
Kubernetes
Prometheus
Splunk
IAM
Cybersecurity
Keycloak
Apply
SDET 1 day ago
$12k – $39k per year (Estimated) • In office • 4+ years exp • Gurgaon
Java
DevOps
AWS
Azure
CI/CD
Jenkins
QA
Appium
JMeter
Playwright
Rest-Assured
Selenium
Apply
$20k – $45k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Python
pySpark
Databases
Databricks
Microsoft Fabric
AI/ML
Hadoop
Spark
DevOps
AWS
Azure
Analytics
Power BI
Tableau
Apply
$80k – $147k per year (Estimated) • Remote/Hybrid • Full-Time • Sydney
DevOps
CI/CD
Apply
$89k – $119k per year • Equity • In office • Full-Time • Atlanta
DevOps
AWS Lambda
AWS
Marketing
Zendesk
Apply
$380k – $445k per year • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • San Francisco
Go
Python
SQL
Databases
DynamoDB
Presto
DevOps
AWS Lambda
AWS
Amazon S3
Apply
$297k – $440k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • San Jose • Bellevue
AI/ML
InfiniBand
DevOps
AWS Lambda
Incident Management
AWS
HPC
Apply
$278k – $325k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • San Jose
Databases
Google BigQuery
DevOps
AWS Lambda
AWS
Cybersecurity
Okta
Management
Google Workspace
Jira
Marketing
Salesforce
Apply
$137k – $183k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp
AI/ML
InfiniBand
DevOps
AWS Lambda
AWS
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 27 min ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.