371,660open jobs
9,621companies
49,388added this week
Browse all
Salary
$109k – $145k per year
Location
Remote/Hybrid (San Jose, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Lambda is a specialized AI infrastructure provider that offers high-performance GPU cloud compute, clusters, and hardware tailored for deep learning and machine learning workloads. The company enables AI developers and research teams to train, fine-tune, and deploy large language models efficiently through scalable cloud instances and dedicated on-premise GPU servers.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

The Operations team plays a critical role in ensuring the seamless end-to-end execution of our AI-IaaS infrastructure and hardware. This team is responsible for sourcing all necessary infrastructure and components, overseeing day-to-day data center operations to maintain optimal performance and uptime, and driving cross company coordination through product management organization to align operational capabilities with strategic goals. By managing the full lifecycle from procurement to deployment and operational efficiency, the Operations team ensures that our AI-driven infrastructure is reliable, scalable, and aligned with business priorities.

What You’ll Do

  • Track, log, and manage all quality issues arising in the data center during deployment and production environment

  • Perform root cause analysis (RCA) for every failure (hardware, software, process)

  • Analyze production system metrics and quality data to detect trends, anomalies, or weak points

  • Improve turnaround time (TAT) for Return Merchandise Authorization (RMA) processes

  • Design, monitor, and drive corrective and preventive actions (CAPA)

  • Implement and verify containment actions to keep systems operational until permanent fixes are applied.

  • Collaborate with operations, hardware, engineering, supply chain, and vendors to resolve quality issues

  • Capture and upload failure analysis (FA) reports and related data into Quality Management Systems (QMS)

  • Verify quality of spares (incoming and outgoing) to avoid repeat failures.

  • Define and track quality KPIs / SLAs and report on quality performance to leadership

  • Oversee MRB (Material Review Board) inventory, rework, disposal decisions

  • Ensure the quality management system (QMS) is up to date, with necessary training rolled out

  • Work cross-functionally during hardware ramp, deployments, and upgrades to ensure quality gates

  • Up to 30% travel may be required for this role.

You

  • Have experience working with hardware / data center / infrastructure systems

  • Are strong at data analysis, statistics, and metrics (you can turn raw data into insight)

  • Are skilled in root cause analysis methods (5 Whys, fishbone, 8D, A3, etc.)

  • Are comfortable managing cross-team communication, stakeholder expectations, and conflict resolution

  • Are detail-oriented, process-driven, and quality-minded

  • Have experience working with quality tools or QMS software (e.g. audit modules, ERP, defect tracking)

  • Communicate clearly in English (both written and verbal)

Nice to Have

  • Experience in the machine learning / AI infrastructure / GPU / HPC / computer hardware industry

  • Exposure to data center standards, certifications (e.g. ISO, Uptime Institute, etc.)

  • Experience working on vendor quality, supply chain quality, or incoming inspections

  • Understanding of firmware, embedded systems, reliability engineering

  • Familiarity with scripting or automation (Python, SQL, etc.) to help with data processing

  • Exposure to cloud or hyperscaler infrastructure operations

  • Experience with “manufacturing-like” quality concepts applied to compute hardware

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
371,660 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$42k – $104k per year (Estimated) • In office • Internship • 5+ years exp
Bash
PowerShell
Python
DevOps
Amazon EC2
Amazon EKS
Ansible
AWS
AWS Lambda
Azure
CentOS Stream
Chef
CI/CD
CloudFormation
Configuration Management
Datadog
FinOps
GCP
Git
GitHub Actions
GitLab CI
Grafana
Hyper-V
Istio
Jenkins
Kubernetes
KVM
Linkerd
Platform Engineering
Prometheus
Puppet
Service Mesh
Splunk
Terraform
Ubuntu
VMWare
Windows Server
Amazon CloudWatch
Amazon ECS
Amazon S3
API Gateway
AWS Step Functions
GitHub
GitLab
IAM
Cybersecurity
GDPR
ISO 27001
SOC 2
Apply
Data Engineer 1 day ago
In office • Full-Time • Singapore
Node JS
Python
SQL
JavaScript
Python
Beautiful Soup
Databases
Apache Kafka
MySQL
PostgreSQL
RabbitMQ
SQLite
AI/ML
Hadoop
Spark
DevOps
AWS
AWS Lambda
Azure
CI/CD
GCP
Amazon ECS
Amazon EventBridge
Amazon S3
Analytics
ETL/ELT
QA
Selenium
Apply
up to $48k per year (net) • Remote/Hybrid • Moscow
Node JS
Python
TypeScript
JavaScript
Node JS
Nest.JS
Python
Django
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
PostgreSQL
Qdrant
RabbitMQ
Redis
AI/ML
Chain-of-Thought
Claude
Claude Code
Copilot
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
RAG
Anthropic
Function Calling
OpenAI
Structured Outputs
Frontend
GraphQL
Next.js
React.js
Redux
Redux Toolkit
Zustand
Mobile
State Management
DevOps
AWS
CI/CD
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Kubernetes
Prometheus
Yandex Cloud
GitHub
GitLab
Apply
$24k – $64k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Chennai
Python
Scala
SQL
TypeScript
JavaScript
Java
Python
pySpark
Java
Spring Boot
Databases
Apache Kafka
Databricks
AI/ML
AI Agents
Copilot
Google ADK
LLM
NLP
Prompt Engineering
Spark
Devin
Model Context Protocol
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
OpenShift
GitHub
Analytics
ETL/ELT
Apply
Backend Engineer 1 day ago
$44k – $56k per year • In office • Full-Time • 3+ years exp • Tokyo
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
Frontend
Next.js
React.js
DevOps
Amazon EC2
AWS
AWS CDK
CI/CD
Docker
GitHub Actions
Vercel
Amazon CloudWatch
Amazon S3
GitHub
IAM
Management
Linear
Apply
$89k – $119k per year • Equity • In office • Full-Time • Atlanta
DevOps
AWS Lambda
AWS
Marketing
Zendesk
Apply
$380k – $445k per year • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • San Francisco
Go
Python
SQL
Databases
DynamoDB
Presto
DevOps
AWS Lambda
AWS
Amazon S3
Apply
$297k – $440k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • San Jose • Bellevue
AI/ML
InfiniBand
DevOps
AWS Lambda
Incident Management
AWS
HPC
Apply
$278k – $325k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • San Jose
Databases
Google BigQuery
DevOps
AWS Lambda
AWS
Cybersecurity
Okta
Management
Google Workspace
Jira
Marketing
Salesforce
Apply
$137k – $183k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Elk Grove Village
AI/ML
InfiniBand
DevOps
AWS Lambda
AWS
Apply
$145k – $180k per year • In office • Full-Time • 5+ years exp • San Jose
Verilog
Apply
Sr System Engineer 2 hours ago
$110k – $171k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Jose
C++
DevOps
Git
Ubuntu
Apply
$148k – $282k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • San Jose
AI/ML
LLM
Prompt Engineering
Apply
$124k – $244k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • San Jose
Python
Cybersecurity
Zero Trust
Zscaler
Apply
$172k – $323k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • San Jose
Python
Cybersecurity
Zero Trust
Zscaler
Apply
See all jobs
This is one of many
371,660 more open roles from verified company boards, updated every day.