375,679open jobs
9,756companies
47,886added this week
Browse all
Salary
$231k – $342k per year
Location
Remote/Hybrid (San Francisco, San Jose, Bellevue, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Lambda is a specialized AI infrastructure provider that offers high-performance GPU cloud compute, clusters, and hardware tailored for deep learning and machine learning workloads. The company enables AI developers and research teams to train, fine-tune, and deploy large language models efficiently through scalable cloud instances and dedicated on-premise GPU servers.

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

The Technical Account Manager owns the technical health of the post-sales relationship for Lambda’s public cloud accounts, spanning AI-native startups, Enterprises, and Fortune 500 companies. Where the Customer Success Manager owns the commercial health of an account, you own its technical health: the customer’s workloads run well, the architecture is right, the SLA story is defensible, and technical risk is found and retired before it threatens revenue. Solutions Engineering carries the account through pre-sales and hypercare; at handoff, you take ownership of the technical relationship for the life of the contract.

This is a hands-on role, not a coordination role. You will understand what customers are actually building (training runs, fine-tuning pipelines, inference services) deeply enough to lead joint POC sessions, design and defend architectures, validate SLA events at the root-cause level, and build the tooling that makes account health measurable. You will be the customer’s most credible technical advocate inside Lambda and Lambda’s most trusted technical voice inside the account.

What You’ll Do

  • Own the technical health of your accounts. Take the technical handoff from Solutions Engineering at the end of hypercare and own the account’s technical outcomes through steady state, expansion, and renewal. Know the state of every cluster and workload you are accountable for, and keep your commercial counterparts ahead of technical risk.

  • Understand customer AI use cases end to end. Map what it means for each customer to train, fine-tune, and serve models on Lambda: frameworks, schedulers, parallelism strategy, data paths, and performance baselines. Build the customer user journey and convert it into value-add opportunities across the platform, documentation, and escalation routing.

  • Lead joint customer POC sessions. Define success criteria with the customer before a node is provisioned: acceptance thresholds, benchmarks, timelines. Own the execution plan, coordinate capacity and provisioning, run or oversee the tests, and drive the POC to a clear verdict: win the workload, close the gap through product, or qualify out.

  • Lead customer architecture designs. Produce and defend reference architectures spanning compute, networking, storage, connectivity, and scheduler integration (Slurm, Kubernetes). Make support boundaries explicit: what is managed and what is not. Own the design as it evolves after handoff, pulling in engineering domain experts with specific, well-framed questions.

  • Own SLA and reliability engineering. Build and own the canonical methodology for uptime, downtime, and credit calculation. Validate breach events at the technical level, down to the specific Ethernet or InfiniBand failure, and arm CSMs and leadership with defensible numbers. Partner with product to standardize SLA language and structure across 1CC, on-demand, and reserved offerings.

  • Build the tooling that makes accounts measurable. Own the technical data surfaces for customer health end to end: dashboards, telemetry and uptime history views, health scoring, and churn early-warning signals. Scope, build, and drive adoption. Replace “escalate to engineering to answer a basic question” with self-serve data for the whole GTM org.

  • Direct technical escalations and incidents. Serve as the technical lead during high-severity events on your accounts: drive root cause, hold the quality bar on RCAs, coordinate engineering, support, and vendors (NVIDIA, storage, networking), and give account teams a technically accurate narrative. Run proactive maintenance and known-issue communication so customers hear about problems from Lambda first.

  • Drive the technical voice of the customer. Run a structured feature-request pipeline into product with committed triage timelines. Audit the platform hands-on by provisioning as a customer and testing known friction points. Lead product-led POCs (for example, NVIDIA NIM) that open new value for customers.

  • Know the market technology landscape. Track GPU roadmaps, competing clouds and neoclouds, and the evolving training and inference stacks. Brief customers on what is coming and internal teams on where Lambda stands, and let that context shape architecture and expansion recommendations.

You

  • 5+ years in technical account management, solutions engineering or architecture, ML engineering, technical program or product management, or infrastructure engineering with significant customer-facing scope, in cloud, HPC, or AI infrastructure.

  • Hands-on fluency with GPU infrastructure: able to provision, benchmark, and debug across compute, networking (InfiniBand, Ethernet), storage, and schedulers (Slurm, Kubernetes), and to read results critically.

  • Working command of AI/ML workloads (training, fine-tuning, inference) sufficient to map a customer’s stack, identify constraints, and lead technical conversations with their ML and infrastructure engineers.

  • Track record leading structured technical engagements: POCs with defined success criteria, architecture designs, benchmark programs, or high-severity escalations.

  • A builder’s toolkit: scripting, SQL, and dashboarding, with a history of turning operational data into tools other people depend on.

  • Executive-grade communication of deeply technical content, in writing and in the room.

  • Comfort with ambiguity and a track record of building methodology where none exists.

Nice to Have

  • Experience at an AI cloud, neocloud, or hyperscaler serving large-scale GPU or HPC customers.

  • Applied LLM experience (fine-tuning, RAG systems, or inference serving) that mirrors the workloads Lambda customers run.

  • Depth in the NVIDIA ecosystem: NIM and NeMo, Base Command / BCM, the CUDA stack, DGX-class systems.

  • Familiarity with SLA structures, service credits, enterprise contract mechanics, and retention metrics (NRR/GRR).

  • Product management or TPM background with experience converting customer evidence into roadmap decisions.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
375,679 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$115k – $224k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Wayne • Charlotte • Dallas
Python
DevOps
Amazon ECS
AWS
AWS Lambda
CI/CD
Docker
Kubernetes
OpenTelemetry
Platform Engineering
SLI/SLO/SLA
Apply
Support Desk Engineer 10 hours ago
$70k – $126k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Huntsville
Bash
PowerShell
Python
DevOps
GitLab
Kubernetes
OpenShift
Red Hat
Management
Confluence
Jira
Apply
$89k – $190k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Toronto
Scala
Databases
Apache Kafka
Databricks
PostgreSQL
Redis
DevOps
Amazon Kinesis
AWS
Incident Management
Kubernetes
Terraform
Apply
$84k – $157k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Arlington
PowerShell
Python
AI/ML
AI Agents
LLM
Model Context Protocol
DevOps
API Gateway
Azure
Azure DevOps
CI/CD
Configuration Management
Docker
IAM
Kubernetes
Platform Engineering
Terraform
Cybersecurity
Least Privilege
Cryptography
Vault
Apply
$73k – $190k per year (Estimated) • In office • Full-Time • Wellington
PowerShell
SQL
Databases
Cassandra
Databricks
MS SQL
DevOps
Ansible
AWS
Azure
Azure DevOps
CI/CD
GitHub
Jenkins
Kubernetes
SaltStack
Terraform
Windows Server
Apply
$105k – $140k per year • Equity • In office • Full-Time • San Jose
DevOps
AWS Lambda
AWS
Marketing
Zendesk
Apply
$180k – $240k per year • Equity • Remote/Hybrid • Full-Time • 7+ years exp • San Jose • San Francisco
AI/ML
CUDA
CUDA Toolkit
Fine-tuning
InfiniBand
NCCL
PyTorch
DevOps
AWS Lambda
HPC
Kubernetes
SLURM
AWS
Apply
$89k – $119k per year • Equity • In office • Full-Time • Atlanta
DevOps
AWS Lambda
AWS
Marketing
Zendesk
Apply
$380k – $445k per year • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • San Francisco
Go
Python
SQL
Databases
DynamoDB
Presto
DevOps
AWS Lambda
AWS
Amazon S3
Apply
$297k – $440k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco • San Jose • Bellevue
AI/ML
InfiniBand
DevOps
AWS Lambda
Incident Management
AWS
HPC
Apply
$133k – $338k per year • Remote/Hybrid • Full-Time • 8+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
Knowledge Graph
DevOps
AWS
Azure
GCP
Platform Engineering
SLI/SLO/SLA
Apply
$208k – $332k per year • Remote/Hybrid • Full-Time • 9+ years exp • Bachelor's Degree • San Francisco
Apply
$197k – $314k per year • Remote/Hybrid • Full-Time • 8+ years exp • PhD • San Francisco
Apex
Java
Apex
MuleSoft
Java
Spring Boot
AI/ML
AI Agents
Model Context Protocol
RAG
Agentforce
DevOps
Docker
Kubernetes
Marketing
Salesforce
Apply
$220k – $413k per year (Estimated) • Remote/Hybrid • 12+ years exp • San Francisco
Java
Python
Ruby
Databases
DynamoDB
ElasticSearch
RocksDB
AI/ML
AI Agents
LLM Guardrails
Apply
$149k – $260k per year • In office • Full-Time • 6+ years exp • PhD • San Francisco
Go
JavaScript
TypeScript
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Prompt Engineering
Agentforce
OpenAI Codex
Frontend
React.js
DevOps
Kubernetes
GitHub
Marketing
Salesforce
Apply
See all jobs
This is one of many
375,679 more open roles from verified company boards, updated every day.