417,786open jobs
14,314companies
60,729added this week
Browse all
Salary
$137k – $287k per year
Location
Remote/Hybrid (Fremont, United States)
Overview
Company
Impact
Profile match
Lam Research is an American semiconductor equipment company founded in 1980 that specialises in the deposition and etch steps used to build transistors and memory cells. Its tools are central to three-dimensional NAND manufacturing, where hundreds of layers must be etched through in high aspect ratio channels, and to the atomic layer deposition and selective etch processes required at advanced logic and DRAM nodes. Headquartered in Fremont, California and listed on Nasdaq, it also runs a large installed base business supplying spares, upgrades and service, which smooths the sharp cycles in equipment purchasing.

The group you’ll be a part of

You will join the Reliability Engineering team within Infrastructure Platform Engineering. The group keeps Lam's global infrastructure estate available and recoverable across Azure, AWS, GCP, compute, storage, network, and high-performance computing, supporting engineering and operations teams in the US, Japan, Singapore, Malaysia, India, and Korea.

The impact you’ll make

As Senior Manager of Reliability Engineering & AIOps, you lead the team that keeps critical infrastructure running and proves it is ready for the next failure. In this role, you will directly contribute to the availability of the systems Lam's engineering, manufacturing, and business teams depend on every day, and you will build the automation that makes outages rare, short, and unremarkable.

What you’ll do

  • Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
  • Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams.
  • Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide.
  • Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise.
  • Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up.
  • Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets.
  • Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work.
  • Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations.
  • This is a full-time role on a standard schedule, with participation in a global on-call rotation

Who we’re looking for

  • Bachelor's degree in Computer Science, Engineering, or a related field with 10 years of related experience; or a Master's degree with 8 years of experience; or equivalent experience.
  • Experience leading or mentoring a reliability or operations team and setting technical direction.
  • Proven incident command on major outages, plus ownership of a postmortem process.
  • Strong background in disaster recovery planning across Azure, AWS, GCP, and core infrastructure platforms, including restore validation, failover testing, recovery-objective definition, and corrective action tracking after DR exercises or production incidents.
  • Hands-on ownership of an incident management and paging platform at scale, such as PagerDuty.
  • Experience with capacity planning, performance trending, utilization analysis, and infrastructure demand forecasting for globally distributed production environments across Azure, AWS, GCP, and on-premises platforms.
  • Track record of defining and defending service level objectives and error budgets in production.
  • Working depth in observability tooling (Prometheus, Grafana, Loki, Tempo or equivalent), infrastructure as code (Terraform), and Python or Go.
  • Practical experience applying AI-assisted engineering and operations tools such as Microsoft Copilot, Cursor, GitHub Copilot, or enterprise LLM platforms to improve troubleshooting, automation, documentation, and engineering productivity across Azure, AWS, GCP, and hybrid infrastructure, with clear guardrails for security, privacy, auditability, and production safety.

Preferred qualifications

  • Experience running global, follow-the-sun operations across multiple regions and time zones.
  • Capacity and performance engineering at multi-region scale, including Azure, AWS, GCP, high-performance computing, large storage estates, hybrid cloud infrastructure, and proactive capacity governance.
  • Policy as code, progressive delivery, and chaos engineering in practice.
  • Experience building or operating AI and agent-assisted automation in operations, with a clear view of its failure modes.
  • Experience designing or operating AI Ops capabilities across Azure, AWS, GCP, and hybrid environments, including LLM-grounded knowledge bases, agent-assisted incident workflows, prompt and evaluation practices, and supervised automation that can recommend or propose operational changes before execution.

Our commitment

We believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results.

Lam Research ("Lam" or the "Company") is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non-discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company's intention to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees.

Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work remotely and fall into two categories - On-site Flex and Virtual Flex. ‘On-site Flex’ you’ll work 3+ days per week on-site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on-site at a Lam or customer/supplier location, and remotely the rest of the time.

#LI-DM1

Salary

CA San Francisco Bay Area Salary Range for this position: $137,000.00 - $287,000.00.

The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors.

Our Perks and Benefits

At Lam, our people make amazing things possible. That’s why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
417,786 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Fremont
Quality Engineer 3 1 day ago
$77k – $118k per year • Remote • 7+ years exp
JavaScript
TypeScript
AI/ML
Copilot
Cursor
Claude
Prompt Engineering
AI Agents
RAG
DevOps
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
Git
GitHub
Management
Microsoft Teams
QA
JMeter
Playwright
k6
Apply
Remote/Hybrid • 5+ years exp • Bachelor's Degree
Python
JavaScript
PHP
TypeScript
SQL
Databases
PostgreSQL
Pinecone
FAISS
AI/ML
LangChain
LlamaIndex
Embeddings
Prompt Engineering
AI Agents
LLM
RAG
OpenAI
Hugging Face
OCR
DevOps
GCP
Azure DevOps
GitHub Actions
Datadog
Prometheus
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Vector
GitHub
Chips/EDA
PoC Library
Management
Outlook
Apply
Remote/Hybrid • 10+ years exp • Bachelor's Degree
JavaScript
PHP
TypeScript
PHP
WordPress
AI/ML
Copilot
Claude
Claude Code
OpenAI Codex
Frontend
GraphQL
Next.js
React.js
DevOps
CI/CD
Git
Bitbucket
GitHub
Analytics
A/B Testing
Apply
Remote/Hybrid • 5+ years exp • Bachelor's Degree
Python
PHP
SQL
AI/ML
Copilot
ChatGPT
Prompt Engineering
AI Agents
Analytics
Power BI
Management
Smartsheet
Jira
Apply
In office • 6+ years exp • Bachelor's Degree
Python
AI/ML
Copilot
Claude
ChatGPT
Hallucination
Analytics
Tableau
Power BI
Management
Power Automate
Apply
Process Engineer 3 14 days ago
Remote/Hybrid • Master's Degree
Python
MATLAB
Apply
In office • Bachelor's Degree
Apply
Remote/Hybrid • Bachelor's Degree
SQL
Analytics
Tableau
Apply
$56k – $94k per year • In office • Contractor • Fremont
Apply
$56k – $94k per year • In office • Contractor • Fremont
Apply
$64k – $74k per year • In office • Contractor • Fremont
Apply
$64k – $74k per year • In office • Contractor • Fremont
Apply
$56k per year • In office • Contractor • Fremont
Apply
See all jobs
This is one of many
417,786 more open roles from verified company boards, updated every day.