730,610open jobs
43,655companies
102,906added this week
Browse all
Salary
$108k – $216k per year
Location
In office (Irvine)
Seniority
Senior · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026. Walmart scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Walmart is a multinational retail corporation operating a vast global chain of hypermarkets, discount department stores, and grocery stores. Founded in 1962 by Sam Walton and headquartered in Bentonville, Arkansas, the company is one of the world's largest businesses by revenue and a dominant force in brick-and-mortar and online retail. Today, it integrates its physical store network with a growing e-commerce ecosystem, offering low prices on groceries, general merchandise, and digital services to millions of customers daily.

What you'll do...

Position: Senior Site Reliability Engineer

Job Location: 39 Tesla, Irvine, CA 92618

Duties: Drive the design and evolution of monitoring and observability frameworks that enable proactive detection, root cause analysis, and rapid resolution of customer-impacting incidents. Lead the development and integration of automation tools to streamline operational workflows, reduce toil, and enhance the reliability of customer service platforms. Participate in on-call rotations, applying deep technical expertise to swiftly diagnose and mitigate production issues, ensuring high availability and minimal disruption to customer support experiences. Collaborate with engineering teams to embed reliability into the software development lifecycle, championing a culture of shared ownership and “you build it, you run it.” Define and manage SLIs, SLOs, and SLAs to align service reliability with business expectations and continuously improve system performance. Apply proven reliability patterns and practices, leveraging hands-on experience to architect resilient systems that scale with customer demand. Lead post-incident reviews and blameless retrospectives, identifying systemic improvements and fostering a culture of continuous learning and operational excellence. Analyze system performance and advocate for cost-effective optimizations, balancing infrastructure efficiency with world-class service reliability. Identify repetitive and routine tasks in (Continuous Integration/Continuous Delivery) CI/CD, testing, or any other process that can be automated. Implement telemetry features as required under guidance. Apply security policy requirements to component/module during code development/configuration. Detect and document defects, bugs, and errors for assigned component/module and conduct analysis to determine the sources under guidance. Troubleshoot performance and availability bottlenecks for assigned application under guidance. Work with business partners to identify and document critical applications. Interpret and follow procedures in contingency plans. Explain the contingency and disaster recovery plans for assigned environment. Execute established procedures necessary to continue operations in an emergency. Participate in the design of a minimum operating environment for a computer-based facility. Utilize established criteria (for example, probability of failure, frequency of failure) to measure site reliability. Monitor site reliability conditions and new reliability requirements. Assist in the design and development of a reliability program plan for a specific site environment. Apply appropriate tools, services, or applications for reliability prediction and other site improvements. Research and assess various reliability models for different site environments. Suggest metrics to monitor software or system performance. Monitor current performance data to ensure compliance with defined SLOs for multiple applications/systems. Determine thresholds for monitoring metrics and triggers alerts based on thresholds. Help with specific procedures to proactively check the health of applications and infrastructure, including a variety of operating systems, hardware, and software. Make recommendations regarding situational awareness and alerting. Make recommendations regarding instrumentation gaps and alerting logic, including a variety of operating systems, hardware, and software.

Minimum education and experience required: Master's degree or the equivalent in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, Electrical Engineering, Information Systems Security, or related area and 1 year of experience in site reliability engineering, site and system administration, infrastructure management, or related area; OR Bachelor's degree or the equivalent in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, Electrical Engineering, Information Systems Security, or related area and 3 years of experience in site reliability engineering, site and system administration, infrastructure management, or related area.

Skills required: Experience building, supporting, and maintaining Databricks Platforms and Databricks Infrastructure. Experience creating and managing AWS services, including s3 buckets, IAM roles and policies, VPC, Subnets, and VPCE Endpoints. Experience providing production support for applications running on data bricks and addressing all infrastructure requests from Data Engineering and Data Science teams. Experience working with Databricks and AWS service teams to understand new features, implement updates, and resolve vendor related issues. Experience investigating, troubleshooting, and optimizing Databricks spark jobs and providing resolutions. Experience supporting applications running on Elastic Kubernetes Service clusters and Kafka clusters on AWS and GCP. Experience building Infrastructure as Code using Terraform to automate the creation and management of Databricks resources. Experience using PagerDuty, Monte Carlo, and Cloud Watch for alerting, incident response, and monitoring. Experience using Airflow for workflow orchestration and job scheduling. Experience with migrations from AWS to GCP, setting up infrastructure, and enabling data movement from AWS s3 to Google Cloud Storage. Experience working with offshore, Central Ops, and QA teams to improve system reliability, reduce operational overhead, enhance observability, ensure safe deployments, reduce incident response time, and align engineering practices with business goals. Employer will accept any amount of experience with the required skills.

Salary Range: $108,000/year to $216,000/year. Additional compensation includes annual or quarterly performance incentives.

Benefits: At Walmart, we offer competitive pay as well as performance-based incentive awards and other great benefits for a happier mind, body, and wallet. Health benefits include medical, vision and dental coverage. Financial benefits include 401(k), stock purchase and company-paid life insurance. Paid time off benefits include PTO (including sick leave), parental leave, family care leave, bereavement, jury duty and voting. Other benefits include short-term and long-term disability, education assistance with 100% company paid college degrees, company discounts, military service pay, adoption expense reimbursement, and more.

Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to a specific plan or program terms. For information about benefits and eligibility, see One.Walmart.com.

Wal-Mart is an Equal Opportunity Employer.

#LI-DNI #LI-DNP

Walmart and its subsidiaries are committed to maintaining a drug-free workplace and has a no tolerance policy regarding the use of illegal drugs and alcohol on the job. This policy applies to all employees and aims to create a safe and productive work environment.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
730,610 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Irvine
Senior Data Engineer 5 hours ago
$25k – $53k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Python
pySpark
Databases
PostgreSQL
Databricks
Apache Iceberg
Delta Lake
AI/ML
Spark
Model Context Protocol
MLFlow
Embeddings
Function Calling
LLM
RAG
Reranking
Anomaly Detection
LLMOps
Feature Store
Red Teaming
LLM Guardrails
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
Rest API
Splunk
Terraform
GCP
Datadog
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Amazon EKS
AWS Lambda
Amazon EC2
FinOps
Amazon S3
IAM
Amazon CloudWatch
Cybersecurity
Least Privilege
Analytics
Dimensional Modeling
Collibra
IoT
Matter
Management
Jira
Agile
Apply
$20k – $57k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Apache Kafka
Amazon Redshift
AI/ML
Spark
Airflow
dbt
DevOps
Terraform
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
Git
AWS
AWS Lambda
Amazon S3
IAM
Analytics
ETL/ELT
Apply
$36k – $81k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Hyderabad
Python
Go
Java
TypeScript
PowerShell
Bash
Databases
Databricks
AI/ML
AWS Bedrock
Amazon SageMaker
DevOps
Splunk
Terraform
OpenTelemetry
AWS CDK
CloudFormation
Datadog
Prometheus
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Chaos Engineering
Amazon EKS
AWS Lambda
Amazon EC2
Progressive Delivery
FinOps
Amazon S3
Amazon CloudWatch
Linux
Cybersecurity
Least Privilege
Management
Jira
Agile
Apply
$60k – $125k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Python
JavaScript
Node JS
AI/ML
Model Context Protocol
AI Agents
LLM
RAG
ISO 42001
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Kubernetes
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Threat Modeling
Apply
$11k – $26k per year (Estimated) • In office • Full-Time • Manila
Python
JavaScript
TypeScript
Node JS
Python
Flask
FastAPI
Node JS
Axios
Frontend
Zustand
Redux
Webpack
GraphQL
Tailwind CSS
React.js
Vite
Material UI
Sass
DevOps
Rest API
Terraform
GitHub Actions
CI/CD
Git
AWS
Docker
GitHub
Amazon S3
Design
Figma
QA
Jest
Apply
$97k – $180k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Bentonville
Apply
Staff Product Manager 5 hours ago
$117k – $220k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bentonville
SQL
Analytics
Tableau
Power BI
A/B Testing
Management
Agile
Scrum
Kanban
Apply
Software Engineer III 5 hours ago
$90k – $180k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Bentonville
Python
JavaScript
Java
TypeScript
C#
Java
Maven
Spring Boot
Spring MVC
Spring Data JPA
Spring Security
Spring Cloud
Gradle
Databases
Databricks
ElasticSearch
AI/ML
AI Agents
Ollama
Hugging Face
OpenAI Agents SDK
Multi-Agent Systems
Frontend
GraphQL
Angular
React.js
DevOps
Splunk
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Analytics
Azure Data Factory
Management
Monday.com
Agile
QA
Selenium
Apply
$90k – $180k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Bentonville
DevOps
Azure
Design
Figma
Adobe XD
FigJam
Management
Trello
Miro
Jira
Agile
Scrum
Apply
$90k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bentonville
Python
JavaScript
Java
SQL
C#
Java
Hibernate
C#
.NET
Databases
Oracle
Mobile
MVP
Dependency Injection
DevOps
Rest API
Azure
CI/CD
AWS
Management
Monday.com
Jira
Agile
QA
Selenium
Apply
$38k per year • In office • PhD • Irvine
Mobile
Adjust
DevOps
Rest API
Apply
$266k – $394k per year • Equity • Remote (likely United States) • Full-Time • 10+ years exp • Bachelor's Degree • Phoenix • Irvine • Glendale • San Jose • Santa Monica
Apply
$86k per year • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Seattle • Los Angeles • San Jose • San Francisco • San Diego
Apply
$129k – $207k per year • Equity • In office • Full-Time • 12+ years exp • Bachelor's Degree • Irvine
Apply
Operator II 1 day ago
$32k – $64k per year • In office • Full-Time • 2+ years exp • High School Diploma • Irvine
AI/ML
Ray
Apply
See all jobs
This is one of many
730,610 more open roles from verified company boards, updated every day.