380,026open jobs
9,934companies
47,963added this week
Browse all
Salary
$204k – $396k per year (Estimated)
Location
In office (Memphis)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.

RESPONSIBILITIES:

  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
  • Run blameless postmortems and drive corrective actions to closed, not filed.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
  • Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.

BASIC QUALIFICATIONS:

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.

PREFERRED SKILLS AND EXPERIENCE:

  • Experience in AI/ML infrastructure or supercomputing environments.
  • Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
  • Experience running game days, dependency mapping, and closed-loop corrective action programs.
  • Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
380,026 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$200k – $288k per year • Remote/Hybrid • Full-Time • 5+ years exp • Menlo Park • Bellevue
C++
Go
Java
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
gRPC
IAM
Kubernetes
Pulumi
Terraform
Cybersecurity
Threat Modeling
Apply
$105k – $130k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Charlotte • Greensboro • Richmond • Atlanta • Raleigh
PowerShell
Python
DevOps
CI/CD
Platform Engineering
Cybersecurity
SBOM
SLSA
Threat Modeling
Apply
$140k – $180k per year • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • Charlotte • Greensboro • Richmond • Atlanta • Raleigh
PowerShell
Python
DevOps
CI/CD
Platform Engineering
Cybersecurity
SBOM
SLSA
Threat Modeling
Apply
$160k – $200k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Charlotte • Greensboro • Richmond • Atlanta • Raleigh
PowerShell
Python
DevOps
CI/CD
Platform Engineering
Cybersecurity
SBOM
SLSA
Threat Modeling
Apply
Data and AI Engineer 4 hours ago
$56k – $141k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Toronto
Python
SQL
Databases
Snowflake
DevOps
CI/CD
GitHub
Analytics
Power BI
Tableau
Management
Confluence
Apply
$150k – $316k per year (Estimated) • In office • 3+ years exp • Memphis
Bash
PowerShell
Python
DevOps
Ansible
Configuration Management
GitOps
Puppet
Terraform
IoT
MQTT
OPC UA
Apply
$167k – $358k per year (Estimated) • In office • Contractor • 5+ years exp • Bachelor's Degree • Memphis
Bash
PowerShell
Python
AI/ML
InfiniBand
NCCL
DevOps
Ansible
GitOps
HPC
Terraform
Apply
$142k – $322k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Dublin
AI/ML
EU AI Act
DevOps
AWS
Azure
CI/CD
GCP
IAM
Cybersecurity
GDPR
ISO 27001
SOC 2
Apply
$209k – $454k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$209k – $454k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$67k – $135k per year (Estimated) • In office • Full-Time • 5+ years exp • Memphis
DevOps
Incident Management
Apply
$140k – $200k per year • In office • PhD • Memphis
Swift
AI/ML
Text-to-Speech
Mobile
Fastlane
SwiftUI
DevOps
CI/CD
Git
Vercel
Management
Google Docs
Stripe
Marketing
LinkedIn
Apply
$140k – $200k per year • In office • Memphis
Java
Kotlin
Mobile
Kotlin Multiplatform
DevOps
GCP
Marketing
LinkedIn
Apply
$140k – $200k per year • In office • Internship • PhD • Memphis
C#
C++
C#
.NET
AI/ML
Text-to-Speech
DevOps
CI/CD
Vercel
Management
Google Docs
Stripe
Marketing
LinkedIn
Apply
$65k – $133k per year (Estimated) • In office • Full-Time • Memphis
Apply
See all jobs
This is one of many
380,026 more open roles from verified company boards, updated every day.