665,767open jobs
38,907companies
100,291added this week
Browse all
Salary
$230k – $469k per year (Estimated)
Location
In office (Memphis)
Seniority
Architect · 10+ years exp
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As the Director of Site Operations, you’ll own node and rack uptime for SpaceXAI's AI supercompute cluster-the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You’ll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We’re looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow.

RESPONSIBILITIES:

  • Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites running 24/7, accountable for customer Service Level Agreements and consistently exceptional uptime on SpaceXAI's supercompute cluster.
  • Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift supervisors, and technicians across four 24/7 shifts, building a culture of excellence and accountability at every layer of the org.
  • Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through command-line and physical intervention, driving mean time to repair to the feasible minimum.
  • Partner Across Functions: Coordinate with facilities operations to limit downtime from power and cooling faults and proactive maintenance; with network engineering on cluster upgrades; and with tenant representatives on node remediation and planned and unplanned downtime.
  • Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires.
  • Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses for node, rack, and cluster issues, and site-wide reliability procedures and fault documentation.
  • Run Data-Driven Improvement: Lead continual improvement and efficiency initiatives, using operational data to balance team resources and raise uptime, repair time, and SLA performance across sites.
  • Command Incidents at Scale: Set the standard for incident response during cluster-impacting events, providing clear direction, fast recovery, and tight communication with internal and external partners.
  • Scale Operations: Standardize best practices across sites and grow the organization in step with cluster expansion, keeping operations consistent as SpaceXAI's footprint scales.

BASIC QUALIFICATIONS:

  • Bachelor’s degree and 7+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams OR 10+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams.

PREFERRED SKILLS AND EXPERIENCE:

  • Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced, high-responsibility settings.
  • Deep expertise in server hardware, cluster reliability, and data center technologies, from deployment through lifecycle management.
  • Experience supporting compute-heavy environments like AI, machine learning, or high-performance computing at scale.
  • A track record of owning uptime, Service Level Agreements, or reliability metrics for large compute clusters.
  • Experience leading site reliability engineering or equivalent reliability-focused teams, including root cause analysis and procedure ownership.
  • Strong analytical skills and the ability to explain technical concepts clearly to diverse audiences, from technicians to executive and tenant partners.
  • A history of partnering with vendors at scale, driving mean time to repair down, and scaling operations across multiple sites.
  • Familiarity with tooling and automation (e.g., Jira, Python, Bash) used to monitor cluster health and improve team efficiency.
  • Enthusiasm for SpaceXAI's mission to accelerate human discovery and unravel the universe.
  • Ability to thrive in a dynamic, mission-focused environment with on-call ownership of cluster-impacting events.

ADDITIONAL REQUIREMENTS:

  • Willingness to travel frequently to data center locations to support operations across sites.
  • Physical capability to handle data center tasks, including lifting up to 50 lbs unassisted, standing for long periods, and occasional ladder use.
  • Must be willing to work extended hours and/or weekends as needed.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
665,767 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$110k – $165k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • Scottsdale • Jersey City
SQL
PowerShell
DevOps
Dynatrace
AWS
SLI/SLO/SLA
Analytics
Power BI
Microsoft Excel
Management
Confluence
Jira
ServiceNow
Power Automate
Apply
System Engineer 1 day ago
$50k – $147k per year (Estimated) • In office • Ghent
Python
Python
Django
Celery
Databases
RabbitMQ
Apache Kafka
DevOps
Terraform
Ansible
Loki
Prometheus
CI/CD
AWS
Docker
Grafana
Configuration Management
SLI/SLO/SLA
Amazon S3
IAM
Cybersecurity
Least Privilege
QA
Sentry
Apply
$178k – $329k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Park
Apex
Apex
MuleSoft
DevOps
Shift-Left
SLI/SLO/SLA
Cybersecurity
Shift-Left Security
Analytics
Informatica
Management
Agile
Apply
$38k – $91k per year (Estimated) • In office • Full-Time • High School Diploma • United States
DevOps
SLI/SLO/SLA
Apply
$103k – $198k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Antonio
DevOps
Incident Management
SLI/SLO/SLA
Robotics
Path Planning
Management
ServiceNow
Apply
$160k per year • In office • Palo Alto
JavaScript
Frontend
Next.js
React.js
Marketing
Salesforce
Apply
$84k – $199k per year (Estimated) • In office • 5+ years exp • High School Diploma • Memphis
Apply
$250k per year • In office • Full-Time • 8+ years exp • Palo Alto
Apply
$61k – $115k per year (Estimated) • In office • 1+ year exp • High School Diploma • Memphis
Apply
Network Engineer 2 days ago
$105k – $274k per year (Estimated) • In office • Dublin
Python
AI/ML
NCCL
DevOps
Terraform
Ansible
HPC
Apply
$53k – $105k per year (Estimated) • Remote/Hybrid • 4+ years exp • High School Diploma • Memphis
DevOps
Splunk
Management
Confluence
Jira
ServiceNow
Apply
$38k – $42k per year • Equity • Remote • Full-Time • Nashville • Dallas • Savannah • Orlando • Tampa
Apply
$69k – $136k per year (Estimated) • In office • Full-Time • Memphis
Design
AutoCAD
Management
Microsoft Project
Apply
$131k – $218k per year • Remote/Hybrid • Full-Time • 8+ years exp • High School Diploma • Tempe • Raleigh • Memphis • Scottsdale • Saint Louis
Python
SQL
PowerShell
Analytics
Tableau
ETL/ELT
Apply
$44k – $105k per year (Estimated) • In office • Full-Time • Memphis
Apply
See all jobs
This is one of many
665,767 more open roles from verified company boards, updated every day.