410,083open jobs
14,230companies
73,879added this week
Browse all
Location
In office (Memphis)
Seniority
Senior · 5+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

We are seeking an exceptional Manager, Operations to lead facilities operations and power generation for SpaceXAI's hyperscale AI compute facilities. This role will own the day-to-day and long-term performance of mission-critical data center operations, including power generation, power distribution, cooling, mechanical, electrical, and environmental systems, while also directing the fiber teams responsible for high-capacity networking and connectivity that support our supercomputing clusters.

You will build and lead high-performing operations, power generation, and fiber teams, drive relentless reliability and efficiency, and ensure seamless 24/7 uptime for the infrastructure powering SpaceXAI's AI training at unprecedented scale. This high-impact position requires deep expertise in data center or hyperscale operations (including power generation), strong leadership in fast-paced environments, and the ability to deliver world-class performance under aggressive growth timelines. This is a full-time, primarily onsite role with significant travel to sites and vendor locations.

RESPONSIBILITIES:

  • Lead and scale the facilities operations and power generation teams responsible for the reliable operation, maintenance, monitoring, and optimization of critical infrastructure including on-site power generation assets, electrical systems, mechanical/HVAC, liquid cooling, power distribution, UPS, generators, and building management systems.
  • Direct the fiber teams overseeing the design, deployment, maintenance, and expansion of high-speed fiber optic networks, dark fiber, and connectivity infrastructure supporting AI compute clusters and data center interconnects.
  • Own key performance metrics such as uptime (targeting 99.999%+), mean time to detect/repair (MTTD/MTTR), power usage effectiveness (PUE), water usage effectiveness (WUE), power generation efficiency, and overall infrastructure availability.
  • Develop and enforce standard operating procedures (SOPs), preventive maintenance programs, incident response protocols, and continuous improvement processes for both facilities and power generation assets to minimize downtime and maximize efficiency.
  • Build, mentor, and grow multidisciplinary teams of operations technicians, power generation engineers and controls specialists while fostering a culture of ownership, safety, and excellence.
  • Partner closely with engineering, construction, procurement, and AI hardware teams to support new facility builds, expansions, commissioning, power integration, and smooth handovers from project to operations.
  • Manage operational budgets, vendor relationships (maintenance contractors, fiber providers, power generation OEMs, fuel suppliers), spare parts inventory, and risk mitigation strategies in a high-velocity environment.
  • Drive innovation in operational practices, automation, predictive maintenance, power generation optimization, and sustainability initiatives to support the extreme power and cooling demands of next-generation AI systems.
  • Provide regular performance reporting, root cause analyses, lessons learned, and strategic recommendations to senior leadership.

BASIC QUALIFICATIONS:

  • 5+ years of progressive experience in data center facilities operations, power generation operations, hyperscale infrastructure management, or mission-critical industrial operations, with at least 2+ years in a management or supervisor role.
  • Proven track record leading large-scale operations teams supporting high-density compute environments with significant on-site or dedicated power generation (AI, HPC, or hyperscaler data centers strongly preferred).
  • Strong experience managing fiber optic networks, dark fiber deployments, or high-bandwidth connectivity infrastructure in large-scale technical environments.
  • Deep knowledge of power generation systems (gas turbines, reciprocating engines, cogeneration, etc.), MEP (mechanical, electrical, plumbing) systems, BMS/SCADA, liquid cooling, power redundancy topologies, and 24/7 operations best practices.
  • Demonstrated success delivering high reliability, rapid incident resolution, and operational excellence under aggressive scaling timelines.
  • Hands-on leadership style with the ability to roll up sleeves while effectively managing teams, budgets, and cross-functional stakeholders.
  • Proficiency with operations tools, CMMS (computerized maintenance management systems), monitoring platforms, and data-driven decision making.

PREFERRED SKILLS AND EXPERIENCE:

  • Direct background in AI or hyperscale data center operations, including liquid cooling systems, high-power GPU/accelerator environments, and on-site power generation.
  • Experience building or scaling fiber infrastructure for low-latency, high-bandwidth interconnects between compute clusters or sites.
  • Familiarity with Uptime Institute Tier standards, ASHRAE guidelines, power generation standards (e.g., IEEE, NFPA), OSHA/EPA compliance, and sustainability practices in critical facilities.
  • Bachelor’s or Master’s degree in Electrical, Mechanical Engineering, Power Systems, Facilities Management, or related field; relevant certifications (CDCP, CDCS, or equivalent) a plus.
  • Track record of implementing automation, predictive analytics, or process improvements that significantly enhanced operational performance and power reliability.

ADDITIONAL REQUIREMENTS:

  • Willingness to be primarily onsite at key facilities (e.g., Memphis region) with on-call responsibilities and travel to other sites as needed.
  • Ability to work in industrial/data center environments and lead teams during high-pressure phases.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
410,083 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$51k per year • Remote • Full-Time
C#
C++
Go
Java
Kotlin
Rust
Solidity
TypeScript
JavaScript
Java
Spring Boot
Kotlin
Ktor
Databases
Amazon Aurora
Frontend
ESLint
MSW
React.js
Rollup
Storybook
Vite
Vue.js
DevOps
Amazon Kinesis
AWS
AWS Lambda
Azure
CI/CD
CloudFormation
GCP
Terraform
Web3
DeFi
Layer 2
Rollup
Smart Contracts
Zk-SNARKs
Zk-STARKs
QA
Vitest
Apply
Remote/Hybrid • PhD
Python
SQL
DevOps
HPC
Apply
In office • PhD
Databases
Databricks
Snowflake
DevOps
AWS
HPC
Robotics
Localization
Apply
$56k – $79k per year • In office • 4+ years exp • PhD
AI/ML
Hugging Face
Multimodal AI
PyTorch
DevOps
Docker
HPC
Apply
$100k – $150k per year • Remote • 6+ years exp • Bachelor's Degree
C++
Go
Python
AI/ML
InfiniBand
Ray
DevOps
CI/CD
FinOps
HPC
Kubernetes
SLURM
Apply
Equity • In office • 3+ years exp • Associate's Degree • Memphis
Apply
$135k per year • Remote/Hybrid • Full-Time • 2+ years exp • Palo Alto
Apply
$135k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
Management
Google Workspace
Notion
Slack
Apply
$220k per year • In office • Full-Time • 12+ years exp • New York
Apply
In office • 8+ years exp • Bachelor's Degree • Memphis
AI/ML
Edge AI
Apply
Equity • In office • 3+ years exp • Associate's Degree • Memphis
Apply
In office • 8+ years exp • Bachelor's Degree • Memphis
AI/ML
Edge AI
Apply
In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$117k – $239k per year (Estimated) • In office • 3+ years exp • Memphis
Apply
In office • Contractor • 3+ years exp • High School Diploma • Memphis
Management
Outlook
Apply
See all jobs
This is one of many
410,083 more open roles from verified company boards, updated every day.