368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$257k – $523k per year (Estimated)
Location
Remote (United States)
Seniority
Staff · 5+ years exp
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As the Hardware Deployment Engineer Lead, you will own the end-to-end bring-up of GPU compute hardware across the world's largest AI training clusters. You will build and lead a dedicated in-house hardware deployment team responsible for L11 integration, hardware bring-up, and post-L11 repair of GB300-class systems across multiple data halls concurrently. Your team's throughput directly determines how fast xAI can compute online - this is one of the most critical path activities in the company. You will set the deployment playbook, hold hardware vendors accountable to SLAs, and institutionalize processes so that cluster deployment is limited only by hardware supply and power, never by deployment velocity. The position is based in Memphis, TN.

RESPONSIBILITIES:

  • Lead, hire, and develop a dedicated hardware deployment team (deployment engineers, deployment technicians, and repair technicians) with full ownership of team structure and staffing.
  • Own L11 rack integration and compute hardware bring-up across multiple data halls concurrently, from delivery dock to healthy production handoff.
  • Drive aggressive bring-up timelines: achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
  • Own post-L11 hardware health: run systematic health pushes to sustain greater than 98% node availability prior to turnover to operations.
  • Internalize non-RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times.
  • Develop and enforce vendor SLAs for OEM and supplier responsibilities; prevent accumulation of unrepaired hardware ("bone piles") and repair backlogs before turnover to operations.
  • Perform root cause analysis of hardware failures discovered during L11 and drive corrective actions with vendors and internal engineering teams.
  • Partner with site operations on hardware debugging and repair, and train site operations teams to support future data center deployments.
  • Build, document, and continuously improve deployment processes, tooling, and training so bring-up capability scales across sites and future hardware generations.

BASIC QUALIFICATIONS:

  • 5+ years of hands-on experience deploying, integrating, or repairing compute/server hardware at data center scale.
  • Direct experience with L11 (rack-level) integration and bring-up of GPU or accelerator-based systems.
  • Demonstrated experience leading technician or engineering teams in a fast-paced deployment, manufacturing, or data center environment.
  • Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems.
  • Willingness to work on-site in Memphis, TN, including extended hours and weekends during critical bring-up phases.

PREFERRED SKILLS AND EXPERIENCE:

  • Experience with NVIDIA GB200/GB300 NVL72 or similar rack-scale liquid-cooled GPU systems.
  • Experience standing up a new team or function, including hiring, training, and process development from scratch.
  • Experience managing OEM/ODM vendor relationships (e.g., Dell, Supermicro), including SLA definition and enforcement.
  • Experience with hardware failure analysis, RMA processes, and component-level repair strategies at fleet scale.
  • Experience with data center automation, burn-in/validation tooling, and hardware health telemetry.
  • Track record of driving step-change improvements in deployment velocity or cost (e.g., insourcing work previously performed by OEMs).

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$117k – $239k per year (Estimated) • In office • 3+ years exp • Memphis
Apply
$440k per year • In office • Palo Alto
C++
AI/ML
CUDA Toolkit
CUDA
Frontend
Sass
Apply
$167k – $358k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Python
SQL
AI/ML
BERT
InfiniBand
DevOps
HPC
Apply
$440k per year • In office • Palo Alto
C++
Python
Rust
AI/ML
LLM
Reinforcement Learning
Apply
$600k per year • In office • Palo Alto
AI/ML
Reinforcement Learning
RLHF
DPO
Post-training
Apply
$183k – $275k per year • Equity • In office • Full-Time • 10+ years exp • High School Diploma • Fort Worth • Minneapolis • Irvine • Jacksonville • Boston
DevOps
IAM
Cybersecurity
Least Privilege
Zero Trust
Apply
$209k – $454k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$209k – $454k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$117k – $239k per year (Estimated) • In office • 3+ years exp • Memphis
Apply
$167k – $358k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Python
SQL
AI/ML
BERT
InfiniBand
DevOps
HPC
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.