368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$124k – $253k per year (Estimated)
Location
In office (Memphis)
Seniority
Junior · 2+ years exp
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Site Reliability Engineer focused on Hardware, you will serve as an expert focused on firmware, hardware specifications, vendor relations, and failure analysis. You will proactively identify and resolve hardware issues, manage RMA processes, and stay ahead of emerging hardware technologies to support SpaceXAI's data center operations. This role demands deep technical expertise in hardware diagnostics, and forward-looking hardware evaluation.

RESPONSIBILITIES:

  • Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's data center environment. Run security scanning and CVE / vulnerability analysis on firmware and related components. Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor.
  • Investigate and diagnose hardware failures, including "grey failures" (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis.
  • Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions.
  • Collaborate with Data Center Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time.
  • Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime.
  • Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base.
  • Participate in on-call rotations and incident response for hardware-related issues in the Memphis data center  

BASIC QUALIFICATIONS:

  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
  • Proven expertise in firmware analysis, hardware specifications review, and release validation.
  • Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.
  • Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.
  • Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including operations technicians.

PREFERRED SKILLS AND EXPERIENCE:

  • Experience in AI/ML infrastructure or supercomputing environments.
  • Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.
  • Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$77k – $215k per year (Estimated) • In office • Contractor • 3+ years exp • Singapore
JavaScript
Python
TypeScript
Python
Django
FastAPI
Databases
PostgreSQL
Redis
AI/ML
AI Agents
LangGraph
LLM
Prompt Engineering
LangChain
Edge AI
Frontend
Angular
React.js
Vue.js
Apply
$15k – $39k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Bengaluru
COBOL
Java
TypeScript
JavaScript
Java
Hibernate
Databases
Apache Solr
AI/ML
AI Agents
Edge AI
Frontend
Angular
DevOps
Azure
Apply
$71k – $170k per year (Estimated) • In office • Full-Time • Netanya
Python
TypeScript
AI/ML
Accelerate
Fine-tuning
LangChain
LLM
NLP
Prompt Engineering
PyTorch
RAG
Edge AI
Hugging Face
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$23k – $47k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Katowice
Python
AI/ML
AI Agents
Edge AI
Apply
$18k – $32k per year (Estimated) • Remote/Hybrid • 2+ years exp • Moscow
Python
SQL
Databases
ClickHouse
Vertica
AI/ML
Airflow
Hadoop
NumPy
DevOps
Grafana
Analytics
Power BI
ETL/ELT
Apply
$117k – $239k per year (Estimated) • In office • 3+ years exp • Memphis
Apply
$440k per year • In office • Palo Alto
C++
AI/ML
CUDA Toolkit
CUDA
Frontend
Sass
Apply
$167k – $358k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Python
SQL
AI/ML
BERT
InfiniBand
DevOps
HPC
Apply
$440k per year • In office • Palo Alto
C++
Python
Rust
AI/ML
LLM
Reinforcement Learning
Apply
$600k per year • In office • Palo Alto
AI/ML
Reinforcement Learning
RLHF
DPO
Post-training
Apply
$117k – $239k per year (Estimated) • In office • 3+ years exp • Memphis
Apply
$167k – $358k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Python
SQL
AI/ML
BERT
InfiniBand
DevOps
HPC
Apply
$204k – $396k per year (Estimated) • In office • 5+ years exp • High School Diploma • Memphis
Bash
Management
Jira
Apply
$150k – $316k per year (Estimated) • In office • 4+ years exp • High School Diploma • Memphis
Bash
Management
Jira
Apply
$150k – $316k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Memphis
MATLAB
MATLAB
Simulink
Design
AutoCAD
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.