386,626open jobs
10,106companies
50,572added this week
Browse all
Salary
$160k – $385k per year (Estimated)
Location
In office (Memphis)
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Network Operations Center (NOC) Specialist, you are the eyes and the voice of the campus - never the hands. You watch campus health signals around the clock, detect and verify site-impacting events, assemble the right responders fast, and run incident communications leadership can trust. You make sure no major incident closes without a timeline, a report, and a tracked corrective project. This is a communications-and-judgment role at the center of site operations, not a junior-engineering holding pen. You do not do wrench work, plant operation, deep root-cause analysis, monitoring design, technical SEV command, or tool building - those belong to SiteOps, Facilities, Hardware Failure Analysis, Site SRE, and Software Platforms.

RESPONSIBILITIES:

  • Staff the console per shift schedule and watch the designated signal surface: cluster health, node availability, network health, facility trend panels, storage alarms, and threshold breaches.
  • Acknowledge every page within SLA; classify (actionable / known / noise) and log disposition; feed noise patterns back to SRE so signal quality keeps improving.
  • Detect, verify, and escalate within time budgets; operate the escalation matrix (NOC → on-call SRE → domain owners) and page correctly the first time.
  • Open and run incident bridges; own stakeholder communications (first update within SLA, then fixed cadence); maintain the incident timeline in real time; call out ownership stalls.
  • Produce first-pass RCA framing (what happened, when, what’s impacted, who’s engaged) and hand it to SRE / Hardware Failure Analysis for depth - the NOC does not publish root cause.
  • Run structured shift handoffs and durable shift logs; maintain cross-site awareness.
  • Write major-incident reports; open corrective projects in Linear and chase them to closure - the NOC is the nag of record.
  • Maintain and continuously improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-run game days.

BASIC QUALIFICATIONS:

  • Experience in a 24/7 operations environment (NOC, SOC, dispatch, mission control, or equivalent).
  • Proven ability to acknowledge, classify, and escalate incidents under SLA in a high-signal environment.
  • Experience opening and running incident bridges, including stakeholder updates on a fixed cadence and live timeline hygiene.
  • Excellent written and verbal communication skills; able to write clear updates while an incident is in progress.
  • Demonstrated pattern recognition across multiple domains (compute, network, storage, and/or facilities signals) and curiosity about how those systems interact.
  • Experience following, maintaining, and improving operational process (runbooks, escalation matrices, handoffs, or similar).
  • Willingness and ability to work a rotating shift schedule, including nights and weekends, as part of continuous campus coverage.

PREFERRED SKILLS AND EXPERIENCE:

  • Prior NOC, data center operations, or campus reliability experience in a high-performance computing, AI/ML infrastructure, or large-scale production environment.
  • Experience writing major-incident reports and driving corrective follow-ups to closed (e.g. tickets, projects, or Linear).
  • Familiarity with Linear or similar work-tracking tools for corrective action programs.
  • Experience partnering with SRE, SiteOps, and Facilities on escalations and post-incident follow-through.
  • Participation in game days, tabletop exercises, or runbook improvement programs.
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,626 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$160k – $304k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Frisco • New York • Toronto • Ann Arbor
AI/ML
AI Agents
Claude
Claude Code
Cursor
Model Context Protocol
Prompt Engineering
DevOps
GitHub
Management
Linear
Slack
Apply
$25k – $56k per year (Estimated) • Remote • Full-Time • Moscow
SQL
Databases
PostgreSQL
DevOps
SLI/SLO/SLA
Apply
$60k – $120k per year (Estimated) • In office • Full-Time • 1+ year exp • High School Diploma • San Antonio
DevOps
SLI/SLO/SLA
Apply
$26k – $59k per year (Estimated) • Remote • Full-Time • 2+ years exp • Moscow
AI/ML
LiteLLM
LLM
LLM Guardrails
NLP
Open WebUI
DevOps
Azure
Azure DevOps
SLI/SLO/SLA
Management
Confluence
Jira
n8n
YouTrack
Apply
$55k – $150k per year (Estimated) • In office • Full-Time • Milton Keynes
Python
SQL
DevOps
SLI/SLO/SLA
Analytics
Power BI
Apply
$140k – $271k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Memphis
C++
Java
Python
Rust
Cybersecurity
CVE
Apply
$178k – $351k per year (Estimated) • In office • 1+ year exp • Bachelor's Degree • Memphis
C#
Python
SQL
C#
.NET
Apply
$150k – $316k per year (Estimated) • In office • 3+ years exp • Memphis
Bash
PowerShell
Python
DevOps
Ansible
Configuration Management
GitOps
Puppet
Terraform
IoT
MQTT
OPC UA
Apply
$167k – $358k per year (Estimated) • In office • Contractor • 5+ years exp • Bachelor's Degree • Memphis
Bash
PowerShell
Python
AI/ML
InfiniBand
NCCL
DevOps
Ansible
GitOps
HPC
Terraform
Apply
$204k – $396k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
C++
Java
Python
Rust
Apply
$67k – $135k per year (Estimated) • In office • Full-Time • 5+ years exp • Memphis
DevOps
Incident Management
Apply
$140k – $200k per year • In office • PhD • Memphis
Swift
AI/ML
Text-to-Speech
Mobile
Fastlane
SwiftUI
DevOps
CI/CD
Git
Vercel
Management
Google Docs
Stripe
Marketing
LinkedIn
Apply
$140k – $200k per year • In office • Memphis
Java
Kotlin
Mobile
Kotlin Multiplatform
DevOps
GCP
Marketing
LinkedIn
Apply
$140k – $200k per year • In office • Internship • PhD • Memphis
C#
C++
C#
.NET
AI/ML
Text-to-Speech
DevOps
CI/CD
Vercel
Management
Google Docs
Stripe
Marketing
LinkedIn
Apply
$65k – $133k per year (Estimated) • In office • Full-Time • Memphis
Apply
See all jobs
This is one of many
386,626 more open roles from verified company boards, updated every day.