379,681open jobs
9,923companies
50,127added this week
Browse all
Salary
$167k – $358k per year (Estimated)
Location
In office (Memphis)
Seniority
Senior · 5+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match

xAI

xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

SpaceXAI is looking for an exceptional network engineer with experience in mission-critical, large-scale production environments to support the design, build-out, and operation of networks that power our AI supercomputer campuses. As a member of the Supercomputer Infrastructure / Network Engineering team, you will provide design and operational support for the fabrics used by GPU training and inference clusters, site operations, automation and controls, and facilities teams. The ideal candidate thrives in intense, high-flux environments, brings a strong sense of urgency balanced with operational excellence, communicates clearly, and demonstrates high technical acumen.

RESPONSIBILITIES:

  • Design and implement highly available, low-latency, high-bandwidth networks, carefully balancing routing, congestion control, and redundancy technologies for AI training fabrics, inference front-ends, storage, and site/OT networks.
  • Design and maintain supercomputer data center and campus networks in accordance with company network standards. Collaborate with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams.
  • Evaluate, procure, and deploy network hardware including data-center class switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond.
  • Contribute to maturing network automation tooling; implement configuration analysis, linting, validation, and scalable deployment frameworks (GitOps / IaC).
  • Plan and coordinate network change windows with stakeholders to perform software updates, hardware refreshes, cluster expansions, and general maintenance (including evenings and weekends when required by compute schedules).
  • Troubleshoot and resolve network-related issues affecting cluster health and job performance; publish root cause analysis (RCA) documentation and host retrospectives.
  • Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns; serve as on-call or networking responsible engineer during operations.
  • Proactively tailor network monitoring and telemetry (fabric health, congestion, packet loss, NCCL/collective performance) so issues are detected before they impact training or inference.
  • Continuously create and update network documentation, including architecture overviews, design drawings, fiber/cable plant records, and operational procedures.
  • Collaborate with cross-functional teams to identify and resolve potential design issues, especially systemic or cascading failure modes and false redundancy in AI fabrics and site networks.
  • Perform job walks with customers, vendors, and contractors to gather requirements and produce implementation plans for new halls, rows, and campus interconnects.
  • Ensure networks are configured and maintained in compliance with industry and cybersecurity standards (e.g., ITAR, ISO, NIST), with particular attention to segmentation between compute fabrics, storage, OT/controls, and corporate networks.

BASIC QUALIFICATIONS:

  • Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience;
    • OR 5+ years of professional network engineering experience in lieu of a degree.
  • Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / data-center environments.
  • Functional experience with multiple network vendors in production or lab environments.
  • Experience with GitOps and Infrastructure as Code frameworks, both as a user and contributor.

PREFERRED SKILLS AND EXPERIENCE:

  • Strong understanding of the OSI model and network standards.
  • Hands-on experience with Cisco, Arista, Juniper, and/or NVIDIA Spectrum-X data-center class switches.
  • Experience with RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience is a plus.
  • Working knowledge of AI training and inference traffic patterns and how they behave on the network (collectives, congestion, ECMP, adaptive routing). Familiarity with NCCL is a plus.
  • Experience with WDM and large-scale single-mode / multimode fiber plants, including OTDR and acceptance testing.
  • Experience with switch port security, network segmentation, QoS, multicast, and redundancy protocols.
  • Familiarity with network monitoring and Layer 1 test tools; experience building operational telemetry and dashboards.
  • Proficiency in scripting (Bash / PowerShell / Python) and automation frameworks (Terraform, Ansible, etc.).
  • Linux and Windows system administration experience professionally or from labs.
  • Industry-standard certifications such as CCNA or CCNP.
  • Experience supporting real-time systems, industrial control / OT networks, or high-reliability environments in data center, energy, aerospace, defense, or similar industries.
  • Excellent communication skills with internal and external customers, vendors, and management in both formal and informal settings.

ADDITIONAL REQUIREMENTS:

  • Ability to pass applicable background checks for site access.
  • Ability to work in tight quarters; physical dexterity is necessary to perform job functions.
  • Availability for extended hours and/or weekends as the schedule varies with cluster build-out and operational needs; flexibility is required.
  • Ability to provide 24x7 on-call support in emergency situations and participate in an after-hours on-call rotation.
  • Willingness to travel (up to 20%) between supercomputer campuses and related sites.
  • Ability to lift 30 lbs.
  • Ability to work at heights.
  • Ability to drive (active valid driver’s license).

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
379,681 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Memphis
$150k – $316k per year (Estimated) • In office • 3+ years exp • Memphis
Bash
PowerShell
Python
DevOps
Ansible
Configuration Management
GitOps
Puppet
Terraform
IoT
MQTT
OPC UA
Apply
$80k – $170k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Montreal
Python
Databases
Databricks
AI/ML
LangChain
LLMOps
MLFlow
OpenAI
DevOps
Azure
CI/CD
Docker
Git
GitOps
Kubernetes
Platform Engineering
Apply
Platform Engineer II 4 hours ago
$90k – $188k per year (Estimated) • Equity • In office • Internship • Bachelor's Degree • New York
DevOps
ArgoCD
AWS
Azure
CI/CD
CloudFormation
GCP
GitHub
GitHub Actions
GitOps
Istio
Jenkins
Kubernetes
Service Mesh
Terraform
Cybersecurity
Zero Trust
Apply
$143k – $295k per year (Estimated) • In office • 5+ years exp • London
DevOps
AWS
Azure
CI/CD
CircleCI
Docker
GCP
GitHub
GitHub Actions
GitOps
Kubernetes
Terraform
Apply
Backend Engineer 7 hours ago
Remote/Hybrid • Full-Time • 5+ years exp • Athens
Python
Scala
SQL
TypeScript
JavaScript
Databases
Apache Kafka
ElasticSearch
Kafka
PostGIS
PostgreSQL
Snowflake
AI/ML
Human-in-the-Loop
Frontend
Vue.js
DevOps
GitOps
Kubernetes
Apply
$150k – $316k per year (Estimated) • In office • 3+ years exp • Memphis
Bash
PowerShell
Python
DevOps
Ansible
Configuration Management
GitOps
Puppet
Terraform
IoT
MQTT
OPC UA
Apply
$204k – $396k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
C++
Java
Python
Rust
Apply
$142k – $322k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Dublin
AI/ML
EU AI Act
DevOps
AWS
Azure
CI/CD
GCP
IAM
Cybersecurity
GDPR
ISO 27001
SOC 2
Apply
$209k – $454k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$209k – $454k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Memphis
Apply
$67k – $135k per year (Estimated) • In office • Full-Time • 5+ years exp • Memphis
DevOps
Incident Management
Apply
$140k – $200k per year • In office • PhD • Memphis
Swift
AI/ML
Text-to-Speech
Mobile
Fastlane
SwiftUI
DevOps
CI/CD
Git
Vercel
Management
Google Docs
Stripe
Marketing
LinkedIn
Apply
$140k – $200k per year • In office • Memphis
Java
Kotlin
Mobile
Kotlin Multiplatform
DevOps
GCP
Marketing
LinkedIn
Apply
$140k – $200k per year • In office • Internship • PhD • Memphis
C#
C++
C#
.NET
AI/ML
Text-to-Speech
DevOps
CI/CD
Vercel
Management
Google Docs
Stripe
Marketing
LinkedIn
Apply
$65k – $133k per year (Estimated) • In office • Full-Time • Memphis
Apply
See all jobs
This is one of many
379,681 more open roles from verified company boards, updated every day.