368,910open jobs
9,449companies
47,822added this week
Browse all
Salary
$179k – $218k per year
Location
In office (San Francisco, Sunnyvale)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Crusoe (formerly Crusoe Energy Systems) is an energy-first AI infrastructure and cloud computing company headquartered in Denver, Colorado. The company specializes in building and operating high-performance AI data centers powered by stranded, wasted, or underutilized energy sources - such as flared natural gas from oil fields and surplus renewable power.

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

The Mission

Crusoe is building the world’s most climate-aligned AI infrastructure. As we scale toward unprecedented power densities and liquid-cooled architectures, the gap between "Data Center Design" and "Silicon Reality" must be bridged.

We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority on GPU platforms within the Data Center Engineering and Operations organization. Your mission is twofold: act as the primary technical consultant to our Data Center Engineering team to ensure future facilities are built for next-gen silicon, and provide the Operations team with the specialized tooling, SOPs, and predictive strategies needed to maintain peak cluster health.

The Strategic Bridge

  • For DC Engineering: You are the internal consultant. You translate upcoming GPU power/thermal roadmaps (NVIDIA/AMD) into design requirements for our next-generation facilities.

  • For Site Operations: You are the "Technical Enabler." You develop the diagnostic tools and technical SOPs that enable field technicians to resolve complex GPU issues with surgical accuracy.

  • For Sourcing: You are the "Technical Strategist." You define the technical sparing requirements and site-level inventory needs based on hardware failure telemetry.

Key Responsibilities

  • Engineering Education & Design Support: Provide deep-dive technical guidance to the Data Center Engineering team on upcoming silicon (e.g., NVIDIA Blackwell/Rubin, AMD MI350/400). Ensure future facility designs for power, cooling, and rack-spacing are ready for 2000W+ per-chip densities.

  • Predictive Operations & Telemetry: Leverage AI/ML methodologies to analyze fleet-wide telemetry (power draws, thermal gradients, and error rates). You will lead the transition from reactive troubleshooting to predictive maintenance, identifying "pre-failure" patterns in HBM or NVLink components before they impact customer training runs.

  • Technical Sparing Architecture: Architect the site-level sparing strategy from a technical perspective. Use failure telemetry and MTBF data to define the "Critical Spares List" and stocking levels required at each site to meet cluster uptime targets, providing these requirements to Sourcing for execution.

  • Operational Tooling & SOPs: Build the "Operational Blueprint" for the field. Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or thermal throttling.

  • Advanced Troubleshooting & RCA: Act as the Tier-3 escalation point for the most complex hardware failures in the production environment. Lead Root Cause Analysis (RCA) on systemic issues that span the boundary between hardware and facility environmental factors.

  • Silicon Roadmap Authority: Maintain a 24-month forward-looking view of NVIDIA and AMD architectures. Educate internal stakeholders on how transitions in HBM4, interconnect speeds, and liquid-cooling will impact Crusoe’s physical infrastructure.

  • Vendor & VAR Technical Lead: Support the technical relationship with OEMs and VARs. Audit their hardware builds, review their technical bulletins, and ensure their hardware roadmaps align with Crusoe’s operational and engineering standards.

Technical Requirements

  • Silicon & Fabric Mastery: Expert-level knowledge of NVIDIA (Hopper/Blackwell/Rubin) and AMD (Instinct) architectures. Mastery of the physical and logical layers of NVLink, NVSwitch, and InfiniBand.

  • Infrastructure Bridge-Building: Ability to translate "Silicon Data Sheets" into "Mechanical Engineering Requirements." You can explain how a GPU's specific heat-load profile affects CDU sizing and secondary loop design.

  • Data-Driven Diagnostics: Proficient in Python, Go, or Bash to build telemetry and health-check tools (utilizing DCGM and ROCm). Experience using large datasets or basic ML frameworks to build "Smart Monitoring" that filters critical health signals from noise.

  • Operational Reliability Analysis: Experience using failure telemetry to inform site-level sparing requirements and field-service workflows.

  • Thermal Management: Deep understanding of the operational realities of Direct-to-Chip (D2C) cooling, including fluid dynamics, pressure-drop curves, and the lifecycle of dripless couplings.

Qualifications

  • 10+ years in Hardware Engineering, Systems Architecture, or Data Center Infrastructure.

  • The "Consultant" Mindset: Proven track record of educating and influencing cross-functional teams (specifically Engineering and Operations).

  • GPU Authority: You have managed or architected GPU clusters at scale (thousands of nodes) at a hyperscaler, a GPU-specialized cloud, or a major silicon vendor.

Education: B.S. or M.S. in Electrical Engineering, Computer Engineering, or a related technical field.

Benefits:

  • Competitive compensation

  • Restricted Stock Units

  • Paid time off & paid holidays

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

Compensation Range

Compensation will be paid in the range of up to $179,000 -$218,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data. (#INDDIG)

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
In office • Internship • Ho Chi Minh City
C++
MATLAB
Perl
Python
MATLAB
Simulink
Apply
$14k – $32k per year (Estimated) • Remote/Hybrid • 3+ years exp • Moscow
Python
DevOps
CI/CD
Git
GitHub Actions
GitLab CI
Jenkins
Rest API
GitHub
GitLab
QA
Playwright
Postman
Pytest
Selenium
TestRail
Apply
Full Stack Engineer 6 days ago
$75k – $135k per year (Estimated) • Remote • Full-Time • Warsaw
Node JS
Python
SQL
JavaScript
AI/ML
LLM
PyTorch
RAG
Anthropic
OpenAI
Frontend
Next.js
React.js
DevOps
Docker
Kubernetes
Apply
$10k – $24k per year (Estimated) • In office • Nizhny Novgorod
Python
DevOps
Bitbucket
Management
Jira
Apply
$64k – $84k per year • Remote/Hybrid • Full-Time • Kraków
Python
Scala
Databases
Apache Kafka
AI/ML
Airflow
Spark
DevOps
CI/CD
Kubernetes
Amazon S3
Apply
$160k – $195k per year • Equity • In office • Full-Time • 7+ years exp • San Francisco • Sunnyvale
Design
AutoCAD
Apply
$185k – $225k per year • Equity • In office • Full-Time • Bachelor's Degree • Denver
C++
Python
AI/ML
CUDA
CUDA Toolkit
LLM
SGLang
vLLM
DevOps
Docker
Kubernetes
Apply
$285k – $335k per year • Equity • In office • Full-Time • 12+ years exp • San Francisco • Sunnyvale
DevOps
Kubernetes
SLURM
Apply
Senior Director, IT 11 days ago
$200k – $250k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Cybersecurity
SOC 2
Apply
$170k – $205k per year • Equity • In office • Full-Time • 5+ years exp • San Francisco • Bellevue • Sunnyvale
C++
Go
Java
Python
Rust
DevOps
Ansible
CI/CD
Configuration Management
gRPC
Kubernetes
Terraform
Apply
$89k – $193k per year (Estimated) • In office • Full-Time • 3+ years exp • High School Diploma • San Francisco
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$170k – $230k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • New York • San Francisco
Go
Java
Kotlin
Python
Scala
Databases
Apache Kafka
Databricks
Delta Lake
Snowflake
AI/ML
Dagster
dbt
Flink
Spark
Knowledge Graph
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 3 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,910 more open roles from verified company boards, updated every day.