Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs
Crusoe (formerly Crusoe Energy Systems) is an energy-first AI infrastructure and cloud computing company headquartered in Denver, Colorado. The company specializes in building and operating high-performance AI data centers powered by stranded, wasted, or underutilized energy sources—such as flared natural gas from oil fields and surplus renewable power.

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role:

As a Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters. You will own the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud. Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters.

San Francisco, Sunnyvale, Bellevue (Onsite)

What You’ll Be Working On:

  • Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe’s AI Cloud Stack.

  • CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.

  • Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.

  • Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc.

  • Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary.

  • Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.

  • Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.

What You’ll Bring to the Team:

  • Education & Experience: 12+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.

  • Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes.

  • Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.

  • CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.

  • Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack.

  • Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios.

  • System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).

  • Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.

  • Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.

Bonus Points:

  • Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.

  • Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).

  • Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $250,000 -$300,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Recommended for you based on this role

Similar stack
Same company
In your city
$250k – $300k per year • Equity • In office • Full-Time • 10+ year exp • San Francisco
Python
Rust
AI/ML
CUDA Toolkit
DevOps
CI/CD
KVM
QEMU
Xen
Apply
$250k – $300k per year • Equity • In office • Full-Time • San Francisco
Java
Python
Rust
AI/ML
AI Agents
PyTorch
DevOps
GCP
Kubernetes
Apply
$215k – $260k per year • Equity • In office • Full-Time • San Francisco
Java
Python
Rust
AI/ML
AI Agents
PyTorch
DevOps
GCP
Kubernetes
Apply
$175k – $200k per year • Equity • In office • Full-Time • 2+ year exp • Bachelor's Degree • San Francisco
AI/ML
RAG
DevOps
Kubernetes
Apply
Equity • In office • Full-Time • 5+ year exp • San Francisco
PowerShell
Python
AI/ML
Claude
DevOps
Rest API
Terraform
Cybersecurity
Okta
Management
Google Workspace
Slack
Apply
$260k – $310k per year • Equity • In office • Full-Time • 15+ year exp • San Francisco
Apply
Equity • In office • Full-Time • 12+ year exp • Bachelor's Degree • San Francisco
AI/ML
LLM
DevOps
AWS
Azure
GCP
Kubernetes
KVM
QEMU
SLURM
Apply
$140k – $165k per year • Equity • Remote/Hybrid • Bachelor's Degree
Java
Python
Rust
AI/ML
AI Agents
PyTorch
DevOps
GCP
Kubernetes
Apply
$170k – $205k per year • Equity • In office • Full-Time • 4+ year exp • San Francisco
Java
Python
Rust
AI/ML
AI Agents
PyTorch
DevOps
GCP
Kubernetes
Apply
$117k – $135k per year • Equity • In office • Full-Time • 1+ year exp • San Francisco
Java
Python
Rust
AI/ML
AI Agents
PyTorch
DevOps
GCP
Kubernetes
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
San Francisco
Remote work
In office
Employment
Full-Time

Compensation

Salary
$250k – $300k per year
Benefits
Life insurance, Parental leave, Professional development, Restricted stock units, Retirement plans, Vision insurance
Equity
Equity stake in a tech company