1,334,991open jobs
78,300companies
214,577added this week
Browse all
Salary
$250k – $300k per year
Location
In office (San Francisco, Sunnyvale, Bellevue)
Seniority
Staff
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 7, 2026. First seen by Alion on Oct 7, 2026. Crusoe scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Crusoe is an energy and cloud company founded in 2018 that builds data centres next to stranded and low-carbon power sources. It started by converting flared natural gas into computing capacity and has become a large supplier of graphics processing capacity for artificial intelligence training and inference. The company develops sites, energy systems and its own managed AI cloud.

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

Build the operating system for the AI datacenter

Crusoe operates one of the world's largest managed GPU fleets, and it is growing fast. A fleet at this scale cannot be run the way GPU clouds have traditionally been run: runbooks, war rooms, and heroics. It has to be run by a unified platform that senses, reasons about, and acts on the entire fleet, so infrastructure that used to take a team to operate takes a service instead. That is what our cloud platform team builds. You will work on the control plane for one of the largest AI fleets in the world, at a point where fleet autonomy is still an open problem: nobody has fully solved this at this scale, in this market.

A true platform, not internal tooling: We want to be explicit about that heading, because it is the thing most infrastructure roles get wrong. Everything we build ships as a product, fleet engineers, SREs, and product teams across Crusoe build their own services and workflows on top of what we ship. A platform team does not scale by doing everyone's work; it scales by making everyone's work self-serve. Concretely:

  • API-first. Every capability is exposed through well-designed, versioned APIs behind a single gateway. If it isn't an API, it doesn't exist. No side doors, including for us.

  • SDKs and paved paths. First-class client libraries, workflow templates, and golden paths so a fleet or SRE engineer can ship a new remediation or lifecycle workflow in days without asking the platform team.

  • Micro frontends and a self-serve portal. Teams plug their own UI surfaces into one developer portal instead of building one-off dashboards. One console for the fleet, extensible by every team.

  • Platform as product. Internal teams are customers. We own contracts, versioning, deprecation policy, quotas, documentation, and support. Adoption is our success metric: the platform wins when other teams choose it because it is the fastest path, not because it is mandated.

What this platform is

Four layers, built as one system:

  • Agents on every site and host that collect telemetry and execute commands.

  • A distributed infra graph: Models system connections down to the rack, fabric, power, and cooling layers. By integrating these connections with telemetry signals, the platform can precisely trace events to identify their blast radius and root cause.

  • A reconciliation core: workflow engine, policy engine, and state reconciler that continuously close the gap between intended state and reality, exposed through the API gateway.

  • Domain services: Services spanning provisioning, firmware upgrade, validation, deployment, repair and RMA, capacity, power and thermal, and Day-2 operations. Built once, run fleet-wide, consumable by any team through APIs and SDKs.

We operate on a continuous autonomy loop-sense, correlate, reason, act, learn-incorporating guardrails that evolve from recommendation to full automation. We treat every recurring manual intervention as a signal to engineer the next automation.

You'll thrive here if you

  • Want to build a platform, not integrate one. This is core distributed-systems engineering: event buses, graph models, reconciliation loops, policy evaluation.

  • Treat internal engineers as customers and sweat API ergonomics, docs, and onboarding the way product teams sweat UX.

  • Like owning a hard abstraction and defending it as ten teams build on top of you.

  • Believe the interesting problems are where physical infrastructure meets software: a firmware counter, a thermal event, and a scheduling decision are one problem, not three.

  • Measure yourself by what stops paging humans, and by how fast another team ships on your platform.

What you'll do

  • Design and build core platform services: RBAC, tenancy, the workflow engine, policy engine, and state reconciler that drive fleet actions safely at scale.

  • Design the public face of the platform: the API gateway, resource model, and versioned API contracts that fleet, SRE, and product teams build against.

  • Build SDKs, workflow templates, and golden paths that make the platform self-serve, plus the developer portal and micro frontend framework that let teams bring their own UI surfaces.

  • Build the inventory and topology graph as the fleet's source of intended truth, and the pipelines that keep it honest against reality (metadata drift is one of our top verified incident root causes; you will kill it).

  • Build site, GPU, and network agents and the event bus that moves fleet telemetry and commands reliably.

  • Deliver the platform roadmap: pilot site on the foundation layer, first site deployed entirely through the platform, zero-downtime firmware upgrades, first fully automated RMA, then 100K+ GPUs on platform with MTTD under 60 seconds and MTTR under 30 minutes.

  • Work with embedded engineers from fleet and production engineering who bring the operational scar tissue, and turn it into services other teams extend.

Requirements

  • 10+ years building distributed systems, control planes, or infrastructure platforms.

  • Strong software engineering skills in Go, Python, or Rust.

  • Experience building platforms other engineers consume: public or internal APIs, SDKs, or developer tooling with real adoption.

  • Depth in at least one of: workflow/orchestration engines (Temporal or similar), event-driven architectures, graph data models, policy/rules engines, or reconciliation-based control loops (Kubernetes operator patterns).

  • Experience running what you build: you have carried a pager for a platform other teams depend on.

  • Systems thinking across the hardware/software boundary.

Bonus experience

  • Internal developer platforms: API gateways, service catalogs, Backstage-style portals, micro frontend architectures.

  • GPU or bare-metal fleet infrastructure: DCGM, Redfish/IPMI, firmware lifecycle.

  • High-cardinality observability platforms (per-GPU telemetry at fleet scale).

  • InfiniBand or RoCE fabrics.

  • AI agents applied to infrastructure triage and autonomous remediation.

About CAPE

Vision. Crusoe's infrastructure runs as a self-aware, self-healing system: anticipating and auto-remediating failures, shaping its own power demand, and tuning silicon-to-orchestration as one instrument. The world's most reliable, efficient, and sustainable AI compute platform.

Mission. We design, build, and operate the world's most reliable and energy-efficient AI infrastructure platform by treating the physical and digital layers as one software-defined system. Every day, for every workload, we automate away the latency, waste, and fragility between stranded energy and delivered intelligence.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $250,000 - $300,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,334,991 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
San Francisco
≈ $179k – $330k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Boston
Java
SQL
Java
Maven
Gradle
Databases
Apache Kafka
Mobile
JUnit
DevOps
Splunk
Ansible
OpenShift
OpenTelemetry
Prometheus
Jenkins
AWS
Kubernetes
Self-Healing
Opsgenie
Cybersecurity
Sysdig Secure
Management
ServiceNow
Agile
Apply
$171k – $255k per year • Equity • Remote (United States, Canada) • Full-Time • San Francisco • Toronto • Los Angeles • Denver • Chicago
Python
Go
Ruby
Scala
DevOps
CI/CD
Apply
Sr Software Developer 2 hours ago
≈ $127k – $213k per year (Estimated) • Remote (United States) • Full-Time • 7+ years exp • Associate's Degree • United States
JavaScript
PHP
SQL
PowerShell
Node JS
Databases
MySQL
Oracle
MS SQL
Frontend
npm
Apply
Software Architect 1 hour ago
$115k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Dallas
DevOps
Terraform
Puppet
Ansible
Chef
Kubernetes
Apply
$125k – $135k per year • Remote (United States) • 7+ years exp • Bachelor's Degree • United States
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
RabbitMQ
Frontend
Angular
React.js
DevOps
CircleCI
CI/CD
Jenkins
Git
AWS
Bitbucket
Management
Slack
Trello
Jira
Microsoft Teams
Agile
Scrum
Kanban
Apply
≈ $132k – $222k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp • San Francisco
JavaScript
Java
TypeScript
SQL
Node JS
Java
Spring Boot
Databases
PostgreSQL
RabbitMQ
Apache Kafka
AI/ML
Copilot
Claude
Machine Learning
Frontend
Angular
React.js
DevOps
Rest API
CI/CD
Jenkins
AWS
Kubernetes
Configuration Management
QA
Swagger
Apply
$250k – $270k per year • Hybrid • Full-Time • 7+ years exp • San Francisco
Go
AI/ML
Model Context Protocol
AI Agents
Tool Use
DevOps
Rest API
Kong
Kubernetes
Apply
≈ $154k – $322k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
C++
DevOps
WebRTC
eBPF
Robotics
Teleoperation
Apply
$90k – $110k per year • Hybrid • Full-Time • 3+ years exp • San Francisco • Los Angeles
Python
SQL
Databases
Snowflake
Amazon Redshift
DevOps
AWS
Amazon S3
Analytics
Tableau
Power BI
ETL/ELT
Alteryx
Looker
Domo
Microsoft Excel
Management
Asana
Smartsheet
Microsoft Teams
Apply
$100k – $300k per year • Equity 1–5% • Remote (United States) • Full-Time • 6+ years exp • San Francisco
Python
Rust
TypeScript
C++
Zig
Analytics
Microsoft Excel
Apply
See all jobs
This is one of many
1,334,991 more open roles from verified company boards, updated every day.