1,461,157open jobs
87,517companies
229,927added this week
Browse all
Salary
$230k – $280k per year
Location
In office (San Francisco, Sunnyvale)
Seniority
Staff · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 9, 2026. Crusoe scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Crusoe is an energy and cloud company founded in 2018 that builds data centres next to stranded and low-carbon power sources. It started by converting flared natural gas into computing capacity and has become a large supplier of graphics processing capacity for artificial intelligence training and inference. The company develops sites, energy systems and its own managed AI cloud.

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About This Role

Crusoe's Cloud Engineering organization is 380 people today and hiring toward 550. When we sell capacity, we make customers two promises: it will be ready on the date we said, and it will work while they use it. At this scale, keeping those promises doesn't happen by accident anymore. This role exists to make sure it happens on purpose.

You'll lead Engineering Operations: the operating system that lets Cloud Engineering leadership agree on outcomes, execute, and keep its promises to customers. You'll own the standing programs that make operations better every week (SLOs, incident follow-up, change safety, capacity delivery tracking, and operational data) and build the systems that feed the weekly executive Engineering Operations review. You don't just report when a metric slips; you find the owner, hold them to the commitment, and make sure a decision happens fast.

You'll start as the first hire in this function, working solo while you prove out the model above. As the mandate holds up, you'll shape how the function grows: what to build, who to bring in, and what to keep lean.

This isn't a role scoped to a single program or a single dashboard. You will be a strategic advisor to leadership on whether teams are actually executing operationally, and the person who makes clear owners, shared metrics, and a steady cadence stick across the org.

The ideal candidate has a technical background (software engineering, SRE, Technical Program Management, or product in an infrastructure context) and is close enough to production systems to hold senior engineering leaders to what they committed.

What You'll Be Working On

Engineering Operations programs

  • Own the standing Engineering Operations programs: SLO/SLI definition and attainment, incident management and follow-up, change safety (change policy, maintenance readiness, and pre-production checks), capacity delivery tracking, and operational data. Partner with the engineering leaders who own each program and hold the line on outcomes.

  • Define and track the metrics that show how the org runs, grows, builds, and enables itself: usable capacity, SLO attainment, MTTD and MTTR, customer-found incidents, on-time capacity delivery, milestone slip, change failure rate, follow-up closure rate, and change policy compliance. Data should tell the story before anyone has to ask.

  • Partner with engineering, SRE, data engineering, and product to keep operational data accurate, with one source of truth for each metric.

  • Partner with engineering, TPM, product, SRE, data scientists, customer success, and data center operations to ensure operations run smoothly across all functional areas.

  • Work toward a single pane of glass that makes the operational health of the organization easy to understand at a glance.

Accountability and operating cadence

  • Own the weekly executive Engineering Operations review: build the systems, data, and pre-reads that feed it, and make sure every risk on the page has an owner, a date, and a next step.

  • Run the weekly operations planning forum, and partner with Product on a monthly business review that ties operational health to customer outcomes.

  • Make ownership explicit: clear accountable owners, swimlanes, and articulated outcomes for every program and metric.

  • Hold teams to what they committed. Drive incident follow-ups, overdue actions, and slipping milestones to closure, and escalate fast when an owner can't fix the problem alone.

Growing the function

  • Operate as an individual contributor first: prove out the programs, metrics, and cadence before asking for headcount.

  • Make the case for more capacity as scope outgrows one person, then help hire and onboard the people who join the function.

  • Own the evolution of the operating system itself, not just the programs inside it: retire process that no longer earns its cost, and keep the overhead on engineering teams as low as possible.

What You'll Bring to the Team

  • 10+ years in software engineering, SRE, technical program management, or a technical product role, close enough to production systems to know what an SLO breach actually means.

  • A track record of leading org-wide programs at the staff level or above, and of driving accountability across senior engineering leaders without formal authority.

  • Comfortable starting as a team of one: you don't need a team in place to start driving impact, and you know when the case for headcount is real versus premature.

  • You've operated in high-growth infrastructure environments where processes are still being built; ambiguity doesn't paralyze you.

  • You're a natural coordinator who works across teams without formal authority. Engineering, SRE, and product leads trust you because you follow through.

  • You can turn messy, multi-source data into a clear picture of organizational health, and you know how to pick the few metrics that drive decisions over the many that fill dashboards.

  • You've designed operating rhythms (weekly reviews, business reviews, incident reviews) that leaders actually use, and you keep them light for the teams that feed them.

  • Scrappy, low-ego, high-drive. You build the program and tooling yourself when it doesn't exist yet, and you care more about the outcome than the credit.

Bonus Points

  • Time inside AWS, GCP, Azure, CoreWeave, Lambda Labs, or a similar cloud provider.

  • Experience running an SLO, incident review, or change management program at a cloud or infrastructure company.

  • Familiarity with incident.io, Opsgenie, PagerDuty, or similar incident management platforms at scale.

  • Background in AI/ML infrastructure.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $230,000 - $280,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,461,157 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
San Francisco
$127k – $191k per year • In office • Full-Time • Bachelor's Degree • Northbrook
AI/ML
Copilot
DevOps
Terraform
GCP
CloudFormation
Azure
AWS
Kubernetes
Incident Management
IAM
Cybersecurity
GDPR
Microsoft Entra ID
Management
ITIL
Apply
≈ $180k – $353k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara • Hillsboro
C
C
MPI
AI/ML
OpenMP
InfiniBand
DevOps
HPC
Apply
$99k – $206k per year • In office • TS/SCI • Full-Time • 8+ years exp • Bachelor's Degree • Sterling
Python
PowerShell
DevOps
Splunk
Ansible
Linux
Windows
DNS
DHCP
Cybersecurity
Nessus
CIS Benchmarks
NIST 800-53
Active Directory
PKI
SIEM
DLP
Apply
≈ $133k – $272k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • United States
DevOps
GCP
Azure
AWS
Apply
$135k – $180k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Iselin
Python
PowerShell
Bash
AI/ML
Copilot
Cursor
LLM
Anomaly Detection
DevOps
Terraform
GCP
Helm
GitHub Actions
Istio
CloudFormation
Linkerd
Prometheus
Pulumi
GitLab CI
Azure
GitOps
Windows Server
ArgoCD
AWS
Kubernetes
Ubuntu
Grafana
Amazon EKS
Google GKE
IAM
Linux
Windows
Cybersecurity
SOC 2
HIPAA
FedRAMP
Zero Trust
Least Privilege
Active Directory
Apply
In office • Kuala Lumpur
C#
C#
.NET
DevOps
Incident Management
SLI/SLO/SLA
Apply
≈ $3k – $8.5k per year (Estimated) • In office • Full-Time • Tashkent
Databases
MySQL
PostgreSQL
DevOps
Zabbix
Prometheus
Incident Management
SLI/SLO/SLA
Linux
Windows
TCP/IP
DNS
DHCP
VPN
Management
Jira
ITIL
ITSM
Apply
≈ $160k – $300k per year (Estimated) • Hybrid • United States
DevOps
SLI/SLO/SLA
Apply
$260k – $300k per year • Equity • Hybrid • 15+ years exp • Raleigh
DevOps
CI/CD
Incident Management
Management
Agile
Apply
Software Architect 7 hours ago
≈ $30k – $74k per year (Estimated) • Equity • In office • 12+ years exp • Bachelor's Degree • Hyderabad
Python
JavaScript
SQL
C++
AI/ML
Model Context Protocol
AI Agents
DevOps
Rest API
Windows Server
Kubernetes
SLI/SLO/SLA
Windows
Cybersecurity
SOC 2
OWASP SAMM
Threat Modeling
SBOM
OWASP
Apply
$170k – $205k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Sunnyvale
Python
Databases
PostgreSQL
AI/ML
CUDA Toolkit
CUDA
NCCL
InfiniBand
NVLink
ROCm
DevOps
Terraform
Puppet
Ansible
Chef
CI/CD
Docker
Kubernetes
SaltStack
Configuration Management
GitLab
Linux
Cybersecurity
osquery
Apply
$215k – $260k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco
DevOps
KVM
QEMU
Xen
HPC
Linux
Cybersecurity
CVE
Apply
$230k – $280k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco • Sunnyvale
AI/ML
CoreWeave
DevOps
GCP
Azure
AWS
Kubernetes
Grafana
AWS Lambda
Management
Confluence
Jira
Apply
$250k – $300k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco • Sunnyvale
DevOps
KVM
QEMU
Xen
HPC
Linux
Cybersecurity
CVE
Apply
$250k – $300k per year • Equity • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco • Bellevue • Sunnyvale
Python
Bash
Databases
PostgreSQL
AI/ML
CUDA Toolkit
CUDA
NCCL
InfiniBand
NVLink
ROCm
DevOps
Terraform
Puppet
Ansible
Chef
CI/CD
Docker
Kubernetes
SaltStack
Configuration Management
AWX
GitLab
Linux
Cybersecurity
osquery
Apply
$205k – $250k per year • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
Red Teaming
DevOps
Terraform
Ansible
GitHub Actions
CI/CD
AWS
AWS Lambda
IAM
Amazon ECS
Linux
DNS
DHCP
Cybersecurity
pfSense
PKI
Apply
≈ $249k – $448k per year (Estimated) • Equity • In office • 8+ years exp • San Francisco
Cybersecurity
HIPAA
Apply
≈ $249k – $448k per year (Estimated) • Equity • In office • 8+ years exp • San Francisco
Cybersecurity
HIPAA
Apply
≈ $249k – $448k per year (Estimated) • Equity • In office • 8+ years exp • San Francisco
Cybersecurity
HIPAA
Apply
$132k – $170k per year • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Washington • San Francisco
Apply
See all jobs
This is one of many
1,461,157 more open roles from verified company boards, updated every day.