368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$160k – $230k per year
Location
In office (Austin, New York, San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Fluidstack is an AI cloud platform that designs, builds, and operates high-performance GPU clusters for frontier AI laboratories, enterprises, and governments. The company provides enterprise-grade bare-metal compute infrastructure - scaled across tens of thousands of state-of-the-art NVIDIA GPUs - specifically optimized for training large language models (LLMs) and running high-throughput inference.

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate

  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

The Data Center Operations Team

Examples of key problems the team is working on

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.

  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.

  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.

Role Scope

  • Own the end-to-end site handover framework: define the gates, acceptance criteria, and sign-off procedures that move a new facility from construction to live operations without dropped terms or late surprises.

  • Embed into design, construction, and due diligence teams early enough to shape maintainability requirements before they become field problems.

  • Drive the cross-functional handover rhythm across training, documentation, systems access, and knowledge transfer, surfacing blockers weeks before they hit the go-live schedule.

  • Build and maintain the SOPs that govern critical datacenter operations across the fleet, with metrics that track adoption, execution quality, and efficiency at each site.

  • Lead incident management and stability improvement programs, including post-incident reviews with root cause analysis, corrective action tracking, and preventive maintenance oversight that reduces unplanned outages across the global footprint.

  • Produce the dashboards and reporting that give leadership visibility into stability metrics and incident trends, and run the CAPA programs that turn that data into durable fixes.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly,tell us where you would.

  • You have run program management in mission-critical environments where a delayed handover or missed SOP had real operational consequences, not just schedule slippage.

  • You have designed operational frameworks from scratch: handover gates, SOP libraries, incident management programs built without a legacy system to copy from.

  • You quarterback across design, construction, supply chain, and site ops teams simultaneously, and other teams call you when a cross-functional workstream is stuck.

  • You write clearly enough to distill a complex operational issue into a decision and a next action for a site lead, an executive, or a counterparty who was not in the room.

  • You track incident trends and CAPA status in live dashboards and follow corrective actions through to closure, not just to initial assignment.

  • You have personally built or maintained SOPs and measured whether they were actually followed, not just whether they existed.

  • Bonus: ITIL, PMP, or PgMP certification. Hyperscale or large colo operator experience. Familiarity with ASHRAE, Uptime Institute, or TIA-942 standards. Exposure to datacenter construction and commissioning processes.

Salary & Benefits

  • Competitive total compensation package (salary + equity)

  • Retirement or pension plan, in line with local norms

  • Health, dental, and vision insurance

  • Generous PTO policy, in line with local norms

Total compensation may also include equity in the form of restricted stock units.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Austin
Sr Quality Engineer 2 days ago
$78k – $158k per year (Estimated) • In office • Full-Time • 4+ years exp • High School Diploma • Puerto Rico
Python
AI/ML
ChatGPT
Cybersecurity
CAPA
Analytics
Power BI
Apply
$24k – $51k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
Node JS
JavaScript
Node JS
Commander.js
Databases
Apache Kafka
OpenSearch
PostgreSQL
Redis
DevOps
AWS
Chaos Engineering
CI/CD
GitOps
Grafana
Helm
Incident Management
Jaeger
Kubernetes
Opsgenie
PagerDuty
Prometheus
SLI/SLO/SLA
Terraform
Zabbix
Cybersecurity
Tcpdump
Management
Jira
Apply
$117k – $251k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
SQL
DevOps
Incident Management
Apply
$133k – $284k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
DevOps
Incident Management
SLI/SLO/SLA
Splunk
Management
Confluence
ServiceNow
Apply
$66k – $156k per year (Estimated) • In office • Full-Time • Singapore
SQL
DevOps
Incident Management
Apply
$173k – $250k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM
LLM Guardrails
Model Context Protocol
Apply
$224k – $279k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco • New York • Seattle • Austin
Python
JavaScript
Databases
PostgreSQL
Redis
AI/ML
Time Series Forecasting
Frontend
Bootstrap
DevOps
Ansible
CI/CD
Docker
Grafana
Incident Management
OpenTelemetry
Platform Engineering
Prometheus
Terraform
Robotics
Digital Twin
Apply
$258k – $300k per year • Remote • Full-Time
Python
SQL
Apply
$173k – $279k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
Python
AI/ML
Time Series Forecasting
IoT
OPC UA
Apply
$224k – $264k per year • In office • Full-Time • Austin • New York • San Francisco • Seattle
Python
SQL
Apply
$96k – $200k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Hillsboro • Austin
C++
AI/ML
NCCL
DevOps
HPC
Apply
In office • Internship • Master's Degree • Austin
C++
Python
C++
PyTorch C++
AI/ML
AI Agents
CUDA
CUDA Toolkit
LLM
NCCL
PyTorch
TensorRT
TensorRT-LLM
Triton
vLLM
Apply
$106k – $145k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Austin
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$120k – $140k per year • In office • Full-Time • Charlotte • Raleigh • Dallas • Boston • New York
SQL
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.