368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$164k – $206k per year
Location
In office (New York, San Francisco, Seattle, Austin)
Employment
Full-Time
Overview
Company
Impact
Profile match
Fluidstack is an AI cloud platform that designs, builds, and operates high-performance GPU clusters for frontier AI laboratories, enterprises, and governments. The company provides enterprise-grade bare-metal compute infrastructure - scaled across tens of thousands of state-of-the-art NVIDIA GPUs - specifically optimized for training large language models (LLMs) and running high-throughput inference.

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate

  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

The Infrastructure Team

Examples of key problems the team is working on

  • Bring gigawatts of accelerators from first power-on to production. Facility availability to ready-for-service across thousands of racks per site, with a new data hall landing every few weeks.

  • Make rack qualification faster than the fleet grows. Firmware baselines, burn-in, and cluster validation proven on every rack before a customer workload touches it, at a pace that never becomes the critical path.

  • Scale by tooling, not headcount. Deployed megawatts grow severalfold next year while the team stays near-flat, because anything done twice by hand becomes software.

Role Scope

  • Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.

  • Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.

  • Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.

  • Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.

  • Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.

  • Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.

  • Ability to travel 30-40% of the time to our Data Centers and Labs, as needed.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly,tell us where you would.

  • You've brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.

  • You work deep in Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups.

  • You've automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.

  • You've worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them.

  • You triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.

  • You travel for turn-up windows when a new data hall comes online.

  • Bonus: Kubernetes-based bare-metal provisioning. Accelerator platform bringup (NVIDIA, AMD, or custom). Burn-in and stress harness design. DCIM and inventory tooling.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$68k – $85k per year • In office • Full-Time • Master's Degree • San Jose
Python
AI/ML
AI Agents
DevOps
Amazon EC2
AWS
AWS Lambda
Bitbucket
CI/CD
CloudFormation
Docker
Git
Kubernetes
Terraform
Amazon S3
IAM
HPC
Cybersecurity
Least Privilege
Apply
$140k – $225k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
Python
TypeScript
Node JS
JavaScript
Python
FastAPI
Node JS
Fastify
Databases
PostgreSQL
Redis
Frontend
React.js
Vite
DevOps
Kubernetes
WebSockets
QA
Playwright
Vitest
Apply
$21k per year • In office • Contractor • Yekaterinburg
Python
SQL
Python
FastAPI
Flask
AI/ML
Claude
Claude Code
Embeddings
Function Calling
LLM
RAG
OpenAI
OpenAI Codex
Structured Outputs
DevOps
Docker
Git
Apply
$140k – $225k per year • Remote • Full-Time • 5+ years exp • Seattle
Python
Lua
Python
FastAPI
Celery
Pydantic
SQLAlchemy
Databases
pgvector
PostgreSQL
Redis
AI/ML
LLM
RAG
Hybrid Search
Reranking
Anthropic
LLM Evaluation
LLM Guardrails
OpenAI
Model Context Protocol
DevOps
OpenTelemetry
Apply
$185k – $260k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
DevOps
AWS
CI/CD
GCP
Kubernetes
GitHub
Cybersecurity
Clair
Dependabot
OWASP Top 10
OWASP ZAP
Snyk
Trivy
Apply
$173k – $250k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM
LLM Guardrails
Model Context Protocol
Apply
$224k – $279k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco • New York • Seattle • Austin
Python
JavaScript
Databases
PostgreSQL
Redis
AI/ML
Time Series Forecasting
Frontend
Bootstrap
DevOps
Ansible
CI/CD
Docker
Grafana
Incident Management
OpenTelemetry
Platform Engineering
Prometheus
Terraform
Robotics
Digital Twin
Apply
$258k – $300k per year • Remote • Full-Time
Python
SQL
Apply
$173k – $279k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
Python
AI/ML
Time Series Forecasting
IoT
OPC UA
Apply
$224k – $264k per year • In office • Full-Time • Austin • New York • San Francisco • Seattle
Python
SQL
Apply
$155k per year • In office • Full-Time • New York
Apply
$268k – $481k per year (Estimated) • Equity • In office • Full-Time • 10+ years exp • New York
Apply
$220k – $350k per year • Remote/Hybrid • Full-Time • 15+ years exp • New York • Princeton
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Azure
Azure DevOps
CI/CD
GitHub
Platform Engineering
Design
Figma
Management
Jira
QA
Playwright
Apply
$320k per year • In office • 10+ years exp • Bachelor's Degree • New York
AI/ML
Anthropic
Claude
Multimodal AI
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • New York
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.