368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$173k – $279k per year
Location
In office (San Francisco, New York, Seattle, Austin)
Employment
Full-Time
Overview
Company
Impact
Profile match
Fluidstack is an AI cloud platform that designs, builds, and operates high-performance GPU clusters for frontier AI laboratories, enterprises, and governments. The company provides enterprise-grade bare-metal compute infrastructure - scaled across tens of thousands of state-of-the-art NVIDIA GPUs - specifically optimized for training large language models (LLMs) and running high-throughput inference.

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate

  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

The Production Engineering Team

Examples of key problems the team is working on

  • Turn every network fault from a mystery into a closed ticket. Link diagnostics across router to router and NIC to router paths, remote command execution across the fleet, and repair visualization that shows what's broken and why, because at 10 GW scale debugging has to be systematic, not artisanal.

  • Build the network repair pipeline that runs itself. Automated fault detection through RMA initiation, ticket integration, transceiver lifecycle tracking, and return to service, across DC fabric, edge, and host to network layers, because a new site comes online every six months and manual repair doesn't keep up.

  • Build the network monitoring platform for a fleet that never stops growing. Alerting lifecycle and health dashboards for infrastructure spanning multiple hyperscale sites today and 10 GW by next year. It doesn't exist yet. We're building it.

Role Scope

  • Carry the on-call pager for the network fleet and run repair end to end: diagnose the fault, execute the fix or drive the RMA, and return the link to service across DC fabric, edge, and host to network layers.

  • Write Python and Go tooling that replaces manual diagnosis, including link tests, remote command execution across the fleet, and repair visualization that shows what's broken and why, so faults resolve in minutes, not hours.

  • Automate the repair pipeline from fault detection through RMA initiation, ticket integration, transceiver and optics tracking, and return to service, so a failure at any site follows the same path without manual handoffs.

  • Maintain the realtime monitoring and alerting that gives every on-call engineer a true picture of network health across all sites, and keep it accurate as new sites and hardware roll in.

  • Validate new sites and hardware into production by running the qualification tests that confirm a network is healthy before it carries traffic.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly,tell us where you would.

  • You've carried a pager for a production network and can walk into an outage, find the fault, fix it or drive the RMA, and write the postmortem without someone walking you through it.

  • You've written Python or Go scripts that replaced a manual, repetitive network task, from link diagnostics to config pushes to fleet-wide command execution.

  • You understand how a transceiver fault, a misconfigured route, and a power event each show up differently in the data, and you can tell them apart from the signals alone.

  • You've worked hands-on with link diagnostics, optics, and network monitoring protocols such as gNMI, gRPC, NETCONF, and SONiC, and you're comfortable at the CLI on switches and routers across a fleet.

  • You treat toil as a bug. If a repair step means SSHing into ten boxes by hand, you script it once and never do it by hand again.

  • You reach real competence in an unfamiliar part of the stack fast, and you document what you learn so the next on-call engineer doesn't start from zero.

  • Bonus: RMA and repair lifecycle automation. Large-scale datacenter fabric (BGP, ECMP, spine-leaf). Out-of-band network management. Fluency with AI coding tools such as Claude Code or Cursor to move faster on scripts and tooling.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$87k – $166k per year (Estimated) • Equity • Remote • Full-Time • PhD
Python
Ruby
SQL
JavaScript
Ruby
RSpec
Ruby on Rails
Databases
PostgreSQL
AI/ML
Claude
Claude Code
AI Agents
OpenAI Codex
Frontend
GraphQL
Vue.js
DevOps
GitLab
Management
Slack
Apply
$55k – $157k per year (Estimated) • Remote • Full-Time • Sydney
C++
Go
Lua
Python
C++
CMake
Databases
ActiveMQ
Aerospike
Apache Kafka
Cassandra
DevOps
Docker
gRPC
Kubernetes
Apply
$28k – $65k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Bash
PowerShell
Python
Node JS
JavaScript
Node JS
Commander.js
AI/ML
AI Agents
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
Kubernetes
Amazon ECS
IAM
Cybersecurity
Crowdstrike
Zero Trust
Apply
$72k – $201k per year (Estimated) • In office • Full-Time • 3+ years exp • PhD • Basel
Python
AI/ML
Multimodal AI
OpenCV
PyTorch
scikit-image
Scikit-learn
TensorFlow
AI Agents
DevOps
Git
Apply
$20k – $52k per year (Estimated) • In office • Full-Time • Gurgaon
SQL
Python
Python
pySpark
AI/ML
Hadoop
Spark
Analytics
Power BI
Tableau
Apply
$173k – $250k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM
LLM Guardrails
Model Context Protocol
Apply
$224k – $279k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco • New York • Seattle • Austin
Python
JavaScript
Databases
PostgreSQL
Redis
AI/ML
Time Series Forecasting
Frontend
Bootstrap
DevOps
Ansible
CI/CD
Docker
Grafana
Incident Management
OpenTelemetry
Platform Engineering
Prometheus
Terraform
Robotics
Digital Twin
Apply
$258k – $300k per year • Remote • Full-Time
Python
SQL
Apply
$173k – $279k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
Python
AI/ML
Time Series Forecasting
IoT
OPC UA
Apply
$224k – $264k per year • In office • Full-Time • Austin • New York • San Francisco • Seattle
Python
SQL
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 2 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.