1,249,613open jobs
72,355companies
214,617added this week
Browse all
Salary
$220k – $320k per year
Location
In office (Seattle)
Seniority
Staff
Visa
H-1B filings in 12 months: 10 · for this role: 3

Confirmed on the employer's own hiring board on Oct 5, 2026. First seen by Alion on Aug 7, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the role

Nscale is hiring a Staff Software Engineer to build Fleet Manager - the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.

This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high - the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.

This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.

What you'll work on

  • Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready - BMC configuration, DHCP reservations, and provisioning state machines.
  • Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
  • Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
  • GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
  • Network configuration: switch lifecycle automation and network state management across the fleet.
  • Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
  • Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.

Responsibilities

  • Domain-level technical direction. Own the architecture for a major Fleet Manager domain - such as provisioning, validation, or remediation - influencing engineers across the team and adjacent squads.
  • Design and build production-grade automation. Implement device provisioning, burn-in testing, network configuration, and hardware health validation workflows in Python.
  • Engineer for reliability and auditability. Treat idempotency, resumability, checkpointing, retries, replay, and failure handling as first-class design concerns.
  • Integrate broadly. Connect Fleet Manager with datacenter infrastructure management systems, cloud orchestration platforms, and bare-metal provisioning tools.
  • Create leverage through standards. Establish shared patterns, libraries, conventions, and operational runbooks that other engineers build on.
  • Own production outcomes. Operate what you build with strong observability, alerting, incident response, and day-2 operational discipline.
  • Mentor and influence. Raise the bar through design reviews, implementation guidance, and operational best practices.
  • Use AI to accelerate delivery while maintaining architectural coherence.

Requirements

  • Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering.
  • Strong proficiency in Python - Fleet Manager is built entirely in Python.
  • Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
  • Track record of delivering automation systems from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization).
  • Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
  • You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.
  • Excellent communication skills to build consensus with stakeholders, both internally and externally, in a fast-paced, high-agency environment.

Preferred

  • Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
  • Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
  • Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE)
  • Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
  • Open-source contributions in infrastructure automation or cloud-native tooling

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000—$320,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,249,613 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Seattle
$110k – $148k per year • In office • Secret • Full-Time • 9+ years exp • Bachelor's Degree • United States
Management
Agile
Scrum
Apply
≈ $144k – $264k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • New York
Java
SQL
Java
Spring Boot
Databases
Oracle
AI/ML
LLM
Anomaly Detection
DevOps
GCP
Azure
CI/CD
GitOps
Jenkins
Git
AWS
Linux
Unix
Apply
$96k – $130k per year • Remote (United States) • Full-Time • 7+ years exp • Columbia
JavaScript
SQL
C#
C#
ASP.NET Core
Databases
Redis
ElasticSearch
AI/ML
LLM
Edge AI
Frontend
React.js
DevOps
Azure
AWS
Cybersecurity
Auth0
Apply
≈ $132k – $242k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Spring
Python
JavaScript
TypeScript
C#
Node JS
Python
Flask
C#
ASP.NET Core
WPF
Databases
Databricks
Delta Lake
Frontend
Angular
React.js
DevOps
Rest API
GCP
GitHub Actions
Azure
CI/CD
AWS
GitHub
Windows
Analytics
ETL/ELT
Apply
≈ $143k – $263k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Warren
Python
C++
DevOps
RTOS
Management
Agile
Scrum
Apply
DevOps Engineer 10 days ago
$30k – $50k per year • Equity 0.1–0.5% • Remote (Indonesia, Taiwan, Turkey, Vietnam) • Contractor • 2+ years exp • Bachelor's Degree • Taiwan
Python
DevOps
GCP
GitLab CI
CI/CD
Git
AWS
Docker
Kubernetes
Linux
Cybersecurity
PKI
Apply
$40k – $70k per year • Equity 0.1–1% • Remote (United States, Canada, Indonesia, Taiwan, Vietnam) • Full-Time • 3+ years exp • Singapore
Python
Java
Scala
Databases
Redis
AI/ML
Cursor
Windsurf
AI Agents
Mobile
Room
DevOps
Docker
Apply
$6k – $24k per year • Equity • Remote (United States, Taiwan) • Internship • Bachelor's Degree
Go
JavaScript
Scala
AI/ML
Cursor
Windsurf
DevOps
Windows
Apply
$30k – $50k per year • Equity 0.1–0.5% • Remote (Taiwan, Vietnam) • Full-Time • 2+ years exp
Python
Frontend
GraphQL
DevOps
Windows
Management
Agile
Scrum
QA
Postman
Apply
$40k – $70k per year • Equity 0.1–0.2% • Remote (United States, Canada, Indonesia, Malaysia, Taiwan) • Full-Time • 3+ years exp
Go
AI/ML
Cursor
Windsurf
Apply
$220k – $265k per year • In office • 8+ years exp • Houston
Go
DevOps
Kubernetes
Platform Engineering
Service Mesh
eBPF
DNS
Apply
≈ $221k – $402k per year (Estimated) • In office • New York
Python
AI/ML
InfiniBand
DevOps
SLURM
Kubernetes
HPC
Apply
≈ $161k – $308k per year (Estimated) • In office • 8+ years exp • London
Go
DevOps
Kubernetes
Platform Engineering
Service Mesh
eBPF
DNS
Apply
≈ $164k – $314k per year (Estimated) • In office • London
Python
AI/ML
InfiniBand
DevOps
SLURM
Kubernetes
HPC
Apply
$220k – $330k per year • In office • 8+ years exp • New York
Python
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
Fine-tuning
SGLang
TensorRT
TensorRT-LLM
TRL
Transformers
PyTorch
LLM
Mixture of Experts
CUDA
Triton
OpenAI
PPO
GRPO
Post-training
NVLink
ROCm
CUTLASS
Speculative Decoding
KV Cache
Tool Use
Reward Modeling
DevOps
Kubernetes
Apply
≈ $144k – $265k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Seattle
Python
C++
AI/ML
Vertex AI
TPU
Edge AI
DevOps
GCP
Apply
$140k – $190k per year • Equity • Hybrid • Full-Time • 5+ years exp • PhD • Seattle
Python
JavaScript
SQL
Python
Flask
FastAPI
Databases
Redis
Firestore
Google BigQuery
BigQuery
Frontend
React.js
DevOps
Rest API
gRPC
Terraform
GCP
PagerDuty
Kubernetes
Google Cloud Run
Analytics
A/B Testing
Apply
$200k – $250k per year • Equity • Hybrid • Full-Time • 8+ years exp • PhD • Seattle
Python
Go
JavaScript
SQL
Python
Flask
FastAPI
Databases
Redis
Firestore
Google BigQuery
BigQuery
Frontend
React.js
DevOps
Rest API
gRPC
Terraform
GCP
PagerDuty
CI/CD
Kubernetes
Google Cloud Run
Management
WhatsApp
Agile
Apply
Superintendent 1 hour ago
≈ $91k – $180k per year (Estimated) • In office • Full-Time • 5+ years exp • Seattle
Apply
≈ $89k – $172k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Seattle
Python
Java
SQL
C++
Management
Outlook
Apply
See all jobs
This is one of many
1,249,613 more open roles from verified company boards, updated every day.