825,970open jobs
53,207companies
135,769added this week
Browse all
Salary
≈ $153k – $299k per year (Estimated)
Location
In office (Houston)
Seniority
Staff · 8+ years exp

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Aug 14, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers.  Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

We're hiring a Staff Software Engineer to build the software, automation, and control-plane capabilities that manage Nscale's fleet of AI infrastructure at scale. Your work will improve the acceptance, performance, and scalability of our AI and high-performance computing environments - driving higher availability, faster capacity delivery, and lower operational load as Nscale grows into one of the world's leading neo-cloud providers.

This is a senior individual-contributor role for an engineer who enjoys solving hard infrastructure problems at the intersection of software, GPUs, networking, and large-scale operations. You will have the autonomy to investigate problems, learn quickly, innovate, and deliver improvements wherever they create meaningful impact for the team and the platform.

You will work closely with teams across Nscale - including Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware engineering - to translate operational challenges into reliable, scalable software. You will not need to own every component to make a difference: strong engineers identify opportunities, build a compelling case for a solution, and work with the right partners to deliver it.

NOTE: We are hiring for various senior experience levels. The final leveling for the role will be based on your overall work experience, experience in AI Infra domain and interview feedback.

What You'll Be Doing

  • Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems, balancing scalability, reliability, and maintainability.
  • Build and operate production-grade software, services, APIs, and automation that manage the lifecycle of GPU compute and supporting network infrastructure.
  • Own end-to-end workflows for fleet inventory, provisioning, configuration, hardware and firmware lifecycle management, validation, health monitoring, remediation, capacity, and reliability at scale.
  • Investigate complex production issues across hardware, GPUs, operating systems, networks, schedulers, and services; turn findings into durable software improvements rather than recurring manual work.
  • Build safe, observable, and auditable control-plane workflows that give operators clear visibility and reliable ways to act.
  • Establish engineering standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness. Use SLOs, telemetry, alerting, and postmortems to drive continuous improvement.
  • Partner with Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware teams to translate operational needs into robust, scalable automation.
  • Influence the evolution of adjacent systems and services through sound technical judgment, clear communication, and practical solutions.
  • Assess the impact of new hardware programmes on the software stack and ensure fleet-management capabilities are ready to support them.
  • Lead technical design reviews and incident deep-dives; mentor other engineers and raise the engineering bar across the organization.
  • Use AI-assisted development tools to increase delivery leverage while maintaining a high bar for correctness, security, and operational safety.

About You

  • 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems.
  • A Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
  • A strong software-engineering foundation in Python and/or Go, Java, C++, or similar languages, including API design, testing, code review, and production debugging.
  • Deep understanding of Linux, distributed systems, networking fundamentals, and systems performance; you are comfortable working across stateful and stateless services.
  • Experience designing and operating reliable automation or control-plane systems for complex infrastructure, large fleets, cloud platforms, or hardware lifecycle management.
  • Proven ability to take ambiguous technical problems from architecture through implementation and production operation, while influencing peers and stakeholders without relying on formal authority.
  • Hands-on experience with observability, monitoring, metrics, logs, tracing, alerting, incident response, capacity planning, and performance analysis.
  • Strong communication skills and sound technical judgment. You can explain trade-offs clearly, build alignment, and move work forward in a fast-changing environment.
  • A curious, pragmatic, high-ownership mindset. You enjoy finding the underlying cause of difficult problems and building the simplest durable solution.

Strong Candidates Will Have

  • Direct experience with AI, GPU, HPC, or large-scale cloud infrastructure, including NVIDIA GPUs, CUDA, NVLink/NVSwitch, NCCL, and workload schedulers such as Slurm and Kubernetes.
  • Experience with high-performance datacentre networking, including InfiniBand, RoCE, Ethernet fabrics, routing, congestion control, topology-aware systems, or GPU Direct RDMA.
  • Experience with bare-metal lifecycle automation and infrastructure management tools such as Redfish, IPMI, PXE, MAAS, Ironic, NetBox, DCIM, OpenStack, or equivalent systems.
  • Experience with workflow orchestration and reliable automation systems such as Temporal, Airflow, Prefect, or event-driven architectures.
  • Experience with Kubernetes, containers, infrastructure as code (Terraform, Pulumi, Ansible), and public-cloud or private-cloud platforms.
  • Experience with observability platforms and high-cardinality telemetry, such as Prometheus, Grafana, OpenTelemetry, ELK, or equivalent.
  • Experience with hardware qualification, burn-in, validation, fleet health, or automated remediation for servers, GPUs, or network equipment.
  • A track record of technical leadership: setting direction, defining reusable patterns, developing other engineers, and improving the effectiveness of multiple teams.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package (base + equity) with reviews every 12 months. 
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. 
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support. 

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
825,970 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Houston
≈ $123k – $237k per year (Estimated) • In office • New York
AI/ML
Claude
ChatGPT
Claude Code
AI Agents
OpenAI Codex
DevOps
Terraform
GCP
Azure DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
GitHub
GitLab
Management
Slack
Google Drive
Apply
$86k – $130k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Wilmington
Analytics
Power BI
Design
SolidWorks
AutoCAD
Apply
≈ $93k – $180k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Plano
PowerShell
Databases
Azure SQL Database
DevOps
Terraform
Azure DevOps
CloudFormation
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Platform Engineering
Bicep
Amazon EKS
Azure AKS
AWS Lambda
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
API Gateway
Apply
≈ $93k – $180k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Plano
PowerShell
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Platform Engineering
Bicep
Incident Management
GitHub
VPN
Apply
$185k – $258k per year • In office • Full-Time • 8+ years exp • Washington • London
DevOps
Azure
Platform Engineering
Management
ServiceNow
Agile
Apply
≈ $96k – $193k per year (Estimated) • In office • Hong Kong
Python
Go
Java
Move
C++
AI/ML
AI Agents
LLM
Recommender Systems
Web3
Smart Contracts
Apply
≈ $71k – $186k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Singapore
Python
PowerShell
Bash
DevOps
Terraform
Ansible
CI/CD
Git
AWS
Docker
Kubernetes
SRE
Platform Engineering
Configuration Management
Linux
Apply
≈ $67k – $179k per year (Estimated) • In office • Full-Time • Seoul
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Quantization
Computer Vision
TensorFlow
PyTorch
Apply
≈ $43k – $74k per year (Estimated) • In office • Internship • New York
Python
Java
SQL
Analytics
Tableau
Apply
Remote (likely United States) • Bachelor's Degree
Python
SQL
Bash
Databases
MySQL
DevOps
Ansible
Debian
Proxmox VE
Linux
Cybersecurity
pfSense
Apply
≈ $140k – $286k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Houston
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
Fine-tuning
PyTorch
Pre-training
Megatron-LM
NCCL
InfiniBand
NVLink
ROCm
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
SLURM
Docker
Kubernetes
Grafana
Configuration Management
HPC
Linux
TCP/IP
BGP
Apply
$160k – $230k per year • In office • 5+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
Apply
≈ $135k – $274k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Houston
AI/ML
InfiniBand
Edge AI
DevOps
CI/CD
Incident Management
HPC
Linux
Management
Agile
Scrum
Apply
≈ $138k – $266k per year (Estimated) • In office • 5+ years exp • Seattle
Python
PowerShell
AI/ML
Gemini
Cybersecurity
Okta
ISO 27001
SOC 2
DLP
Management
Google Workspace
Gmail
Apply
$190k – $260k per year • In office • 6+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
SLURM
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
HPC
Apply
$166k – $200k per year • Remote (United States) • 20+ years exp • Bachelor's Degree • Houston
JavaScript
Frontend
npm
Apply
≈ $105k – $237k per year (Estimated) • Remote (United States) • Bachelor's Degree • Houston
Apply
$87k – $105k per year • Remote (United States) • 5+ years exp • Bachelor's Degree • Houston
Python
C++
Fortran
DevOps
Git
GitHub
GitLab
Apply
Project Manager 4 1 day ago
$124k – $150k per year • In office • 10+ years exp • Bachelor's Degree • Houston
Apply
Project Manager 3 1 day ago
$99k – $120k per year • In office • 5+ years exp • Bachelor's Degree • Houston
Apply
See all jobs
This is one of many
825,970 more open roles from verified company boards, updated every day.