893,443open jobs
55,503companies
149,331added this week
Browse all
Salary
≈ $140k – $289k per year (Estimated)
Location
In office (Houston)
Seniority
Principal · 10+ years exp

Confirmed on the employer's own hiring board on Sep 28, 2026. First seen by Alion on Sep 3, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

Overview

As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs.

Key Responsibilities

  • Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs - open-source and customer-representative models - across thousands of accelerators to validate cluster behavior under genuine load rather than synthetic proxies alone, and translate what those runs reveal into fleet-wide fixes.
  • Lead deep diagnosis of large-scale cluster failures and performance regressions, isolating root cause across the full stack: GPU and NIC firmware, PCIe/NVLink topology and NUMA placement, InfiniBand/RoCE fabric health, congestion control and routing, storage and data-loader throughput, scheduler placement, and framework/communication-library behavior. Serve as the final escalation point for the hardest slow-job and stalled-job investigations.
  • Design and build the validation and burn-in systems that qualify nodes, racks, and full pods at scale - NCCL/RCCL collective sweeps, HPL/HPCG and MLPerf-style benchmarks, thermal and power soak tests, straggler and flapping-link detection - and automate them so that qualification is a repeatable pipeline, not a manual campaign.
  • Drive cluster optimization end to end, tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput.
  • Partner with Infrastructure, Platform, SRE, and customer-facing teams to translate operational and customer performance needs into durable engineering solutions, and to feed diagnostic signal back into provisioning, remediation, and capacity workflows.
  • Build production-grade Python systems and performance tooling for automated triage, telemetry correlation, and regression detection, leveraging AI tools to accelerate delivery. Assess impact to the team's software and validation stack from new hardware product programs, and explore AI-driven process improvement and automation.
  • Establish engineering standards for reliability, observability, benchmarking methodology, and operational excellence across all services, and raise the diagnostic capability of the wider organization through mentorship, runbooks, and post-incident technical write-ups.

Required Qualifications

  • Education: Bachelor's or higher degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Experience: 10+ years of relevant experience building, operating, or debugging large-scale compute infrastructure, including significant time at staff or principal level owning cross-team technical direction.
  • AI Workload Expertise: Hands-on experience running real AI compute jobs at scale - pre-training, fine-tuning, or large-scale inference of open-source or proprietary models - including practical familiarity with distributed training strategies (data, tensor, pipeline, and expert parallelism) and frameworks such as PyTorch, Megatron-LM, DeepSpeed, or equivalent.
  • Cluster Validation: Demonstrated experience validating and accepting large clusters (thousands of GPUs) for performance and reliability, with a working command of benchmark methodology and the ability to defend a number to both engineers and customers.
  • Performance Debugging: Proven ability to diagnose distributed performance problems - stragglers, collective stalls, link flaps, thermal throttling, silent data corruption, ECC and Xid errors, noisy-neighbor and storage-bound bottlenecks - using tools such as NCCL debug tracing, Nsight Systems/Compute, PyTorch Profiler, perf, and fabric telemetry.
  • Networking: Deep understanding of high-performance fabrics - InfiniBand and/or RoCEv2, RDMA, GPUDirect, adaptive routing, congestion control, and rail-optimized topologies - and of networking fundamentals (TCP/IP, BGP).
  • Systems & Programming: Deep Linux systems expertise (kernel tunables, NUMA, PCIe, IRQ and memory behavior) and strong production Python, plus experience with C/C++ or Go and with configuration management tooling (e.g., Ansible, Terraform).
  • Schedulers: Experience operating and debugging AI workloads under SLURM and/or Kubernetes at scale.

Preferred Qualifications

  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience bringing up and qualifying a greenfield GPU supercluster from first rack to production traffic, including firmware, driver, and topology standardization across a heterogeneous fleet.
  • Direct experience with NVIDIA GPU platforms (H200/GB200/GB300-class), NVLink and NVSwitch domains, DCGM, SHARP, UFM, and the NVIDIA software stack; or equivalent depth on AMD Instinct and ROCm/RCCL.
  • Published or presented benchmark, scaling, or post-mortem work - MLPerf submissions, scaling studies, or public technical write-ups on large-cluster behavior.
  • Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) applied to high-cardinality GPU and fabric telemetry, including automated anomaly and regression detection.
  • Experience with high-throughput parallel storage (Lustre, GPFS, WEKA, VAST) and with diagnosing data-pipeline-bound training jobs.
  • Familiarity with cloud-native technologies (Kubernetes, Docker), infrastructure-as-code principles, and integration with infrastructure tooling such as DCIMs, NetBox, and bare metal APIs (MAAS, Ironic, IPMI, Redfish).
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
  • Familiarity with SLOs/metrics measurement and logs/telemetry/metrics integration with tools for enhanced operator experience.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
893,443 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Houston
≈ $122k – $267k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Coppell
DevOps
Azure DevOps
GitHub Actions
Azure
CI/CD
Platform Engineering
Shift-Left
GitHub
Cybersecurity
Shift-Left Security
Apply
$161k – $241k per year • In office • Secret • Full-Time • Bachelor's Degree • Colorado Springs
DevOps
Red Hat
VMWare
Linux
Management
Confluence
Apply
$94k – $141k per year • In office • Secret • Full-Time • 8+ years exp • Bachelor's Degree • Colorado Springs
DevOps
Red Hat
Windows Server
Linux
Cybersecurity
Active Directory
LDAP
Apply
$130k – $140k per year • Remote (United States) • 10+ years exp • Dallas
Python
Databases
MySQL
Oracle
MS SQL
DevOps
Splunk
Ansible
Zabbix
OpenShift
Helm
Prometheus
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Configuration Management
Bitbucket
Amazon EC2
GitLab
Amazon S3
IAM
Amazon CloudWatch
Apply
$131k – $175k per year • In office • Confidential • Full-Time • Master's Degree • Taylor
Python
SQL
Robotics
Digital Twin
Analytics
Power BI
Apply
≈ $119k – $265k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Dublin
Python
C++
AI/ML
Claude
Claude Code
vLLM
CUDA Toolkit
Function Calling
SGLang
TensorRT
TensorRT-LLM
LLM
CUDA
Triton
OpenAI
OpenAI Codex
CUTLASS
Speculative Decoding
KV Cache
Tool Use
Machine Learning
Apply
≈ $64k – $99k per year (Estimated) • Equity • Remote (Poland) • Full-Time • Warsaw
Python
Databases
Databricks
AI/ML
Anthropic
DevOps
Rest API
Terraform
GCP
Azure
CI/CD
AWS
Platform Engineering
Cybersecurity
Microsoft Entra ID
Management
Agile
Scrum
Apply
≈ $110k – $243k per year (Estimated) • Equity • Remote (Poland) • Full-Time • 5+ years exp
Python
SQL
Databases
Snowflake
AI/ML
Embeddings
AI Agents
AWS Bedrock
RAG
Anthropic
LLM Guardrails
DevOps
GCP
Prometheus
Azure
Git
AWS
Platform Engineering
Cortex
Management
Agile
Apply
$53k – $90k per year • Equity • Remote (Poland) • Full-Time • 8+ years exp • Katowice
Python
Python
Pydantic
Databases
Databricks
AI/ML
MLFlow
Anthropic
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
GitLab
IAM
Management
Agile
Apply
≈ $121k – $252k per year (Estimated) • Equity • Remote (Poland) • Full-Time
AI/ML
LangChain
Semantic Kernel
OpenAI
Anthropic
DevOps
Terraform
GCP
Azure DevOps
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Azure AKS
FinOps
GitHub
Management
Agile
Apply
$160k – $230k per year • In office • 5+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
Apply
≈ $135k – $278k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Houston
AI/ML
InfiniBand
Edge AI
DevOps
CI/CD
Incident Management
HPC
Linux
Management
Agile
Scrum
Apply
≈ $151k – $300k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Houston
Python
Java
C++
AI/ML
CUDA Toolkit
Prefect
CUDA
NCCL
InfiniBand
NVLink
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
Pulumi
SLURM
CI/CD
Kubernetes
Grafana
OpenStack
HPC
Linux
Apply
≈ $123k – $239k per year (Estimated) • In office • 5+ years exp • Houston
Python
PowerShell
AI/ML
Gemini
Cybersecurity
Okta
ISO 27001
SOC 2
DLP
Management
Google Workspace
Gmail
Apply
$190k – $260k per year • In office • 6+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
SLURM
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
HPC
Apply
≈ $93k – $202k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Houston
DevOps
VPN
BGP
OSPF
Apply
$96k – $144k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • New York • Pasadena • San Francisco • Los Angeles • Portland
Analytics
Microsoft Excel
Apply
$126k – $182k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Pasadena • Portland • Frisco • Irvine
Analytics
Microsoft Excel
Apply
≈ $60k – $114k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • El Paso • Houston • Dallas
SQL
Databases
Oracle
Mobile
Maestro
DevOps
Linux
Unix
Analytics
Talend
Apply
≈ $71k – $178k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Louisville • New York • Allen • Houston • Salt Lake City
Apply
See all jobs
This is one of many
893,443 more open roles from verified company boards, updated every day.