819,440open jobs
52,731companies
133,157added this week
Browse all
Salary
≈ $135k – $274k per year (Estimated)
Location
In office (Houston)
Seniority
Principal · 5+ years exp

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Aug 20, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

Role Overview

As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will be responsible for driving complex, cross-functional programs that ensure the stability, availability, and growth of our cutting-edge GPU fleet and Infiniband network fabrics. This role requires a blend of deep technical understanding, rigorous program management, and a relentless focus on delivering against key operational metrics (SLAs, Uptime, Availability). You will bridge the gap between engineering execution and strategic business goals, directly impacting our ability to serve customer workloads at scale.

Key Responsibilities

  • Program Leadership: Own the planning, execution, and delivery of strategic operational programs, including new data center AI infrastructure build-outs, large-scale fleet software/firmware rollouts, and the implementation of new operational tooling (in partnership with SRE).
  • Metrics and Reporting: Establish, track, and drive accountability against critical infrastructure KPIs, specifically focusing on Availability (Target 97.5%) and Uptime (Target 99%). Develop clear dashboards and communication rhythms to provide leadership with real-time visibility into operational health, program status, and risk.
  • Process Engineering: Analyze and optimize operational workflows across Fleet Operations, Network Operations, and SRE teams. Drive the standardization of incident management, change management, and postmortem processes to reduce toil and improve Mean Time to Recovery (MTTR).
  • Cross-Functional Coordination: Serve as the primary liaison between engineering teams (Hardware, Compute Platform, Network), Data Center Operations, and external vendors (GPU, Network hardware). Proactively identify and resolve dependencies, risks, and roadblocks.
  • Capacity and Readiness: Partner with Data Science/Operation Programs to translate capacity planning models into actionable infrastructure delivery and readiness roadmaps. Ensure that new hardware (GPUs, NICs, switches) is successfully integrated into the operational control plane and meets go-live criteria.
  • Risk Management: Proactively identify technical, schedule, and resource risks related to AI infrastructure scaling and stability. Develop mitigation strategies and communicate impacts clearly to stakeholders.

Required Qualifications

  • Experience: 5+ years of experience in a Technical Program Management role, successfully driving large-scale, complex infrastructure or software engineering programs.
  • Technical Domain Knowledge: Strong foundational understanding of data center infrastructure, distributed systems, Linux, and networking concepts.
  • Program Management Rigor: Proven expertise in modern program management methodologies (Agile, Scrum, PMP certification preferred). Exceptional organizational, communication, and presentation skills.
  • Metrics-Driven Approach: Demonstrable experience in defining, tracking, and improving system performance based on operational metrics (e.g., Uptime, Availability, MTTR, SLOs/SLIs).
  • Execution in Ambiguity: Ability to thrive in a fast-paced, high-growth environment, managing multiple priorities and adapting to evolving technical requirements.

Preferred Qualifications

  • Direct experience managing programs related to data center infrastructure build-outs and hardware commissioning processes.
  • Specific domain knowledge of AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand/RDMA networks, and the challenges of tightly-coupled systems.
  • Experience in a hyperscale or public cloud environment supporting 24/7 mission-critical services.
  • Familiarity with SRE principles, automation tooling, and continuous integration/continuous deployment (CI/CD) pipelines for infrastructure.
  • A Bachelor's or Master's degree in a technical field (Computer Science, Engineering, etc.) or equivalent practical experience.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package (base + equity) with reviews every 12 months. 
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. 
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
819,440 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Houston
≈ $143k – $254k per year (Estimated) • Remote (United States) • Full-Time • 10+ years exp • Bachelor's Degree • Indianapolis
Management
Agile
Apply
≈ $128k – $260k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Jose
Python
Java
PHP
C++
DevOps
Self-Healing
Linux
Apply
$195k – $285k per year • Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
InfiniBand
NVLink
DevOps
Splunk
Terraform
Ansible
GCP
Datadog
Prometheus
SLURM
Azure
AWS
Kubernetes
Grafana
Self-Healing
Progressive Delivery
FinOps
AIOps
SLI/SLO/SLA
HPC
Linux
Management
ITIL
Apply
≈ $143k – $278k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Sunnyvale
Apply
$183k – $247k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Boston
Databases
Databricks
AI/ML
AI Agents
AWS Bedrock
Amazon SageMaker
AWS Bedrock AgentCore
Machine Learning
DevOps
AWS
AWS Step Functions
Robotics
Digital Twin
IoT
OPC UA
Apply
$121k – $201k per year • Hybrid • 8+ years exp • Bachelor's Degree • San Diego
Python
SQL
Databases
SAP HANA
Snowflake
AI/ML
Dagster
dbt
Prefect
Supervision
DevOps
Terraform
Ansible
CloudFormation
CI/CD
Git
AWS
Platform Engineering
AWS Lambda
Amazon S3
Analytics
Dimensional Modeling
Management
Jira
Agile
Scrum
Microsoft Office
Apply
$112k – $147k per year • Remote (likely United States) • 2+ years exp
DevOps
Rest API
Management
Agile
Scrum
Apply
$137k – $244k per year • Remote (United States) • Public Trust • 10+ years exp • Bachelor's Degree
SQL
Analytics
Tableau
Power BI
Management
Confluence
Jira
Agile
Scrum
Apply
Software Developer 1 day ago
$115k – $190k per year • In office • TS/SCI • 4+ years exp • Bachelor's Degree • Quantico
Python
Java
SQL
C#
C++
Game Dev
Unity
Unreal Engine
Management
Agile
Scrum
Apply
$136k – $226k per year • Hybrid • 10+ years exp • Bachelor's Degree • San Diego
Python
ABAP
ABAP
SAP Fiori
SAP BTP
ABAP Objects
DevOps
Rest API
Ansible
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Incident Management
GitHub
SOAP
Apply
≈ $140k – $286k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Houston
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
Fine-tuning
PyTorch
Pre-training
Megatron-LM
NCCL
InfiniBand
NVLink
ROCm
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
SLURM
Docker
Kubernetes
Grafana
Configuration Management
HPC
Linux
TCP/IP
BGP
Apply
$160k – $230k per year • In office • 5+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
Apply
≈ $153k – $299k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Houston
Python
Java
C++
AI/ML
CUDA Toolkit
Prefect
CUDA
NCCL
InfiniBand
NVLink
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
Pulumi
SLURM
CI/CD
Kubernetes
Grafana
OpenStack
HPC
Linux
Apply
≈ $138k – $266k per year (Estimated) • In office • 5+ years exp • Seattle
Python
PowerShell
AI/ML
Gemini
Cybersecurity
Okta
ISO 27001
SOC 2
DLP
Management
Google Workspace
Gmail
Apply
$190k – $260k per year • In office • 6+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
SLURM
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
HPC
Apply
$166k – $200k per year • Remote (United States) • 20+ years exp • Bachelor's Degree • Houston
JavaScript
Frontend
npm
Apply
≈ $105k – $237k per year (Estimated) • Remote (United States) • Bachelor's Degree • Houston
Apply
Simulator Engineer 3 6 hours ago
$87k – $105k per year • Remote (United States) • 5+ years exp • Bachelor's Degree • Houston
Python
C++
Fortran
DevOps
Git
GitHub
GitLab
Apply
Project Manager 4 6 hours ago
$124k – $150k per year • In office • 10+ years exp • Bachelor's Degree • Houston
Apply
Project Manager 3 6 hours ago
$99k – $120k per year • In office • 5+ years exp • Bachelor's Degree • Houston
Apply
See all jobs
This is one of many
819,440 more open roles from verified company boards, updated every day.