967,129open jobs
58,386companies
159,186added this week
Browse all
Salary
≈ $95k – $193k per year (Estimated)
Location
In office (London)
Seniority
Senior

Confirmed on the employer's own hiring board on Sep 30, 2026. First seen by Alion on Sep 6, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About The Role

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

What You'll Be Doing

Platform Operations & Engineering

  • Build and improve shared cloud-native platform capabilities used by internal engineering teams to run AI applications and services.
  • Own significant parts of the platform area, including Kubernetes cluster operations, workload runtime configuration, deployment workflows, observability foundations, or environment automation.
  • Improve the reliability, scalability, and supportability of platform services through practical engineering and operational enhancements.
  • Develop automation, tooling, and configuration that reduce manual effort, improve consistency, and make the platform easier to use and operate.
  • Apply software engineering where it creates leverage, including scripts, services, CI/CD automation, operational tooling, and platform integrations.

Reliability, Operability & Automation

  • Improve incident prevention, detection, response, and recovery across the platform areas you support.
  • Build and refine observability for platform services, including metrics, logs, tracing, dashboards, alerts, and other useful operational signals.
  • Strengthen rollout safety, capacity awareness, failure handling, and recovery procedures for production environments.
  • Debug and resolve complex issues spanning Kubernetes, Linux, networking, storage, workload runtime behaviour, and cloud or datacentre infrastructure dependencies.
  • Enhance operational playbooks, runbooks, and engineering practices to reduce toil and increase service resilience.

Team Technical Contribution

  • Contribute to design discussions, code reviews, and operational standards within the platform engineering team.
  • Collaborate with software engineering, infrastructure, and SRE teams to deliver platform capabilities that are practical, supportable, and aligned to operational needs.
  • Define sensible defaults, paved roads, and supportable patterns for service deployment and runtime operations.
  • Mentor less experienced engineers in platform engineering fundamentals, operational judgement, and good automation practices.

KPIs

  • Platform reliability and service resilience
  • Reduction in manual operational toil
  • Incident detection, response, and recovery effectiveness
  • Observability and operational readiness of platform services

About You

  • Strong hands-on experience operating and improving Kubernetes-based platforms in production.
  • Solid experience with infrastructure automation, CI/CD, configuration management, or GitOps-style workflows.
  • Strong understanding of reliability engineering principles, including observability, incident response, failure analysis, and operational readiness.
  • Experience writing production-quality automation, tooling, or backend code in Go, Python, Bash, or similar languages.
  • Good Linux fundamentals, including processes, filesystems, cgroups, service behaviour, and system debugging.
  • Good networking fundamentals, including TCP/IP, DNS, routing, load balancing, and container or overlay networking concepts.
  • Experience debugging complex production issues across multiple system layers.
  • Ability to work independently on substantial technical problems while collaborating effectively with adjacent teams.
  • Experience mentoring or supporting less experienced engineers through practical technical guidance.

What We Can Offer You

You’ll have the opportunity to help shape the operating standards behind a next-generation AI cloud platform, working on complex infrastructure challenges with real ownership and impact. This is a chance to play a meaningful role in scaling high-performance, sustainable data centre operations in a fast-moving environment.

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
967,129 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
London
≈ $81k – $164k per year (Estimated) • Hybrid • Full-Time • Southampton
Python
C++
MATLAB
MATLAB
Simulink
Apply
≈ $77k – $157k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • London
PowerShell
DevOps
Terraform
VMWare
Azure
Windows Server
Kubernetes
Platform Engineering
SLI/SLO/SLA
Linux
Management
Agile
Apply
≈ $70k – $141k per year (Estimated) • In office • Full-Time • London
Python
Bash
Databases
DynamoDB
DevOps
Terraform
GitHub Actions
CI/CD
Git
AWS
Docker
Configuration Management
AWS Fargate
AWS Lambda
Incident Management
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Linux
Apply
≈ $69k – $141k per year (Estimated) • In office • Full-Time • Richmond
Python
PowerShell
Bash
Databases
OpenSearch
DevOps
Splunk
Terraform
Ansible
GCP
Red Hat
GitHub Actions
CloudFormation
Debian
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Amazon EKS
Azure AKS
AWS Fargate
Amazon EC2
FinOps
Incident Management
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Linux
Cybersecurity
Microsoft Entra ID
Management
Agile
Apply
≈ $74k – $150k per year (Estimated) • In office • Full-Time • Bachelor's Degree • United Kingdom
Python
PowerShell
C#
AI/ML
Ray
DevOps
Azure DevOps
Azure
CI/CD
Incident Management
Apply
In office • 2+ years exp • Bengaluru
Python
Go
DevOps
Rest API
GCP
Azure
CI/CD
Git
AWS
Docker
Linux
Apply
≈ $13k – $25k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
AI/ML
Hadoop
Spark
XGBoost
Reinforcement Learning
Scikit-learn
Prompt Engineering
TensorFlow
PyTorch
RAG
Feature Store
Recommender Systems
Machine Learning
DevOps
CI/CD
Analytics
A/B Testing
Apply
≈ $18k – $39k per year (Estimated) • In office • 4+ years exp • Pune
Python
JavaScript
Ruby
PowerShell
Node JS
DevOps
Terraform
CloudFormation
AWS
AWS Lambda
Management
Agile
Apply
In office • Internship • Bachelor's Degree • Austin
Python
Java
SQL
C++
Apply
In office • Internship • Bachelor's Degree • Vancouver
Python
JavaScript
C
C++
Lua
C
SDL
Game Dev
Godot
Unity
Apply
In office • 3+ years exp • Bachelor's Degree • London
Python
JavaScript
SQL
Apex
Apex
MuleSoft
AI/ML
Edge AI
DevOps
Rest API
Azure
CI/CD
IAM
Cybersecurity
Okta
ISO 27001
SOC 2
Analytics
ETL/ELT
Master Data Management
Management
Zapier
ITSM
Apply
≈ $124k – $278k per year (Estimated) • In office • Seattle
Python
Bash
DevOps
Terraform
Ansible
Kubernetes
Proxmox VE
KVM
OpenStack
Linux
Unix
Apply
≈ $100k – $244k per year (Estimated) • In office • Norway
AI/ML
CUDA Toolkit
CUDA
DevOps
HPC
Apply
In office • 2+ years exp • Singapore
Python
Bash
AI/ML
Edge AI
DevOps
Rest API
GCP
Prometheus
Azure
AWS
Docker
Kubernetes
Grafana
HPC
Linux
Unix
Management
Jira
ServiceNow
ITIL
Apply
≈ $140k – $289k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • New York
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
Fine-tuning
PyTorch
Pre-training
Megatron-LM
NCCL
InfiniBand
NVLink
ROCm
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
SLURM
Docker
Kubernetes
Grafana
Configuration Management
HPC
Linux
TCP/IP
BGP
Apply
$97k – $107k per year • Hybrid • Confidential • Full-Time • Manchester • Bristol • Edinburgh • London
AI/ML
AI Agents
LLM
DevOps
Terraform
GCP
CI/CD
Kubernetes
Service Mesh
Google GKE
IAM
Cybersecurity
Least Privilege
Apply
≈ $98k – $173k per year (Estimated) • In office • Full-Time • London
JavaScript
TypeScript
Node JS
AI/ML
Time Series Forecasting
Frontend
Tailwind CSS
RxJS
React.js
Vite
Radix UI
shadcn/ui
Recharts
esbuild
DevOps
Rest API
Grafana
Apply
≈ $71k – $129k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • London
Analytics
Power BI
Microsoft Excel
Management
Outlook
Apply
≈ $67k – $110k per year (Estimated) • In office • London
Management
Microsoft Office
Apply
≈ $99k – $218k per year (Estimated) • Hybrid • London
Apply
See all jobs
This is one of many
967,129 more open roles from verified company boards, updated every day.