1,338,580open jobs
78,412companies
206,623added this week
Browse all
Salary
≈ $161k – $308k per year (Estimated)
Location
In office (London)
Seniority
Staff · 8+ years exp
Visa
Licensed UK visa sponsor

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Sep 6, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

We're hiring a Staff Cloud Native Software Engineer to build, operate, and improve the cloud-native software integrations that connect AI applications and networking components at scale.

In this software engineering role, you'll work on shared Kubernetes-based platforms, deployment patterns, observability foundations, infrastructure architecture, and operational tooling that help internal teams run services safely and efficiently on GPU-backed infrastructure. You'll partner closely with platform engineering, infrastructure, and product teams to ensure capabilities meet real developer and operational needs.

This role is important to the reliability, scalability, and usability of Nscale's software integrations. As a Staff engineer, you'll take ownership of significant components and set technical direction across teams, deliver complex technical work independently, and raise the quality of operations and engineering through practical improvements, sound technical judgement, and mentoring.

What You’ll Do

Cloud Native Software Engineering

  • Design, build, and operate Kubernetes-native software - controllers, operators, custom resources (CRDs), and admission webhooks - that connects AI applications with core networking components on GPU-backed infrastructure.
  • Extend Kubernetes control-plane capabilities to support AI workload requirements, including network policy controllers, CNI/service-mesh integrations, and resource/scheduling extensions.
  • Own significant components end-to-end and set the technical direction for how they're designed, deployed, and operated across the team.
  • Build reconciliation loops, informers, and client-go-based tooling that keep infrastructure state consistent between the API server, networking systems, and AI runtime components.
  • Develop operational tooling and automation that make Kubernetes-native services easier for internal teams to deploy, run, and support.

Infrastructure Architecture, Reliability & Observability

  • Drive infrastructure architecture decisions around how AI applications and networking components integrate across the platform, weighing trade-offs at a cross-team level.
  • Build observability foundations for controller and operator software - metrics, structured events, tracing, and status reporting surfaced through the Kubernetes API and platform dashboards.
  • Design systems that degrade gracefully and self-heal, using controller patterns (reconciliation, backoff, status conditions) to reduce manual intervention.
  • Debug and resolve complex issues spanning the Kubernetes control plane, networking (CNI, service mesh, kube-proxy/eBPF datapaths), and workload runtime behavior on GPU-backed infrastructure.
  • Define standards for safe rollout of controller and platform changes, including versioning, compatibility, and staged deployment.

Team Technical Leadership

  • Set technical direction for how the team builds Kubernetes-native software, establishing patterns for controller design, CRD schema evolution, and testing strategy.
  • Lead design discussions and code reviews, holding a high bar for Kubernetes API conventions and idiomatic client-go usage.
  • Partner with platform engineering, infrastructure, and product teams to translate real developer and operational needs into clean CRDs, APIs, and controller-managed abstractions.
  • Define reusable patterns, shared libraries, and scaffolding that let other teams build correctly on the platform without reinventing integration logic.
  • Mentor engineers in Kubernetes internals, controller-runtime patterns, and sound operational judgement.

KPIs

  • Reliability, scalability, and usability of AI infrastructure-networking software integrations
  • Correctness and maintainability of Kubernetes controllers and operators in production
  • Reduction in manual operational effort and config drift across supported components
  • Adoption of shared patterns/frameworks and effectiveness of observability tooling across teams

Who You are

  • At least 8 years of experience in production-level software development.
  • Deep hands-on experience building and operating Kubernetes-native software: custom controllers, operators, CRDs, or admission webhooks - using controller-runtime, client-go, or equivalent.
  • Strong understanding of Kubernetes internals: the API server, informer/lister patterns, reconciliation loops, and the object model.
  • Strong networking fundamentals - CNI, service mesh, kube-proxy/eBPF datapaths, DNS, load balancing - and experience building software that integrates with these systems.
  • Proficiency in Go (strongly preferred) or a similar language, with a track record of shipping well-tested, production-quality code at scale.
  • Experience with observability practices - metrics, tracing, structured logging - built into software rather than added afterward.
  • Comfortable owning components independently end-to-end, from design through operation, while setting direction for adjacent teams.
  • Experience with or strong interest in GPU-backed infrastructure and AI workload patterns is a plus.
  • Track record of leading technical design at a staff level and mentoring engineers through practical technical guidance.

What We Can Offer You

You’ll have the opportunity to help shape the operating standards behind a next-generation AI cloud platform, working on complex infrastructure challenges with real ownership and impact. This is a chance to play a meaningful role in scaling high-performance, sustainable data centre operations in a fast-moving environment.

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,338,580 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
London
≈ $67k – $135k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Belfast
Java
SQL
Java
Spring Boot
Databases
Apache Kafka
AI/ML
Copilot
Cursor
Machine Learning
DevOps
Rest API
Azure
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Apply
≈ $91k – $178k per year (Estimated) • In office • Full-Time • 6+ years exp
JavaScript
TypeScript
Node JS
AI/ML
Agentic Workflows
Frontend
React.js
Management
Agile
Apply
$180k per year • In office • Contractor • 10+ years exp • London
Python
Java
Databases
KDB+
Apache Kafka
AI/ML
Time Series Forecasting
Apply
$116k – $147k per year • Hybrid • Full-Time • London
JavaScript
Java
TypeScript
Java
Spring Framework
AI/ML
Machine Learning
Frontend
React.js
DevOps
CI/CD
Apply
≈ $88k – $176k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • London
SQL
C#
Databases
Redis
Databricks
DynamoDB
RabbitMQ
Apache Kafka
DevOps
CloudFormation
CI/CD
AWS
Kubernetes
Amazon EventBridge
Analytics
Tableau
Apply
≈ $40k – $108k per year (Estimated) • Remote (Canada) • Montreal
Python
Java
Databases
PostgreSQL
DevOps
Datadog
Git
AWS
Kubernetes
Platform Engineering
Linux
QA
JMeter
Playwright
Apply
$122k – $214k per year • Remote (United States) • Full-Time • Bachelor's Degree
Python
AI/ML
Multimodal AI
AI Agents
LLMOps
TPU
Edge AI
Machine Learning
DevOps
GCP
Azure
AWS
Kubernetes
Google GKE
SLI/SLO/SLA
Apply
In office • 4+ years exp • Bachelor's Degree
Rust
Scala
Databases
PostgreSQL
Apache Kafka
DevOps
CI/CD
Docker
Kubernetes
Apply
≈ $24k – $50k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Hyderabad
AI/ML
Model Context Protocol
AI Agents
LLM
LLM Guardrails
Frontend
GraphQL
DevOps
gRPC
Platform Engineering
Apply
In office
Java
Java
Spring Boot
DevOps
Rest API
Helm
Jenkins
Git
Kubernetes
Apply
≈ $164k – $314k per year (Estimated) • In office • London
Python
AI/ML
InfiniBand
DevOps
SLURM
Kubernetes
HPC
Apply
$240k – $400k per year • In office • 12+ years exp
Python
AI/ML
Cursor
Claude
Prefect
InfiniBand
DevOps
Terraform
GCP
Pulumi
AWS
Kubernetes
Self-Healing
OpenStack
HPC
Apply
$290k – $520k per year • In office • 10+ years exp • New York
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
Quantization
Knowledge Distillation
SGLang
TensorRT
TensorRT-LLM
TRL
Transformers
LLM
Mixture of Experts
CUDA
Triton
DPO
GRPO
Post-training
Megatron-LM
ROCm
Speculative Decoding
KV Cache
Tool Use
Model Distillation
Reward Modeling
DevOps
SLI/SLO/SLA
Apply
$220k – $330k per year • In office • 8+ years exp • New York
Python
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
Fine-tuning
SGLang
TensorRT
TensorRT-LLM
TRL
Transformers
PyTorch
LLM
Mixture of Experts
CUDA
Triton
OpenAI
PPO
GRPO
Post-training
NVLink
ROCm
CUTLASS
Speculative Decoding
KV Cache
Tool Use
Reward Modeling
DevOps
Kubernetes
Apply
$220k – $265k per year • In office • 8+ years exp • Houston
Go
DevOps
Kubernetes
Platform Engineering
Service Mesh
eBPF
DNS
Apply
≈ $43k – $110k per year (Estimated) • In office • Full-Time • London • Edinburgh
Apply
≈ $62k – $139k per year (Estimated) • In office • Full-Time • 4+ years exp • London • Manchester • Bristol
Marketing
Instagram
Apply
≈ $67k – $116k per year (Estimated) • Hybrid • Full-Time • London
Python
SQL
Python
Beautiful Soup
Databases
PostgreSQL
Databricks
MS SQL
Amazon Redshift
AI/ML
Airflow
Prompt Engineering
TensorFlow
Pandas
PyTorch
RAG
Time Series Forecasting
Context Engineering
Machine Learning
DevOps
Azure DevOps
Azure
CI/CD
Jenkins
Git
AWS
Platform Engineering
Configuration Management
AWS Lambda
Amazon EC2
Amazon S3
IAM
Linux
Analytics
Tableau
Power BI
ETL/ELT
Management
Agile
Scrum
Kanban
QA
Selenium
Apply
≈ $73k – $137k per year (Estimated) • Hybrid • Full-Time • London
SQL
Management
Agile
Apply
$57k per year • Hybrid • Full-Time • London
Python
DevOps
CI/CD
GitHub
GitLab
Management
Confluence
Jira
Agile
Waterfall
Service Desk
QA
Selenium
Robot Framework
Apply
See all jobs
This is one of many
1,338,580 more open roles from verified company boards, updated every day.