1,264,521open jobs
72,980companies
210,977added this week
Browse all
Salary
$220k – $265k per year
Location
In office (Houston)
Seniority
Staff · 8+ years exp
Visa
H-1B filings in 12 months: 10 · for this role: 3

Confirmed on the employer's own hiring board on Oct 6, 2026. First seen by Alion on Aug 27, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

We're hiring a Staff Cloud Native Software Engineer to build, operate, and improve the cloud-native software integrations that connect AI applications and networking components at scale.

In this software engineering role, you'll work on shared Kubernetes-based platforms, deployment patterns, observability foundations, infrastructure architecture, and operational tooling that help internal teams run services safely and efficiently on GPU-backed infrastructure. You'll partner closely with platform engineering, infrastructure, and product teams to ensure capabilities meet real developer and operational needs.

This role is important to the reliability, scalability, and usability of Nscale's software integrations. As a Staff engineer, you'll take ownership of significant components and set technical direction across teams, deliver complex technical work independently, and raise the quality of operations and engineering through practical improvements, sound technical judgement, and mentoring.

What You’ll Do

Cloud Native Software Engineering

  • Design, build, and operate Kubernetes-native software - controllers, operators, custom resources (CRDs), and admission webhooks - that connects AI applications with core networking components on GPU-backed infrastructure.
  • Extend Kubernetes control-plane capabilities to support AI workload requirements, including network policy controllers, CNI/service-mesh integrations, and resource/scheduling extensions.
  • Own significant components end-to-end and set the technical direction for how they're designed, deployed, and operated across the team.
  • Build reconciliation loops, informers, and client-go-based tooling that keep infrastructure state consistent between the API server, networking systems, and AI runtime components.
  • Develop operational tooling and automation that make Kubernetes-native services easier for internal teams to deploy, run, and support.

Infrastructure Architecture, Reliability & Observability

  • Drive infrastructure architecture decisions around how AI applications and networking components integrate across the platform, weighing trade-offs at a cross-team level.
  • Build observability foundations for controller and operator software - metrics, structured events, tracing, and status reporting surfaced through the Kubernetes API and platform dashboards.
  • Design systems that degrade gracefully and self-heal, using controller patterns (reconciliation, backoff, status conditions) to reduce manual intervention.
  • Debug and resolve complex issues spanning the Kubernetes control plane, networking (CNI, service mesh, kube-proxy/eBPF datapaths), and workload runtime behavior on GPU-backed infrastructure.
  • Define standards for safe rollout of controller and platform changes, including versioning, compatibility, and staged deployment.

Team Technical Leadership

  • Set technical direction for how the team builds Kubernetes-native software, establishing patterns for controller design, CRD schema evolution, and testing strategy.
  • Lead design discussions and code reviews, holding a high bar for Kubernetes API conventions and idiomatic client-go usage.
  • Partner with platform engineering, infrastructure, and product teams to translate real developer and operational needs into clean CRDs, APIs, and controller-managed abstractions.
  • Define reusable patterns, shared libraries, and scaffolding that let other teams build correctly on the platform without reinventing integration logic.
  • Mentor engineers in Kubernetes internals, controller-runtime patterns, and sound operational judgement.

KPIs

  • Reliability, scalability, and usability of AI infrastructure-networking software integrations
  • Correctness and maintainability of Kubernetes controllers and operators in production
  • Reduction in manual operational effort and config drift across supported components
  • Adoption of shared patterns/frameworks and effectiveness of observability tooling across teams

Who You are

  • At least 8 years of experience in production-level software development.
  • Deep hands-on experience building and operating Kubernetes-native software: custom controllers, operators, CRDs, or admission webhooks - using controller-runtime, client-go, or equivalent.
  • Strong understanding of Kubernetes internals: the API server, informer/lister patterns, reconciliation loops, and the object model.
  • Strong networking fundamentals - CNI, service mesh, kube-proxy/eBPF datapaths, DNS, load balancing - and experience building software that integrates with these systems.
  • Proficiency in Go (strongly preferred) or a similar language, with a track record of shipping well-tested, production-quality code at scale.
  • Experience with observability practices - metrics, tracing, structured logging - built into software rather than added afterward.
  • Comfortable owning components independently end-to-end, from design through operation, while setting direction for adjacent teams.
  • Experience with or strong interest in GPU-backed infrastructure and AI workload patterns is a plus.
  • Track record of leading technical design at a staff level and mentoring engineers through practical technical guidance.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000—$265,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,264,521 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Houston
$102k – $202k per year • In office • 5+ years exp • Bachelor's Degree • Redmond
Python
JavaScript
C#
C++
DevOps
GCP
Azure
CI/CD
AWS
Apply
$120k – $235k per year • In office • 5+ years exp • Bachelor's Degree • Redmond
Python
JavaScript
C#
C++
DevOps
Platform Engineering
Management
Outlook
Apply
≈ $142k – $287k per year (Estimated) • In office • 12+ years exp • Bachelor's Degree • Redmond
Python
JavaScript
C#
C++
AI/ML
AI Agents
LLM
Edge AI
DevOps
GCP
Azure
CI/CD
AWS
Cybersecurity
Threat Modeling
Apply
$120k – $235k per year • In office • 8+ years exp • Bachelor's Degree • Washington • Atlanta • San Antonio • Columbia • Phoenix
Python
JavaScript
TypeScript
C#
C++
C#
ASP.NET Core
Blazor
Databases
Microsoft Fabric
AI/ML
Copilot
Function Calling
Semantic Kernel
RAG
OpenAI
Structured Outputs
Context Engineering
Multi-Agent Systems
Tool Use
Frontend
Vue.js
Angular
React.js
DevOps
Rest API
Azure
CI/CD
Cybersecurity
Least Privilege
Microsoft Entra ID
Apply
$120k – $235k per year • In office • 8+ years exp • Bachelor's Degree • Redmond
Python
JavaScript
Rust
C#
C++
AI/ML
Fine-tuning
DevOps
Helm
Istio
KEDA
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Platform Engineering
Bicep
Azure AKS
Apply
Tech Lead 2 days ago
≈ $32k – $83k per year (Estimated) • In office • Cape Town
Python
Go
JavaScript
Java
TypeScript
SQL
Node JS
Databases
RabbitMQ
Apache Kafka
Frontend
Vue.js
GraphQL
Angular
React.js
DevOps
Rest API
gRPC
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
Management
Agile
Apply
≈ $18k – $51k per year (Estimated) • In office • Cape Town
JavaScript
C#
C#
.NET
DevOps
Rest API
Azure
CI/CD
Platform Engineering
Management
Power Automate
Power Apps
Agile
Apply
≈ $72k – $147k per year (Estimated) • In office • 2+ years exp • Morrisville
AI/ML
CUDA Toolkit
ClearML
CUDA
InfiniBand
DevOps
Docker
Kubernetes
Apply
≈ $53k – $97k per year (Estimated) • In office • 5+ years exp • Singapore
Python
Bash
DevOps
Red Hat
VMWare
Debian
Azure
AWS
Docker
Kubernetes
Ubuntu
KVM
CentOS Stream
IAM
Linux
TCP/IP
Cybersecurity
CIS Benchmarks
Apply
≈ $69k – $155k per year (Estimated) • Hybrid • Bachelor's Degree • Morrisville
Python
JavaScript
TypeScript
SQL
AI/ML
LangGraph
LangChain
LlamaIndex
Model Context Protocol
Prompt Engineering
AI Agents
NLP
LLM
RAG
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Management
ServiceNow
Apply
$290k – $520k per year • In office • 10+ years exp • New York
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
Quantization
Knowledge Distillation
SGLang
TensorRT
TensorRT-LLM
TRL
Transformers
LLM
Mixture of Experts
CUDA
Triton
DPO
GRPO
Post-training
Megatron-LM
ROCm
Speculative Decoding
KV Cache
Tool Use
Model Distillation
Reward Modeling
DevOps
SLI/SLO/SLA
Apply
$220k – $330k per year • In office • 8+ years exp • New York
Python
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
Fine-tuning
SGLang
TensorRT
TensorRT-LLM
TRL
Transformers
PyTorch
LLM
Mixture of Experts
CUDA
Triton
OpenAI
PPO
GRPO
Post-training
NVLink
ROCm
CUTLASS
Speculative Decoding
KV Cache
Tool Use
Reward Modeling
DevOps
Kubernetes
Apply
$220k – $267k per year • In office • New York
Go
TypeScript
DevOps
Terraform
IAM
Apply
$220k – $293k per year • In office • 8+ years exp • New York
AI/ML
Fine-tuning
LLM Guardrails
DevOps
Terraform
Platform Engineering
SLI/SLO/SLA
API Gateway
Apply
≈ $221k – $402k per year (Estimated) • In office • New York
Python
AI/ML
InfiniBand
DevOps
SLURM
Kubernetes
HPC
Apply
≈ $83k – $164k per year (Estimated) • In office • Bachelor's Degree • Houston
Python
JavaScript
PHP
SQL
C#
C++
C#
.NET
Databases
MySQL
PostgreSQL
MS SQL
MariaDB
AI/ML
Machine Learning
DevOps
Git
AWS
Docker
Linux
Windows
Unix
TCP/IP
Management
Agile
Scrum
Apply
In office • Full-Time • Houston
Apply
≈ $42k – $69k per year (Estimated) • In office • Internship • Houston
Apply
≈ $56k – $118k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Houston
Apply
≈ $70k – $147k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Houston
Analytics
Microsoft Excel
Management
Microsoft Office
Apply
See all jobs
This is one of many
1,264,521 more open roles from verified company boards, updated every day.