823,562open jobs
53,068companies
134,459added this week
Browse all
Salary
$190k – $260k per year
Location
In office
Seniority
Staff · 6+ years exp

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Jun 9, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency while contributing to the technology that powers the future.

About the Role

As a Staff Observability Platform Engineer, you'll play a critical role in building and evolving Nscale's observability platform, enabling deep visibility into GPU clusters, AI workloads, and the infrastructure that powers them.

You view observability as a product, not simply a collection of tools. You'll help define and implement scalable, reliable observability solutions that empower engineering teams to understand system behavior, diagnose issues quickly, and operate complex distributed systems with confidence.

You'll combine technical leadership with hands-on engineering, partnering across SRE, infrastructure, platform, and AI/ML teams to improve reliability, operational efficiency, and developer experience. You'll influence architectural decisions, establish engineering best practices, and help drive the evolution of observability capabilities across the organization.

This is a role for someone who enjoys solving difficult infrastructure problems, building platforms that scale, and helping engineering teams succeed through better visibility and operational insight.

What You'll Do

  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.

  • Lead the implementation of scalable observability solutions that support Nscale's growing GPU and AI infrastructure.

  • Partner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.

  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.

  • Develop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.

  • Identify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.

  • Contribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.

  • Lead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.

  • Mentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.

  • Participate in incident investigations and postmortems, translating operational learnings into durable platform improvements.

  • Evaluate new observability technologies and practices, balancing innovation with operational simplicity and long-term maintainability.

About You

  • 6+ years of experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.

  • Strong experience building and operating observability platforms in cloud-native, distributed environments.

  • Deep hands-on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.

  • Strong software engineering skills with proficiency in Go, Python, or equivalent languages.

  • Experience operating and troubleshooting Kubernetes-based platforms at scale.

  • Strong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.

  • Experience designing systems with scalability, reliability, performance, and operational simplicity in mind.

  • Proficiency with Infrastructure-as-Code tools such as Terraform, Ansible, or equivalent.

  • Ability to lead technical initiatives and influence engineering decisions across multiple teams.

  • Excellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.

Preferred

  • Experience operating observability systems in GPU, AI/ML, HPC, or large-scale compute environments.

  • Familiarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.

  • Experience with high-volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.

  • Knowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.

  • Experience mentoring engineers and helping grow technical capability across teams.

Equal Opportunities Statement

We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds.

If there's anything we can do to accommodate your specific situation, please let us know.

Note:  Responsibilities outlined are not exhaustive and may evolve as business needs change.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$190,000—$260,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
823,562 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
$200k – $300k per year • In office • Full-Time • New York
Management
Stripe
Apply
≈ $115k – $212k per year (Estimated) • Remote (United States) • Full-Time • 6+ years exp
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
AI/ML
Prompt Engineering
AI Agents
Frontend
React.js
Apply
$125k – $156k per year • Remote (United States) • Full-Time • United States
DevOps
Terraform
Azure DevOps
Azure
CI/CD
GitOps
Kubernetes
Platform Engineering
Azure AKS
Apply
$180k – $243k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Austin
Apply
$60k – $160k per year • In office • Phoenix
AI/ML
Claude
ChatGPT
Claude Code
AI Agents
OpenAI Codex
DevOps
Terraform
GCP
Azure DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
GitHub
GitLab
Management
Slack
Google Drive
Apply
≈ $40k – $110k per year (Estimated) • In office • Internship • Paris
Python
SQL
AI/ML
Pandas
Analytics
Matplotlib
Management
Agile
Microsoft Office
Apply
Software Developer 2 days ago
≈ $40k – $105k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Porto
Python
Java
SQL
Java
Spring Boot
AI/ML
NLP
TensorFlow
PyTorch
LLM
DevOps
CI/CD
Docker
Kubernetes
Management
Agile
Apply
IAR Data Analyst 1 day ago
$75k – $113k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Salt Lake City
Python
SQL
Databases
Databricks
DevOps
GitHub
Analytics
Tableau
Power BI
ETL/ELT
Apply
Data Engineer 5 2 days ago
$209k – $239k per year • In office • Full-Time • 9+ years exp • Bachelor's Degree • Richmond
Python
Java
SQL
Scala
Databases
Snowflake
Databricks
Cassandra
DynamoDB
Amazon Redshift
AI/ML
Spark
Dagster
Machine Learning
DevOps
Splunk
GCP
Azure
AWS
Management
Agile
Apply
$88k – $100k per year • In office • Full-Time • 3+ years exp • High School Diploma • Riverwoods
Python
PowerShell
DevOps
GCP
VMWare
Azure
AWS
Incident Management
Windows
TCP/IP
DNS
DHCP
Management
ITIL
Apply
See all jobs
This is one of many
823,562 more open roles from verified company boards, updated every day.