791,964open jobs
50,470companies
123,351added this week
Browse all
Salary
$160k – $230k per year
Location
In office
Seniority
Senior · 5+ years exp

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Aug 21, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

Senior Observability Platform Engineer - Nscale

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency while contributing to the technology that powers the future.

About the Role

As a Senior Observability Platform Engineer, you’ll play a key role in designing, building, and scaling Nscale’s observability platform. You’ll focus on delivering reliable, high-quality visibility into GPU clusters, AI workloads, and the infrastructure that powers them.

You approach observability as a product-balancing usability, scalability, and operational efficiency. You build systems that reduce cognitive load for engineers, surface meaningful signals, and enable fast, confident debugging when things go wrong.

You’ll contribute to platform direction, implement critical systems, and collaborate closely with SRE, infrastructure, and AI/ML teams to ensure observability is embedded into everything we run.

This is a hands-on engineering role with meaningful influence over platform design and evolution.

What You’ll Do

  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
  • Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies
  • Improve signal quality by reducing noise, managing cardinality, and refining alerting practices
  • Help identify and address observability gaps before they impact reliability
  • Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms
  • Develop reusable patterns, libraries, and best practices that improve consistency across teams
  • Participate in incident response and postmortems, driving actionable improvements
  • Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
  • Support and mentor engineers within the team through code reviews and knowledge sharing

About You

  • 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
  • Experience operating and scaling observability systems in production environments
  • Strong understanding of monitoring concepts: metrics, logs, traces, alerting, and SLOs
  • Hands-on experience with several of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic
  • Solid programming skills (Python, Go, or similar) with the ability to build and maintain production systems
  • Experience working with Kubernetes-based infrastructure
  • Familiarity with Infrastructure-as-Code (Terraform, Ansible, or similar)
  • Pragmatic mindset with a focus on simplicity, reliability, and maintainability
  • Strong collaboration skills and ability to work across teams

Preferred

  • Experience with observability data pipelines (Kafka, Vector, Fluent Bit, etc.)
  • Exposure to AI/ML infrastructure or GPU-based systems
  • Familiarity with performance monitoring for distributed systems
  • Experience improving developer experience through observability tooling

Equal Opportunities Statement

We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

Note: Responsibilities outlined are not exhaustive and may evolve as business needs change

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$160,000—$230,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
791,964 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
$87k – $165k per year • In office • TS/SCI • Full-Time • 5+ years exp • Bachelor's Degree • United States
Java
SQL
C++
DevOps
Splunk
Red Hat
Grafana
Linux
TCP/IP
Cybersecurity
LDAP
PKI
Apply
≈ $133k – $261k per year (Estimated) • Equity • Remote (United States) • 5+ years exp
Databases
Apache Kafka
AI/ML
AI Agents
Ray
DevOps
Terraform
Azure
CI/CD
Docker
Kubernetes
Apply
$180k – $240k per year • Equity • In office • Full-Time • 10+ years exp • PhD • Seattle
Python
TypeScript
AI/ML
AI Agents
OpenAI
Hugging Face
DevOps
Terraform
GCP
Helm
Crossplane
Kustomize
Pulumi
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
Service Mesh
IAM
Linux
DNS
Cybersecurity
SOC 2
HIPAA
FedRAMP
Apply
≈ $81k – $168k per year (Estimated) • In office • 3+ years exp • Houston
Python
Java
C#
Java
Spring Boot
C#
.NET
DevOps
Splunk
Terraform
Datadog
Dynatrace
Prometheus
CI/CD
Jenkins
Docker
Kubernetes
Grafana
SRE
SLI/SLO/SLA
GitLab
Amazon ECS
Apply
≈ $81k – $169k per year (Estimated) • In office • Wilmington
SQL
Databases
Apache Kafka
DevOps
Splunk
Terraform
Datadog
Dynatrace
Prometheus
CI/CD
Jenkins
AWS
Docker
Kubernetes
Grafana
Amazon EKS
SLI/SLO/SLA
GitLab
Amazon CloudWatch
Management
ServiceNow
ITIL
Apply
≈ $71k – $186k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Singapore
Python
PowerShell
Bash
DevOps
Terraform
Ansible
CI/CD
Git
AWS
Docker
Kubernetes
SRE
Platform Engineering
Configuration Management
Linux
Apply
≈ $25k – $61k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Go
JavaScript
Kotlin
TypeScript
Ruby
C#
Scala
AI/ML
Copilot
Cursor
ChatGPT
Claude Code
DevOps
GCP
Azure
CI/CD
AWS
Apply
≈ $96k – $193k per year (Estimated) • In office • Hong Kong
Python
Go
Java
Move
C++
AI/ML
AI Agents
LLM
Recommender Systems
Web3
Smart Contracts
Apply
Remote (likely United States) • Bachelor's Degree
Python
SQL
Bash
Databases
MySQL
DevOps
Ansible
Debian
Proxmox VE
Linux
Cybersecurity
pfSense
Apply
≈ $67k – $179k per year (Estimated) • In office • Full-Time • Seoul
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Quantization
Computer Vision
TensorFlow
PyTorch
Apply
≈ $140k – $286k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Houston
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
Fine-tuning
PyTorch
Pre-training
Megatron-LM
NCCL
InfiniBand
NVLink
ROCm
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
SLURM
Docker
Kubernetes
Grafana
Configuration Management
HPC
Linux
TCP/IP
BGP
Apply
≈ $135k – $274k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Houston
AI/ML
InfiniBand
Edge AI
DevOps
CI/CD
Incident Management
HPC
Linux
Management
Agile
Scrum
Apply
≈ $153k – $299k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Houston
Python
Java
C++
AI/ML
CUDA Toolkit
Prefect
CUDA
NCCL
InfiniBand
NVLink
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
Pulumi
SLURM
CI/CD
Kubernetes
Grafana
OpenStack
HPC
Linux
Apply
≈ $138k – $266k per year (Estimated) • In office • 5+ years exp • Seattle
Python
PowerShell
AI/ML
Gemini
Cybersecurity
Okta
ISO 27001
SOC 2
DLP
Management
Google Workspace
Gmail
Apply
$190k – $260k per year • In office • 6+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
SLURM
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
HPC
Apply
See all jobs
This is one of many
791,964 more open roles from verified company boards, updated every day.