1,434,312open jobs
83,785companies
217,826added this week
Browse all
Salary
$190k – $300k per year
Location
In office
Seniority
Principal · 8+ years exp
Visa
H-1B filings in 12 months: 10 · for this role: 7

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 9, 2026. Nscale scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

.

Principal Observability Platform Engineer - Nscale

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency while contributing to the technology that powers the future.

About the Role

As a Principal/Staff Observability Platform Engineer, you'll own the technical direction of Nscale's observability platform: the systems that give us deep visibility into GPU clusters, AI workloads, and the infrastructure running them. You treat observability as a product and a discipline, not a tooling exercise. You'll set the architectural roadmap, raise the engineering bar across teams, and ensure our platform scales ahead of the business, not behind it.

You understand that complexity is a cost. Solutions that require constant babysitting don't scale, and neither does operational burden. The platforms you build should be simple to operate, easy to understand, and self-evidently correct when something goes wrong.

This isn't a "maintain and operate" role. It's a "define, build, and lead" role.

What You'll Do

  • Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at scale.
  • Drive platform decisions that have multi-year impact: tooling, data models, ingestion patterns, retention, cardinality management.
  • Identify systemic gaps before they become incidents; design platforms that make failure visible and fast to diagnose.
  • Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how Nscale builds and operates.
  • Define standards and patterns that other engineers adopt, not by mandate, but because they're clearly better.
  • Mentor and technically grow the observability team; raise the ceiling on what the team can build and own.
  • Lead incident postmortems and use them to drive durable platform improvements.
  • Evaluate and introduce tooling that meaningfully improves signal quality, operational efficiency, or scalability, and retire what doesn't.

About You

  • 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
  • You've operated observability infrastructure at serious scale. You know what breaks at 10x and you design for it.
  • You have a strong bias toward simplicity. You've seen over-engineered observability stacks collapse under their own weight and you build accordingly.
  • Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic.
  • Strong engineering fundamentals, proficient in Python, Go, or similar; comfortable owning complex systems end to end.
  • Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus.
  • You can architect systems, write the code, review others' work, and explain the tradeoffs clearly, all in the same week.
  • Infrastructure-as-Code is default, not optional (Terraform, Ansible, or equivalent).
  • You influence without authority. Teams want your opinion because it makes their work better.

Preferred

  • Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.).
  • Background in AI/ML infrastructure observability: GPU utilisation, training job visibility, inference latency.
  • Prior experience defining observability strategy at an organisation level.

Equal Opportunities Statement

We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

Note:  Responsibilities outlined are not exhaustive and may evolve as business needs change.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$190,000—$300,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,434,312 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
$135k – $206k per year • Hybrid • 12+ years exp • Bachelor's Degree • Chicago
Apply
$118k – $180k per year • Hybrid • 10+ years exp • Bachelor's Degree • Albuquerque
Apply
$135k – $206k per year • Hybrid • 12+ years exp • Bachelor's Degree • Albuquerque
Apply
$118k – $180k per year • Hybrid • 10+ years exp • Bachelor's Degree • Oak Ridge
Apply
$135k – $206k per year • Hybrid • 12+ years exp • Bachelor's Degree • Oak Ridge
Apply
DevOps Engineer 4 hours ago
$122k – $211k per year • Remote (United States) • Secret • 5+ years exp • Bachelor's Degree • United States
Python
PowerShell
Databases
OpenSearch
DevOps
Splunk
Terraform
Helm
Azure DevOps
GitHub Actions
Rancher
CloudFormation
Prometheus
GitLab CI
Azure
CI/CD
GitOps
Jenkins
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Amazon CloudWatch
DNS
Management
Agile
Apply
≈ $40k – $93k per year (Estimated) • Remote (Spain) • Full-Time • Spain • Estonia
Python
SQL
Databases
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
dbt
AI Agents
LLM
LLM Guardrails
DevOps
Terraform
GCP
GitHub Actions
Azure
CI/CD
AWS
Platform Engineering
IAM
Apply
≈ $19k – $36k per year (Estimated) • In office • Internship • Bachelor's Degree • Rio de Janeiro
Python
C++
Apply
≈ $129k – $245k per year (Estimated) • In office • TS/SCI • 7+ years exp • Fort Belvoir
Python
AI/ML
Machine Learning
DevOps
Splunk
VMWare
Azure
AWS
AIOps
Management
ServiceNow
ITSM
Apply
Full Stack Engineer 4 4 hours ago
$179k – $205k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Richmond
Python
JavaScript
Rust
TypeScript
C#
Node JS
Scala
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
Management
Agile
Apply
≈ $148k – $290k per year (Estimated) • In office • New York
AI/ML
Claude
ChatGPT
Gemini
DevOps
Platform Engineering
Cybersecurity
Okta
ISO 27001
SOC 2
Microsoft Entra ID
Management
Slack
Confluence
Notion
Jira
Google Workspace
Apply
≈ $126k – $284k per year (Estimated) • In office • Seattle
Python
Bash
DevOps
Terraform
Ansible
Kubernetes
Proxmox VE
KVM
OpenStack
Linux
Unix
Apply
≈ $165k – $340k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • New York
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
Fine-tuning
PyTorch
Pre-training
Megatron-LM
NCCL
InfiniBand
NVLink
ROCm
Edge AI
DevOps
Terraform
Ansible
OpenTelemetry
Prometheus
SLURM
Docker
Kubernetes
Grafana
Configuration Management
HPC
Linux
TCP/IP
BGP
Apply
$160k – $230k per year • In office • 5+ years exp
Python
Go
Databases
ClickHouse
Apache Kafka
DevOps
Terraform
Ansible
Loki
OpenTelemetry
Fluent Bit
Prometheus
VictoriaMetrics
Kubernetes
Grafana
Platform Engineering
Thanos
Vector
Apply
≈ $158k – $327k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • New York
AI/ML
InfiniBand
Edge AI
DevOps
CI/CD
Incident Management
HPC
Linux
Management
Agile
Scrum
Apply
See all jobs
This is one of many
1,434,312 more open roles from verified company boards, updated every day.