683,840open jobs
39,578companies
97,862added this week
Browse all
Salary
$145k – $180k per year
Location
Remote/Hybrid (United States)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

.

Operational Data & Observability Engineer

About the Role

We're looking for an Operational Data & Observability Engineer to build and evolve the monitoring, logging, and observability capabilities that power our production environments. In this role, you'll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights.

You'll partner closely with DevOps, Site Reliability Engineering (SRE), platform, and software engineering teams to develop modern observability solutions, improve incident response, and enable data-driven operational excellence.

What You'll Do

Design & Build Observability Solutions

  • Design and implement enterprise observability strategies across infrastructure, services, and applications.

  • Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility.

  • Build and maintain centralized logging and log analysis pipelines.

  • Implement distributed tracing to improve visibility across microservices and complex application workflows.

  • Establish performance baselines and develop anomaly detection strategies.

Operational Data Engineering

  • Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems.

  • Design and manage operational data pipelines that support monitoring and analytics.

  • Develop APIs and integrations that enable operational data consumption across teams.

  • Ensure data quality, consistency, retention, and cost-efficient storage practices.

Reliability & Operations

  • Troubleshoot production issues using monitoring, logging, and tracing data.

  • Participate in an on-call rotation and support incident response activities.

  • Create and maintain operational documentation, runbooks, and troubleshooting guides.

  • Partner with engineering teams to improve platform reliability, scalability, and operational readiness.

  • Continuously optimize observability infrastructure for performance and resilience.

Platform & Tool Administration

  • Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar technologies.

  • Evaluate emerging observability tools and recommend improvements.

  • Automate monitoring deployments, instrumentation, and platform configuration.

  • Perform ongoing maintenance, upgrades, and lifecycle management of observability infrastructure.

What You'll Bring

Required Qualifications

  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering.

  • Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent.

  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions.

  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent.

  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM).

  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies.

  • Solid understanding of application, infrastructure, networking, database, and storage performance monitoring.

  • Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem-solving.

Preferred Qualifications

  • Experience supporting microservices-based architectures.

  • Expertise across multiple observability platforms.

  • Experience with incident management, root cause analysis, and post-incident reviews.

  • Infrastructure as Code experience using Terraform, Ansible, or similar tools.

  • Familiarity with eBPF or low-level Linux performance monitoring.

  • Experience building custom telemetry, ETL, or operational data pipelines.

  • Understanding of security monitoring, audit logging, and compliance requirements.

What Success Looks Like

Success in this role will be measured by your ability to:

  • Improve platform visibility and operational health.

  • Reduce Mean Time to Resolution (MTTR) during incidents.

  • Increase alert quality while reducing unnecessary noise.

  • Deliver highly available, scalable observability platforms.

  • Improve engineering productivity through actionable monitoring and operational insights.

  • Optimize observability infrastructure performance and cost efficiency.

Work Environment

  • Participate in a rotating on-call schedule to support production environments.

  • Support mission-critical systems with occasional after-hours or incident response responsibilities.

  • Hybrid or remote work arrangements available, depending on business needs.

Why Join Us?

You'll play a critical role in building the operational intelligence that keeps our platforms running at scale. If you're passionate about observability, automation, reliability, and empowering engineering teams with meaningful operational insights, we'd love to hear from you.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$145,000—$180,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
683,840 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$80k per year • Remote/Hybrid • Full-Time • London
Python
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
Frontend
GraphQL
React.js
WebAssembly
DevOps
Terraform
GCP
Azure
CI/CD
Git
AWS
Kubernetes
Incident Management
GitHub
Management
Agile
Apply
$124k – $250k per year (Estimated) • Remote • Full-Time • New York
Python
JavaScript
SQL
Node JS
Frontend
D3.js
Chart.js
DevOps
GCP
AWS
Analytics
Plotly
Management
n8n
Zapier
Apply
$51k – $84k per year (Estimated) • Remote
Python
SQL
Bash
Databases
Apache Kafka
DevOps
gRPC
Kibana
CI/CD
Docker
Kubernetes
QA
Pytest
Apply
$31k – $81k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Warsaw
Python
AI/ML
Fine-tuning
Prompt Engineering
Speech Recognition
RAG
Apply
$126k – $173k per year • Remote • Full-Time • 11+ years exp • High School Diploma • United States
Python
Java
Rust
SQL
Databases
Snowflake
Databricks
AI/ML
Copilot
Claude Code
AI Agents
LLM
OpenAI
OpenAI Codex
DevOps
Terraform
Puppet
Ansible
Chef
Pulumi
Azure
AWS
Apply
$80k – $186k per year (Estimated) • In office • 6+ years exp • London
DevOps
SLI/SLO/SLA
Apply
$160k – $290k per year • In office • 10+ years exp • Houston
AI/ML
InfiniBand
Apply
$30k – $76k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree
Analytics
Microsoft Excel
Apply
$220k – $260k per year • In office • 5+ years exp • Bachelor's Degree • New York
AI/ML
InfiniBand
DevOps
Datadog
Prometheus
Grafana
OpenStack
Incident Management
Management
Jira
ServiceNow
Apply
$170k – $300k per year • In office • 7+ years exp • Houston
Apply
See all jobs
This is one of many
683,840 more open roles from verified company boards, updated every day.