368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$88k – $183k per year (Estimated)
Location
Remote/Hybrid (London, United Kingdom)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
AbcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ ABOUT PHOTO APP What you write here is totally up to you, there is no right or wrong way to complete your 'About Us' page. We advise making your language friendly and approachable, while being informative and professional.

About Us

We’re a fast-growing GPU-as-a-Service provider, delivering scalable, high-performance compute infrastructure purpose-built for AI and HPC workloads. Operating across global data centres, we run mission-critical environments where uptime, throughput, and ultra-low latency are non-negotiable.

Role Overview

We are seeking an Infrastructure Tooling & Observability Engineer to act as a key engineering force within our global Infrastructure Operations organisation. Working closely with our SRE teams, you will translate high-level reliability objectives into scalable, production-ready systems that directly improve the resilience, efficiency, and performance of our global infrastructure.

This role goes beyond traditional monitoring. You will help design and build the internal control plane that enables operations at scale across a rapidly growing GPU fleet. Your work will focus on transforming complex, high-volume telemetry-spanning logs, metrics, and events across HPC, networking, and platform layers-into actionable insight that drives operational excellence and proactive reliability.

A core part of your responsibility will be developing intelligent observability and automation systems, including advanced alerting strategies, anomaly detection, and AI-driven tooling that reduces L1/L2 escalations and removes operational toil. You will also contribute to Continual Service Improvement (CSI) initiatives by building frameworks for reliability measurement, automated remediation, and system health evaluation.

In addition, you will play a central role in turning SRE reliability initiatives into scalable engineering solutions. This includes designing and delivering capabilities such as inventory management systems, performance testing frameworks, and automated performance result collection. You will also help eliminate manual workflows involved in onboarding new regions, facilities, and clusters, embedding automation and standardisation into every stage of infrastructure deployment.

As the organisation scales, you will act as a critical interface between operations and engineering teams. You will evaluate and mature internally built tooling-from capacity planning systems to autonomous remediation pipelines-and help integrate these capabilities into core infrastructure platforms to ensure consistent, high-performance, and highly reliable global operations.

What’s In it for you?

Join a team building the internal platforms that enable large-scale infrastructure to operate reliably, efficiently, and at speed. As an Infrastructure Tooling & Observability Engineer, you will design and develop the systems that power visibility, automation, and operational intelligence across complex distributed environments.

This role goes beyond traditional monitoring. You will build the internal control plane that transforms high-volume telemetry-logs, metrics, and events-into actionable insight for engineering and operations teams. Your work will improve observability across infrastructure systems, strengthen signal quality, and help teams understand and respond to system behaviour in real time.

Working closely with SRE and infrastructure engineering teams, you will translate reliability goals into scalable, production-grade tooling. This includes frameworks for observability, alerting, anomaly detection, capacity planning, and service health tracking.

A key focus of the role is automation. You will help eliminate manual processes across infrastructure operations, including environment provisioning, cluster onboarding, inventory management, and recurring operational workflows. You will also contribute to performance engineering initiatives, building tooling for testing, benchmarking, and automated results collection at scale.

You will play a central role in turning SRE reliability initiatives into reusable engineering solutions, including automated remediation systems and tooling that reduces operational toil while improving system resilience.

You can also expect:

Exposure to large-scale distributed infrastructure systems

Opportunities to shape foundational internal platforms

A collaborative, engineering-led culture with strong ownership

High-impact work spanning observability, automation, and reliability

Close partnership with SRE and infrastructure engineering teams

A fast-moving environment where tooling directly improves operational performance

Key Responsibilities

  • Design, build, and evolve internal tooling and observability platforms that support large-scale infrastructure operations across distributed environments.

  • Develop systems that turn high-volume telemetry (logs, metrics, events) into actionable insight, improving visibility, alerting quality, and operational decision-making.

  • Translate SRE reliability requirements into scalable, production-ready software solutions, including automation for incident detection, prevention, and remediation.

  • Drive automation across infrastructure operations, reducing manual effort in areas such as environment provisioning, cluster onboarding, inventory management, and lifecycle workflows.

  • Build tooling for capacity management, performance testing, benchmarking, and automated collection and analysis of results.

  • Contribute to Continual Service Improvement (CSI) initiatives by identifying operational inefficiencies and delivering durable engineering solutions.

  • Work closely with SRE and infrastructure engineering teams to embed observability and reliability into core platform workflows.

  • Interface with Platform Engineering teams to ensure tooling aligns with broader orchestration and infrastructure strategy.

  • Integrate and extend existing systems written in Ruby/Rails and Go, contributing to a consistent and maintainable engineering ecosystem.

  • Develop and maintain automation workflows using Ansible and AWX.

  • Support CI/CD-driven operational tooling, including GitHub Actions and self-hosted runners.

Essential Skills & Experience

  • Degree in Computer Science/Software Engineering, or equivalent experience

  • 6-8 years of experience in infrastructure engineering, DevOps, SRE, and/or software engineering roles, with a strong focus on operational systems.

  • Proven experience in at least one recent DevOps or software engineering role, building or maintaining production infrastructure tooling or platform systems.

  • Experience working in large-scale or distributed infrastructure environments (hyperscale, enterprise, or similarly complex systems).

  • Strong programming ability in at least one of: Ruby (Rails), Go, or similar systems languages, with willingness and ability to work across multiple languages and codebases.

  • Hands-on experience with infrastructure automation tools such as Ansible and orchestration platforms such as AWX.

  • Strong experience with observability systems, including the Grafana stack (Prometheus, Loki, Mimir, and Grafana Alloy).

  • Familiarity with low-level telemetry and infrastructure protocols such as SNMP and syslog.

  • Experience working with Kubernetes or similar orchestration platforms in production environments.

  • Understanding of API design and integration patterns, particularly REST-based services and service-to-service communication.

  • Experience building and maintaining CI/CD pipelines, including GitHub Actions and self-hosted runners.

  • Strong understanding of operational reliability concepts, including monitoring, alerting, capacity planning, and incident response.

  • Comfortable working closely with SRE, Platform Engineering, and infrastructure teams to translate operational needs into maintainable software systems.

Preferred Qualifications

  • Kubernetes Certified Administrator

  • Cloud-native observability training courses, attendance at industry conferences in this field

  • CompTIA+ Security Qualifications

  • LPI/LPIC certification

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$35k per year • In office • Internship • Bachelor's Degree • Freiburg im Breisgau
C++
Java
Node JS
Python
C#
JavaScript
C#
.NET
AI/ML
AI Agents
DevOps
CI/CD
GitLab
Apply
$19k – $49k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Hyderabad
SQL
C#
TypeScript
JavaScript
C#
.NET
Databases
MS SQL
Frontend
Angular
DevOps
CI/CD
Git
GitHub
Jenkins
Apply
$72k – $119k per year • In office • Full-Time • Bachelor's Degree • Boca Raton • Alpharetta • Dayton
Java
DevOps
CI/CD
Git
Apply
$184k – $307k per year • Remote • Full-Time • PhD • United States
DevOps
CI/CD
Apply
$137k – $229k per year • Remote • Full-Time • PhD • United States
AI/ML
Human-in-the-Loop
DevOps
CI/CD
Apply
$80k – $191k per year (Estimated) • In office • Full-Time • London
DevOps
Incident Management
Apply
Engineering Manager 25 days ago
$126k – $227k per year (Estimated) • Remote/Hybrid • Full-Time • London
DevOps
CI/CD
Kubernetes
Apply
Cluster Architect 27 days ago
$91k – $215k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • London
C++
AI/ML
CUDA
CUDA Toolkit
InfiniBand
DevOps
Docker
Docker Swarm
Kubernetes
SLURM
HPC
Apply
DevOps Engineer 1 month ago
$69k – $171k per year (Estimated) • Remote/Hybrid • Full-Time • London
Go
Python
DevOps
Ansible
Blue-Green Deployment
CI/CD
Git
GitHub Actions
Helm
Kubernetes
Progressive Delivery
Terraform
GitHub
Apply
Senior SDET 1 month ago
$59k – $146k per year (Estimated) • Remote/Hybrid • Full-Time • London
Go
Databases
PostgreSQL
AI/ML
KServe
LLM
vLLM
DevOps
CI/CD
gRPC
Helm
Kubernetes
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
$87k – $159k per year (Estimated) • In office • Contractor • 5+ years exp • London • Stockholm
Design
Figma
Marketing
Zendesk
Apply
$105k – $204k per year (Estimated) • In office • Full-Time • London
Python
SQL
Databases
Snowflake
AI/ML
Dagster
dbt
Analytics
A/B Testing
Apply
$27k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • London
Analytics
Power BI
Tableau
Marketing
Salesforce
Apply
$34k – $85k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
AI/ML
AI Agents
Edge AI
QA
Appium
Cucumber
Cypress
Selenium
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.