368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$89k – $184k per year (Estimated)
Location
Remote/Hybrid (London, United Kingdom)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
AbcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ ABOUT PHOTO APP What you write here is totally up to you, there is no right or wrong way to complete your 'About Us' page. We advise making your language friendly and approachable, while being informative and professional.

About Us

We’re a fast-growing GPU-as-a-Service provider, delivering scalable, high-performance compute infrastructure purpose-built for AI and HPC workloads. Operating across global data centres, we run mission-critical environments where uptime, throughput, and ultra-low latency are non-negotiable.

Role Overview

We are looking for a senior Infrastructure Site Reliability Engineer with deep experience operating large-scale distributed systems and recent hands-on expertise in high-performance computing (HPC) and AI infrastructure. This is an operations-first SRE role, working in a 24/7/365 on-call environment, responsible for ensuring reliability, performance, and continuous improvement of mission-critical infrastructure. This role sits within a cross-functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations.

The ideal candidate has progressed through large-scale, globally distributed or multi-site infrastructure environments and has more recently specialised in GPU-accelerated HPC systems. This role provides exposure to the latest high-density AI compute platforms, including next-generation GPU infrastructure at significant scale. You will bring strong breadth across bare metal, networking, storage, virtualisation, and orchestration, alongside deep HPC experience including NVIDIA GPU ecosystems, RDMA networking (RoCE and InfiniBand), and performance validation and benchmarking. Strong Linux and distributed systems expertise is essential.

Alongside operational ownership, this is a deeply technical Infrastructure SRE role centred on advanced operational troubleshooting and performance evaluation across large-scale HPC systems. You will investigate complex, cross-layer issues spanning GPU compute, networking, storage, and orchestration, building a clear understanding of system behaviour under real production AI and HPC workloads.

A key responsibility is performance evaluation, testing, and operational acceptance of new HPC environments, ensuring platforms meet defined reliability, scalability, and performance expectations before entering production. You will work across hardware, network, and software layers to validate readiness of high-density GPU infrastructure and support safe, predictable deployment at scale.

You will also play a central role in continuous service improvement (CSI)-reducing operational toil, increasing automation, and improving reliability, consistency, and operational efficiency across the platform. This includes strengthening observability, refining operational workflows, and eliminating repetitive or failure-prone processes.

Over time, you will help shape future infrastructure design and deployment approaches, feeding operational insight back into infrastructure engineering decisions and ensuring production learnings directly influence next-generation HPC platform evolution.

What’s in It for You

Join a team operating some of the world’s most advanced high-performance computing infrastructure. As a HPC Infrastructure SRE, you’ll work hands-on with cutting-edge GPU and CPU platforms - including the latest NVIDIA architectures - powering dense, large-scale compute environments used for AI, machine learning, and next-generation workloads.

This is an opportunity to build expertise at the forefront of modern infrastructure, where reliability, scale, and performance matter every day. You’ll collaborate with experienced engineers across a globally distributed organisation that values openness, inclusion, technical excellence, and continuous learning.

We move quickly, solve meaningful challenges, and give people the space to make an impact. If you thrive in fast-paced environments, enjoy working with advanced technology, and want to help shape the future of high-performance compute, you’ll find both challenge and opportunity here.

You can also expect:

  • Exposure to industry-leading GPU and AI infrastructure

  • Opportunities to grow alongside a rapidly scaling global business

  • A collaborative, inclusive, and supportive engineering culture

  • Real ownership and the ability to influence operational excellence

  • Work that sits at the intersection of people, performance, and technology

  • A modern, flexible, globally connected workplace with ambitious goals

  • Key Responsibilities

  • Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment

  • Participate in a 24x7x365 on-call rotation, supporting mission-critical systems and incident response

  • Troubleshoot complex issues across compute, networking, storage, and orchestration layers in GPU-accelerated environments

  • Lead performance evaluation, testing, and operational acceptance of new HPC infrastructure before production release

  • Drive continuous service improvement (CSI), reducing toil through automation, tooling, and process refinement

  • Build and maintain infrastructure automation and tooling (IaC and scripting) to improve reliability and operational efficiency

  • Optimise Linux systems for performance, including kernel, BIOS/firmware, and storage tuning for HPC workloads

  • Configure and operate bare-metal infrastructure using IPMI, iLO, iDRAC, Redfish, and related tooling

  • Partner with infrastructure tooling and observability teams to improve telemetry, alerting, and system visibility at scale

  • Own ITIL-aligned processes across Incident, Major Incident, Problem, and Change Management, ensuring strong execution and continuous improvement

  • Lead root cause analysis and ensure corrective actions are implemented and automated where possible

  • Play a key role in designing and delivering future HPC cluster and site builds, shaping global consistency and operational standards

  • Collaborate closely with Platform Engineering, Network Engineering, Infrastructure Tooling, and Data Centre Operations to improve reliability and deployment quality

  • Feed operational insight back into infrastructure design to influence next-generation HPC platform evolution

  • Mentor engineers and act as a technical authority for operational best practices across teams

  • Communicate clearly with technical and non-technical stakeholders, translating complex issues into actionable outcomes

  • Uphold a culture of: do, document, automate

  • Essential Skills & Experience

  • 8+ years experience in Site Reliability Engineering, Infrastructure Engineering, or similar roles in large-scale distributed production environments operating a 24/7 support model

  • 2-3+ years recent experience in HPC and/or AI infrastructure, including GPU-based compute environments at scale

  • Strong Linux expertise (preferably Ubuntu), including deep systems administration and production troubleshooting

  • Proven experience in performance tuning across compute systems, including kernel, BIOS/firmware, and storage subsystem optimisation

  • Strong hands-on experience with bare-metal infrastructure and out-of-band management tooling (IPMI, iLo, iDRAC, Redfish or equivalent)

  • Solid networking fundamentals including TCP/IP, DNS, DHCP, VLANs, routing, and switching, with exposure to high-performance networking environments

  • Exposure to NVIDIA GPU ecosystems, including CUDA-based workloads and GPU-accelerated compute environments, including the NVIDIA AI reference architecture.

  • Familiarity with high-performance networking technologies such as InfiniBand and RoCE

  • Strong experience with infrastructure automation and scripting (e.g. Bash, Python, Ansible or similar IaC/tooling approaches)

  • Understanding of observability principles and practical use of monitoring and telemetry systems (e.g. Prometheus, Grafana or equivalents)

  • Understanding of workload schedulers and running workloads across multiple systems in parallel.

  • Practical experience with at least one parallel storage platform.

  • Experience working in ITIL-aligned environments, including Incident, Major Incident, Problem, and Change Management

  • Strong troubleshooting skills in high-pressure operational environments, with a track record of incident ownership and resolution

  • Strong communication skills with the ability to work across engineering teams and interface with non-technical stakeholders

  • Ability and willingness to collaborate closely with Platform SRE teams, including exposure to and learning of Kubernetes-based orchestration environments (not a core requirement)

  • Experience contributing to or influencing infrastructure design, reliability improvements, or operational best practices

Bonus / Highly Desirable:

  • Deep experience with HPC workloads and GPU-accelerated infrastructure at scale

  • Experience with InfiniBand, RoCE, or other HPC-grade networking fabrics in production environments

  • Experience with HPC benchmarking, validation, or performance testing (e.g. linpac, fio, NCCL, ibdiagnet)

  • Exposure to large-scale multi-site or global infrastructure deployments

Preferred Qualifications

  • Bachelor or Masters Level degree in Computer Science, Engineering or related field, or equivalent ‘on-the-job’ experience.

  • LPIC Certifications

  • ITIL Foundation level qualification or equivalent experience

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
Bash
C++
Java
Python
DevOps
CI/CD
Git
Apply
$217k – $304k per year • Equity • Remote • Full-Time • 8+ years exp
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$80k – $191k per year (Estimated) • In office • Full-Time • London
DevOps
Incident Management
Apply
Engineering Manager 25 days ago
$126k – $227k per year (Estimated) • Remote/Hybrid • Full-Time • London
DevOps
CI/CD
Kubernetes
Apply
Cluster Architect 27 days ago
$91k – $215k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • London
C++
AI/ML
CUDA
CUDA Toolkit
InfiniBand
DevOps
Docker
Docker Swarm
Kubernetes
SLURM
HPC
Apply
DevOps Engineer 1 month ago
$69k – $171k per year (Estimated) • Remote/Hybrid • Full-Time • London
Go
Python
DevOps
Ansible
Blue-Green Deployment
CI/CD
Git
GitHub Actions
Helm
Kubernetes
Progressive Delivery
Terraform
GitHub
Apply
Senior SDET 1 month ago
$59k – $146k per year (Estimated) • Remote/Hybrid • Full-Time • London
Go
Databases
PostgreSQL
AI/ML
KServe
LLM
vLLM
DevOps
CI/CD
gRPC
Helm
Kubernetes
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
$87k – $159k per year (Estimated) • In office • Contractor • 5+ years exp • London • Stockholm
Design
Figma
Marketing
Zendesk
Apply
$105k – $204k per year (Estimated) • In office • Full-Time • London
Python
SQL
Databases
Snowflake
AI/ML
Dagster
dbt
Analytics
A/B Testing
Apply
$27k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • London
Analytics
Power BI
Tableau
Marketing
Salesforce
Apply
$34k – $85k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
AI/ML
AI Agents
Edge AI
QA
Appium
Cucumber
Cypress
Selenium
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.