436,067open jobs
15,276companies
63,340added this week
Browse all
Location
In office (London)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Northern Data Group operates high-performance computing and data centre capacity that grew out of large-scale crypto mining infrastructure. It now runs GPU cloud services for artificial intelligence alongside its remaining digital asset operations. The listed German company builds and manages sites across Europe and North America.

Job Description

The Operations Engineering Manager will own the operational reliability of the company’s GPU-accelerated HPC infrastructure and lead ateam of Operations Engineers. This role combines technical leadership with people management and Agile delivery ownership.

Thesuccessful candidatewill set the vision for operational excellence, manage and develop the internal team, and work closely with Platform, Network, and Infrastructure teamsto build operational excellence. The role will also play a key part in implementing and maturing the company’s scaled Agile Framework across OperationsEngineering, ensuring alignment, transparency, and continuous improvement.

YOUR RESPONSIBILITIES

Team Leadership & Management

  • Lead, coach, and develop an internal team of Operations Engineers.

  • Set clear goals, priorities, and expectations for the team.

  • Manage workload and resource allocation across the team.

  • Support regular 1:1s, performance reviews, and development plans.

  • Build a collaborative team culture focused on ownership and accountability.

Operational Ownership & Reliability

  • Own the reliability, performance, and availability of the GPU-accelerated HPC infrastructure from an Operations perspective.

  • Oversee proactive system monitoring, incident trend analysis, and root cause analysis.

  • Define, track, and report on key operational metrics.

  • Ensure strong operational control through effective processes, runbooks, and change management.

  • Drive proactive improvements in reliability, automation, and performance.

Agile Ways of Working & Framework Oversight

  • Champion and oversee the scaled Agile Framework within the Operations function.

  • Collaborate with Product, Platform, and Network teams to align priorities and manage backlogs.

  • Support Agile ceremonies such as planning, stand-ups, reviews, and retrospectives.

  • Improve delivery flow, predictability, and cross-team coordination.

  • Ensure work is prioritised and delivered in line with Agile principles.

Process, Automation, and Documentation

  • Own the creation and maintenance of operational documentation, SOPs, and troubleshooting guides.

  • Drive automation initiatives to reduce manual effort and improve consistency.

  • Promote best practices in scripting, configuration management, and observability.

  • Maintain effective knowledge sharing across the team.

  • Stay informed on relevant trends and assess opportunities for improvement.

People & Stakeholder Communication

  • Provide regular status updates and reporting to leadership.

  • Represent Operations in cross-functional planning and strategic discussions.

  • Act as an escalation point for major incidents and complex technical issues.

  • Coordinate with Platform, Network, and third-party support teams during critical events.

  • Communicate clearly to support alignment, risk management, and decision-making.

YOUR QUALIFICATIONS

Required

  • 5+ years in infrastructure/operations, with 2+ years managing a technical team.

  • Advanced Linux administration in production, ideally at scale.

  • Proven experience running incident/problem management and working with third-party or external support teams.

  • Hands-on with automation (Ansible or equivalent) and monitoring/observability tools (e.g. Grafana, Prometheus).

  • Experience with Agile ways of working and exposure to scaled Agile frameworks.

  • Excellent communication and stakeholder management skills, able to work closely with Platform, Network, and leadership.

Nice to Have

  • Experience in HPC or GPU-accelerated environments (NVIDIA GPUs, InfiniBand/RDMA, parallel file systems).

  • Scripting skills in Python and/or Bash for automation and tooling.

  • Understanding of performance tuning for HPC/GPU systems.

  • Experience with CI/CD pipelines and modern DevOps tooling.

  • Background designing or improving on-call rotations, runbooks, and incident readiness.

WHAT WE OFFER

With us, you will work towards the future of HPC: From new, sustainable building methods for data centers to cooling concepts to software solutions for accelerated compute.

Your approaches count: In official exchange formats or spontaneously at the coffee machine. At Northern Data, it's the best idea that counts - not the hierarchy. We’re looking forward to getting your inputs!

You make the difference in the company: Unlike in established corporations, at Northern Data you will really help shape things. From implementing new departments, to optimizing processes and culture.

Best-in-class partners: The best work with Northern Data. This means a knowledge and time advantage from which your career and our customers benefit equally.

Green by heart: Sustainability is at the core of Northern Data. With us, you actively work on the carbon neutrality of datacenters worldwide. Beginning with our infrastructure and continuing with the solutions for our clients, we work towards a green future.

Home Office facts: Work with our international and virtual team flexible from home. And of course, your hardware wishes will be fulfilled to make your ideas for next level HPC come true.

Your wellness matters: At Northern Data we have regularwellbeinginitiatives that are designed to promote wellness, diversity, inclusion, and much more, ensuring a supportive and enriching environment for our global team.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
436,067 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
DevOps/SRE 1 day ago
$39k – $112k per year (Estimated) • In office • Full-Time
Python
Java
SQL
DevOps
Terraform
Ansible
GCP
Prometheus
CI/CD
AWS
Docker
Kubernetes
Grafana
Amazon S3
IAM
Cybersecurity
Keycloak
Least Privilege
Analytics
Apache NiFi
Apply
$89k – $222k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Python
Go
AI/ML
Fine-tuning
AI Agents
NVLink
DevOps
gRPC
Terraform
Helm
GitHub Actions
OpenTelemetry
Kustomize
Prometheus
GitLab CI
SLURM
CI/CD
GitOps
ArgoCD
Kubernetes
Grafana
SRE
Platform Engineering
GitHub
GitLab
HPC
Cybersecurity
Kyverno
Apply
In office • Full-Time • Moscow
Python
SQL
AI/ML
Weights & Biases
MLFlow
Scikit-learn
Computer Vision
NLP
BentoML
TensorFlow
Pandas
NumPy
PyTorch
LLM
RAG
Whisper
KServe
OpenAI
DevOps
GCP
GitHub Actions
Prometheus
Yandex Cloud
GitLab CI
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
GitHub
GitLab
Analytics
ETL/ELT
Apply
$44k – $110k per year (Estimated) • In office • Full-Time • 4+ years exp • Seoul
Databases
PostgreSQL
Redis
ElasticSearch
Apache Kafka
DevOps
Terraform
Puppet
Ansible
GCP
Apply
$12k – $27k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Cyberjaya
DevOps
HPC
Web3
Bitcoin
Apply
$79k – $175k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Frankfurt am Main
Python
PowerShell
DevOps
Platform Engineering
IAM
HPC
Cybersecurity
SOC 2
GDPR
Management
Confluence
Jira
SharePoint
Apply
$51k – $136k per year (Estimated) • In office • Full-Time • London
DevOps
HPC
Apply
$105k – $175k per year (Estimated) • In office • Full-Time • London
Apply
$66k – $177k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
Marketing
Salesforce
Apply
Editor / Speechwriter 12 hours ago
$64k – $169k per year (Estimated) • In office • Full-Time • London
Apply
$73k – $122k per year (Estimated) • In office • Full-Time • 5+ years exp • London
Analytics
Power BI
Management
Outlook
Apply
$81k – $152k per year (Estimated) • Remote/Hybrid • Full-Time • London
Analytics
Power BI
Microsoft Excel
Apply
See all jobs
This is one of many
436,067 more open roles from verified company boards, updated every day.