1,431,395open jobs
83,518companies
216,630added this week
Browse all
Salary
$124k – $214k per year
Location
In office (London)
Seniority
Middle · 4+ years exp

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Oct 7, 2026. Microsoft scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Microsoft is an American multinational technology corporation founded in 1975 by Bill Gates and Paul Allen and headquartered in Redmond, Washington. It built the personal computing era around the Windows operating system and the Office productivity suite, and now generates the largest share of its revenue from Azure cloud infrastructure and commercial subscriptions. The company also owns GitHub, LinkedIn and the Xbox gaming business, and has invested heavily in artificial intelligence through its partnership with OpenAI and the Copilot assistants embedded across its products.
Overview

As Microsoft continues to push the boundaries of AI, we are on the lookout for passionate individuals to work with us on the most interesting and challenging AI questions of our time. Our vision is bold and broad - to build systems that have true artificial intelligence across agents, applications, services, and infrastructure. It’s also inclusive: we aim to make AI accessible to all - consumers, businesses, developers - so that everyone can realize its benefits.

We’re looking for an experienced AI Reliability Engineer to join our High Performance Computing (HPC) infrastructure team. In this role, you’ll blend software engineering and systems engineering to keep our large-scale distributed AI infrastructure reliable and efficient. You’ll ensure that AI systems stay efficient and reliable with very high uptimes.

Microsoft AI Our mission is to build AI that amplifies human potential and empowers people around the world. We strive to deliver breakthroughs that advance science, education, productivity, and global well-being. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in-come and join us as we work on our next generation of models! MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.

Responsibilities

  • Reliability & Availability: Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.
  • Observability: Design and maintain monitoring, alerting, and logging systems to provide real-time visibility into all aspects of HPC systems including GPU, clusters, storage and networking.
  • Automation & Tooling: Build automation for deployments, incident response, scaling, and failover in CPU+GPU environments.
  • Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements.
  • Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments.
  • Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows.
  • Embody our Culture and Values.

Qualifications

Required Qualifications:

  • Bachelor’s Degree in Computer Science, or related technical discipline AND 4+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering OR equivalent experience

Preferred Qualifications:

  • Master’s Degree in Computer Science, or related technical discipline AND 2+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering
  • OR equivalent experience
  • Experience with Kubernetes, Docker, container orchestration, and CI/CD pipelines for ML training or inference workloads.
  • Experience with public cloud platforms such as Azure, AWS, or GCP, including infrastructure-as-code.
  • Experience with monitoring and observability tools such as Grafana, Datadog, or OpenTelemetry.
  • Programming or scripting experience in Python, Go, or Bash.
  • Experience with distributed systems, networking, storage, and high-performance computing (HPC).
  • Experience operating GPU clusters and workload schedulers for ML/AI workloads.
  • Experience with ML training or inference pipelines.
  • Experience with capacity planning and cost optimization for GPU-based infrastructure.

Software Engineering IC5 - The typical base pay range for this role across United Kingdom is £ 93,500.00 - £ 161,800.00 per year. Certain roles may be eligible for benefits and other compensation.

Find additional benefits and pay information here:

https://careers.microsoft.com/v2/global/en/corporate-pay/united-kingdom-corporate-pay.html

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,431,395 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
London
≈ $128k – $256k per year (Estimated) • In office • 3+ years exp • PhD • London
Python
AI/ML
Reinforcement Learning
SciPy
Computer Vision
AI Agents
NumPy
PyTorch
Post-training
Vision-Language-Action
Machine Learning
Game Dev
Unreal Engine
Lumen
Nanite
Management
Agile
Apply
≈ $97k – $194k per year (Estimated) • Hybrid • Master's Degree • Bristol
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
JAX
PyTorch
CUDA
Triton
InfiniBand
NVLink
Machine Learning
DevOps
Kubernetes
HPC
Apply
≈ $104k – $209k per year (Estimated) • Hybrid • Master's Degree • London
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
JAX
PyTorch
CUDA
Triton
InfiniBand
NVLink
Machine Learning
DevOps
Kubernetes
HPC
Apply
≈ $103k – $206k per year (Estimated) • Hybrid • Master's Degree • Cambridge
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
JAX
PyTorch
CUDA
Triton
InfiniBand
NVLink
Machine Learning
DevOps
Kubernetes
HPC
Apply
≈ $118k – $237k per year (Estimated) • Hybrid • Full-Time • London • Munich • Zurich
Python
AI/ML
Diffusion Models
Computer Vision
PyTorch
World Models
Machine Learning
DevOps
Datadog
Design
Webflow
Management
Miro
Stripe
Apply
Receptionist 1 day ago
≈ $13k – $24k per year (Estimated) • In office • Johannesburg
Management
Microsoft Office
Apply
≈ $18k – $34k per year (Estimated) • In office • Internship • Singapore
Analytics
Microsoft Excel
Design
Adobe Photoshop
Adobe Illustrator
AutoCAD
Management
Microsoft Office
Apply
≈ $30k – $48k per year (Estimated) • In office • Internship • Bachelor's Degree
Management
Microsoft Office
Apply
≈ $56k – $130k per year (Estimated) • Hybrid • London
DevOps
HPC
Analytics
Power BI
Management
BPMN
Apply
≈ $30k – $53k per year (Estimated) • In office • Sattledt
Management
Microsoft Office
Apply
$106k per year • In office • PhD • Cambridge
AI/ML
AI Agents
LLM Evaluation
Machine Learning
Apply
$106k per year • In office • PhD • Cambridge
AI/ML
JAX
AI Agents
PyTorch
LLM
Machine Learning
Apply
$99k – $162k per year • In office • PhD • Cambridge
AI/ML
AI Agents
Machine Learning
Apply
$102k – $202k per year • In office • 5+ years exp • Bachelor's Degree • Washington • Mountain View
Python
R
R
dplyr
AI/ML
Spark
Machine Learning
Apply
$120k – $235k per year • In office • 4+ years exp • Bachelor's Degree • Cambridge
AI/ML
PyTorch
Post-training
Model Distillation
Machine Learning
Apply
≈ $34k – $76k per year (Estimated) • In office • Full-Time • London
Apply
≈ $34k – $76k per year (Estimated) • In office • Full-Time • London
Apply
≈ $51k – $148k per year (Estimated) • In office • 5+ years exp • London • Derry • Belfast • Birmingham
Management
Outlook
Agile
Apply
≈ $88k – $179k per year (Estimated) • In office • Full-Time • London • Derry • Belfast • Birmingham
Apply
≈ $109k – $201k per year (Estimated) • In office • Full-Time • London
Apply
See all jobs
This is one of many
1,431,395 more open roles from verified company boards, updated every day.