489,618open jobs
16,287companies
72,295added this week
Browse all
Salary
$97k – $201k per year (Estimated)
Location
In office (Oxford)
Seniority
Senior
Overview
Company
Impact
Profile match
Ellison Medical Institute. The Ellison Medical Institute strives to spark innovation, leverage technology, and drive interdisciplinary, patient-centered research to continually enhance health, reimagine and redefine cancer care, and transform lives.

Join us at EIT:

At the Ellison Institute of Technology (EIT), we’re on a mission to translate scientific discovery into real world impact. We bring together visionary scientists, technologists, policy makers, and entrepreneurs to tackle humanity’s greatest challenges in four transformative areas:

  • Health, Medical Science & Generative Biology
  • Food Security & Sustainable Agriculture
  • Climate Change & Managing CO₂
  • Artificial Intelligence & Robotics

This is ambitious work - work that demands curiosity, courage, and a relentless drive to make a difference. At EIT, you’ll join a community built on excellence, innovation, tenacity, trust, and collaboration, where bold ideas become real-world breakthroughs. Together, we push boundaries, embrace complexity, and create solutions to scale ideas for lab to society. Explore more at www.eit.org

Your Role:

Join our SciComp team to build the cloud and compute foundation that enables scientific breakthroughs. Deliver reliable, secure platforms and self-service guardrails that accelerate experimentation and turn ideas into results - faster, at scale, and with confidence.

Your Responsibilities:

  • Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management.
  • Drive systems design and implementation for high-throughput data paths, optimising I/O, caching, and data locality across compute and storage (including our current Lustre implementation).
  • Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference.
  • Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments.
  • Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines.

Requirements

Essential Skills, Qualifications & Experience:

  • Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale
  • A proactive, autonomous approach to systems design and the proven ability and desire to ideate, co-create and implement optimal solutions
  • Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems
  • Expertise with high-throughput storage systems for ML/HPC workloads
  • Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks
  • A solid grasp of IaC and CI/CD practices (e.g., Terraform, Argo CD)

Benefits

We offer the following salary and benefits:

  • Competitive salary (dependent on experience) + travel allowance + bonus
  • Enhanced holiday. Our annual leave allowance is 25 days plus 8 bank holidays and an additional 3 days between Christmas and New Year. You will also have the opportunity to purchase an additional 5 days annual leave in January and July.
  • Pension - Employer contribution 7.5%, minimum employee contribution 5%
  • Life Assurance.
  • Income Protection
  • Private Medical Insurance as standard for you, your partner and any dependents. Including hospital Cash Plan
  • Employee discounts
  • Electric car scheme
  • Nursery Salary Sacrifice scheme
  • Cycle to Work Scheme
  • Family Planning
  • Neurodiversity support including advise and assessments
  • Coaching & Therapy services

Working together - what it involves:

You must have the right to work permanently in the UK with a willingness to travel as necessary. In certain cases, we can consider sponsorship, and this will be assessed on a case-by-case basis.

You will live in, or within easy commuting distance of, Oxford (or be willing to relocate) and can commit to being onsite at our Oxford office, a minimum of 3 days per working week.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
489,618 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Oxford
$179k – $205k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • McLean • Plano
Python
Java
Scala
Python
Dask
AI/ML
Spark
Scikit-learn
TensorFlow
PyTorch
Explainable AI
DevOps
GCP
Azure
CI/CD
AWS
Management
Agile
Apply
$315k – $359k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • McLean • Plano
Python
C++
Scala
Python
Dask
AI/ML
Scikit-learn
Pandas
NumPy
DevOps
HPC
Apply
$269k – $307k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Plano • McLean
Python
C++
Scala
Python
Dask
AI/ML
Scikit-learn
Pandas
NumPy
DevOps
HPC
Apply
$155k – $256k per year • Equity • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Park
Python
DevOps
GCP
CI/CD
Docker
Kubernetes
Management
Agile
Scrum
QA
JMeter
Locust
Apply
$212k – $308k per year • Equity • Remote • Full-Time • 5+ years exp • Philadelphia • New York
DevOps
Terraform
AWS
Chips/EDA
PoC Library
Apply
$89k – $171k per year (Estimated) • In office • Full-Time • Oxford
Python
Java
DevOps
Terraform
GCP
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Cybersecurity
Kyverno
Apply
$69k – $186k per year (Estimated) • In office • Full-Time • Oxford
DevOps
Vector
Apply
In office • Full-Time • Oxford
Apply
In office • Full-Time • Oxford
Apply
In office • Full-Time • Oxford
DevOps
Vector
Apply
Associate Manager 2 days ago
$49k – $61k per year • In office • Full-Time • 5+ years exp • Master's Degree • Oxford
Apply
In office • PhD • Oxford
C#
C++
C#
.NET
AI/ML
Text-to-Speech
DevOps
Vercel
CI/CD
Management
Google Docs
Stripe
Apply
In office • PhD • Oxford
AI/ML
Text-to-Speech
DevOps
GCP
Vercel
Azure
AWS
Docker
Kubernetes
Management
Google Docs
Stripe
Apply
$35k – $72k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Oxford
Apply
$77k – $145k per year (Estimated) • In office • Full-Time • Oxford
Apply
See all jobs
This is one of many
489,618 more open roles from verified company boards, updated every day.