1,035,818open jobs
60,918companies
173,892added this week
Browse all
Salary
≈ $93k – $210k per year (Estimated)
Location
In office (Austin)
Seniority
Staff

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026. Graphcore scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Graphcore is a British semiconductor company founded in Bristol in 2016 that designs the Intelligence Processing Unit, a processor architected specifically for machine learning rather than adapted from graphics. Its chips place large amounts of memory directly on the die and expose fine-grained parallelism, an approach aimed at sparse and irregular models that map poorly onto conventional accelerators. The company sells IPU systems and the Poplar software stack to research and enterprise customers, and has operated as a wholly owned subsidiary of SoftBank Group since its acquisition in 2024.

About Graphcore

How often do you get the chance to build a technology that transforms the future of humanity?

Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business.

Graphcore recently joined SoftBank Group - bringing large and ongoing investment from one of the world’s leading backers of innovative AI companies.

Job Summary

We are seeking an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. The environment combines highly customized compute, high-performance networking, storage and supporting infrastructure, and will grow through multiple phases of deployment.

This is a rare opportunity to establish the reliability function for a new platform from the ground up. The platform and its operational model are being developed in parallel and will ultimately support a 24x7x365 production service with stringent availability requirements.

You will take the SRE organization from initial formation through production launch, stabilization and scale. This includes hiring and developing the team, defining the operating model, establishing production readiness and incident-management practices, and ensuring reliability and operability are engineered into the platform from the outset.

SRE is responsible for the operational capability required to run the platform reliably in production, while partnering with engineering teams that remain accountable for the reliability and operability of the systems they build.

This is not a purely managerial position. During the development and early production phases, the SRE Manager will be expected to work directly with engineering teams, develop a deep understanding of the platform, and participate in troubleshooting and incident response.

Over time, success will increasingly mean building the people, processes, automation, tooling, and operational discipline that allow the organization to operate effectively without depending on you for day-to-day escalation.

Responsibilities and Duties

  • Build the SRE Organization:
  • Build and develop the team from its initial formation through full 24x7x365 production operations, including defining roles, interviewing and hiring team members, establishing career expectations, and developing future technical leaders.
  • Mentor engineers and team leads, develop successors, and build an organization capable of operating effectively without depending on any single individual.
  • Work with leadership to forecast staffing requirements as the platform grows from initial deployment through full production scale.
  • Establish the Production Operating Model:
  • Define the operating model for an SRE organization, including staffing and coverage model, escalation paths, on-call responsibilities, incident management, handoffs, production access, change management, and operational readiness requirements.
  • Establish clear operational interfaces with Datacenter Operations, engineering teams, vendors, and other service owners.
  • Establish and continuously improve production readiness standards, runbooks, operational procedures, failure-mode documentation, escalation processes, and incident response practices with a strong emphasis on automation and engineering over manual operational work.
  • Develop training, cross-training, simulation, and production incident-response exercises to ensure the team can operate independently and confidently.
  • Engineer Reliability Into the Platform:
  • Embed with platform engineering teams during development to gain deep technical knowledge of the system and ensure reliability, serviceability, observability, and operational requirements are incorporated into the platform before production.
  • Lead the development of SLOs, operational health indicators, alerting standards, incident severity definitions, and reliability reporting appropriate for a large-scale production infrastructure service.
  • Build a culture in which recurring operational problems are engineered out through automation, improved observability, better platform design, and elimination of unnecessary toil.
  • Ensure the SRE organization can rapidly diagnose and mitigate issues across compute, networking, storage, and supporting infrastructure.
  • Develop strong technical depth within the team while maintaining access to specialist expertise in critical areas such as high-performance networking, storage, observability, and platform scheduling.
  • Lead Production Operations:
  • Lead or participate in major production incidents as necessary, particularly during platform development, launch, and early production.
  • Establish a blameless post-incident review process focused on identifying systemic improvements and ensuring corrective actions are completed.
  • Serve as the senior operational authority for the SRE organization and represent production reliability concerns in engineering and leadership discussions.

Required Skills and Experience

  • Significant experience leading or building an SRE, Production Engineering, Infrastructure Reliability, or comparable function supporting large-scale, highly available production infrastructure.
  • Experience taking a new or rapidly evolving platform through production readiness, launch, stabilization, and ongoing operation.
  • Experience building and operating sustainable 24x7x365 production support or on-call organizations.
  • Strong understanding of modern Site Reliability Engineering principles, including SLOs, incident management, observability, automation, toil reduction, capacity management, and production readiness.
  • Strong incident leadership experience, including managing high-severity, multi-team production incidents under significant time pressure.
  • A strong systems engineering background with a working understanding of Linux, networking, storage, distributed systems, automation, and production infrastructure.
  • Ability to operate effectively at both the leadership and hands-on technical levels, moving comfortably between organizational design, architecture discussions, troubleshooting, and incident response.
  • Demonstrated ability to lead, hire, mentor, develop, and retain a growing team of strong technical engineers at various levels of capability
  • Strong judgment regarding when to solve an immediate operational problem and when to invest in eliminating the underlying failure mode.
  • Strong written and verbal communication skills, particularly during incidents and when communicating technical risk to senior leadership.
  • Ability to manage conflicting priorities under a pressured environment

Desired but Not Required

Candidates are not expected to have experience in all the areas below. Experience in several would be particularly valuable:

  • High-performance networking technologies such as InfiniBand, RDMA, RoCE, or large-scale Ethernet fabrics.
  • Large-scale parallel or distributed storage systems.
  • Workload schedulers and orchestration platforms such as Kubernetes, Slurm, or comparable systems.
  • Production observability and service-level monitoring for complex distributed infrastructure.
  • Custom or early-generation hardware, firmware, or environments where hardware and software are developed concurrently.
  • Operational relationships and escalation processes involving datacenter operations and infrastructure/hardware/networking vendors.
  • Rapid capacity expansion, datacenter migration, or transitions between temporary and permanent production environments.

What Success Looks Like

  • The SRE organization is staffed, trained and capable of supporting the platform through production launch and subsequent growth.
  • Production readiness, observability, incident management and operational engineering practices are embedded into how the platform is developed and operated.
  • The team has the technical depth, automation and operating discipline to resolve most production issues independently, with leadership escalation reserved for genuinely exceptional events.

 

In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,035,818 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Management
Similar stack
Same company
Austin
$197k – $276k per year • Equity • In office • Full-Time • 5+ years exp • PhD • Seattle
Management
Agile
Apply
≈ $108k – $213k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Dallas • Phoenix • Columbus • Austin • Atlanta
Apply
Design Project Manager 4 months ago
≈ $97k – $192k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Dallas • Austin
Management
Microsoft Office
Apply
≈ $89k – $201k per year (Estimated) • In office • Full-Time • 1+ year exp • High School Diploma • Owasso
Apply
≈ $88k – $198k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Houston
MATLAB
MATLAB
Simulink
Management
Agile
Apply
≈ $107k – $223k per year (Estimated) • In office • 15+ years exp • Master's Degree • Louisville
C#
C#
.NET
AI/ML
AutoGen
LangChain
AI Agents
LLM
RAG
Multi-Agent Systems
DevOps
Azure
Docker
Kubernetes
Platform Engineering
Cybersecurity
HIPAA
Apply
≈ $67k – $163k per year (Estimated) • In office • Ipswich
Python
Java
PowerShell
Bash
DevOps
Terraform
Puppet
Ansible
VMWare
Kubernetes
Linux
Windows
Apply
$231k – $323k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Littleton • Huntsville • San Diego • Cupertino • Los Angeles
DevOps
Nomad
Docker
Kubernetes
Service Mesh
Apply
AI Engineer 2 days ago
≈ $120k – $227k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Dallas
SQL
C#
C#
.NET
Databases
Redis
AI/ML
LangGraph
AutoGen
LangChain
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
OpenAI Agents SDK
Agentic Workflows
Tool Use
DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
≈ $126k – $266k per year (Estimated) • In office • 15+ years exp • Louisville
Databases
Databricks
AI/ML
AI Agents
LLM Guardrails
DevOps
Azure
CI/CD
Platform Engineering
Cybersecurity
HIPAA
Apply
≈ $97k – $219k per year (Estimated) • In office • Bachelor's Degree • Milpitas
DevOps
HPC
Apply
Programme Manager 16 days ago
≈ $43k – $106k per year (Estimated) • In office • Bristol
Management
Agile
Scrum
Waterfall
Apply
≈ $41k – $101k per year (Estimated) • In office • Bristol
Analytics
Microsoft Excel
Management
Microsoft Office
Apply
≈ $137k – $256k per year (Estimated) • In office • Bachelor's Degree • Austin
C++
DevOps
CI/CD
Apply
≈ $60k – $98k per year (Estimated) • In office • Bristol
DevOps
CI/CD
Apply
≈ $104k – $236k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Databases
PostgreSQL
Oracle
AI/ML
Replicate
DevOps
Splunk
Dynatrace
AWS
Apply
$175k – $291k per year • Remote (United States) • Full-Time • 12+ years exp • Bachelor's Degree • Austin • Saint Louis
AI/ML
AI Agents
Feature Store
Knowledge Graph
LLM Guardrails
DevOps
GCP
Azure
CI/CD
AWS
Apply
$187k – $308k per year • Equity • Remote (United States) • Full-Time • 10+ years exp • San Francisco • Austin • Chicago • Seattle
DevOps
gRPC
OpenTelemetry
Datadog
Kustomize
Prometheus
GitLab CI
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Amazon EKS
Linux
DNS
Cybersecurity
PKI
Apply
$243k – $364k per year • Equity • Remote (United States) • Full-Time • 15+ years exp • Bachelor's Degree • San Jose • Austin • Seattle
Management
Agile
Apply
$229k – $345k per year • Equity • Remote (United States) • Full-Time • 8+ years exp • Austin • Dallas • Houston • Phoenix • Nashville
Python
JavaScript
Perl
DevOps
Linux
TCP/IP
DNS
BGP
Apply
See all jobs
This is one of many
1,035,818 more open roles from verified company boards, updated every day.