368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$139k – $282k per year (Estimated)
Location
In office (Austin)
Seniority
Principal
Overview
Company
Impact
Profile match
Graphcore is a British semiconductor company founded in Bristol in 2016 that designs the Intelligence Processing Unit, a processor architected specifically for machine learning rather than adapted from graphics. Its chips place large amounts of memory directly on the die and expose fine-grained parallelism, an approach aimed at sparse and irregular models that map poorly onto conventional accelerators. The company sells IPU systems and the Poplar software stack to research and enterprise customers, and has operated as a wholly owned subsidiary of SoftBank Group since its acquisition in 2024.

About us

Graphcore is one of the world’s leading innovators in Artificial Intelligence compute.

It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.

As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.

Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects.

Job Summary

We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines.

The Team

The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers.

The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale.

As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems in environments where components and interfaces are still evolving. The role requires strong technical judgement, adaptability and an ability to work effectively across engineering teams.

Responsibilities and Duties

  • Convert agreed System Management direction into technical plans, engineering priorities and deliverable work for assigned areas.
  • Provide technical leadership for architecture and design decisions, building alignment across collaborating teams and documenting important technical trade-offs.
  • Act as a technical authority for assigned areas of System Management, leading the delivery of large and complex engineering initiatives and coordinating technical plans, dependencies, risks and decisions.
  • Provide technical direction and mentoring to engineers working across system management, hardware lifecycle management, deployment automation and production operations.
  • Take responsibility for key technical outcomes across the full software lifecycle, including design, implementation, automated testing, integration, deployment, observability and production readiness.
  • Identify systemic reliability, scalability and operability issues and lead practical improvements across the platform.
  • Collaborate with Hardware, Firmware, Platform Software and Datacenter Operations teams to diagnose system-level issues and improve end-to-end product behaviour.
  • Improve engineering standards and working practices, including CI/CD, Infrastructure-as-Code, automated testing, release safety and learning from operational incidents.
  • Act as a senior technical escalation point for complex issues while creating reusable knowledge, tooling and automation that reduce future operational effort.

Candidate Profile

Essential

  • Bachelor’s degree or equivalent practical experience in a relevant subject.
  • Substantial experience designing, building and operating Linux-based infrastructure or distributed systems.
  • Demonstrated experience providing technical leadership for complex engineering initiatives involving multiple teams or stakeholder groups.
  • Experience influencing architecture and technical decisions across team boundaries without relying on formal authority.
  • Experience translating broad technical goals into scoped plans, milestones, technical decisions, risks and delivery priorities.
  • Strong experience developing RESTful APIs and programming in Go, with Bash and Python used for systems automation.
  • Deep practical experience with Kubernetes, container runtimes and operating production workloads.
  • Hands-on experience with Infrastructure-as-Code, source control and CI/CD technologies such as Terraform/OpenTofu, Ansible, GitLab, GitHub Actions and Git.
  • Experience with hardware-management interfaces such as Redfish, IPMI or equivalent management systems.
  • Strong Linux systems engineering, troubleshooting and operational debugging capability.
  • Demonstrated ability to develop other engineers through technical mentoring, design reviews and coaching.
  • Clear communication skills, with the ability to persuade, build alignment and bring stakeholders together around practical technical outcomes.

Desirable

  • Experience using AI coding assistants effectively within professional engineering workflows.
  • Experience developing Kubernetes operators and custom resources.
  • Experience with High Performance Computing environments using SLURM, LSF or similar workload-management systems.
  • Experience with virtualisation technologies such as Open vSwitch, KVM and QEMU.
  • Experience with distributed object, block and file storage technologies such as Ceph.
  • Experience with monitoring and observability platforms such as Grafana, Prometheus, OpenSearch/Elasticsearch, Loki, Mimir or OpenTelemetry.
  • Experience configuring managed network switches using technologies such as EOS, SONiC or DNOS.
  • Experience supporting AI infrastructure or PyTorch workloads.

In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Austin
$71k – $149k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Jacksonville • King of Prussia
Go
Python
Java
Java
Flyway
Liquibase
DevOps
Amazon EKS
ArgoCD
CI/CD
Git
GitHub Actions
GitLab CI
Helm
Jenkins
JFrog Artifactory
Kubernetes
Platform Engineering
SRE
Terraform
AWS
GitHub
GitLab
Cybersecurity
SonarQube
Apply
$35k – $87k per year (Estimated) • In office • Freelance • 6+ years exp • Bachelor's Degree • Vienna
JavaScript
Kotlin
SQL
Swift
Mobile
Espresso
JUnit
Realm
DevOps
AWS
CI/CD
GitHub
GitHub Actions
Cybersecurity
GDPR
HIPAA
Management
Confluence
QA
Appium
BrowserStack
Charles Proxy
JMeter
Postman
TestNG
TestRail
XCUITest
Apply
$100k – $145k per year • In office • Full-Time • 5+ years exp • New York
JavaScript
SQL
TypeScript
Java
Java
Spring Boot
Databases
MySQL
Oracle
PostgreSQL
Frontend
Angular
React.js
Vue.js
DevOps
Git
Apply
$142k – $215k per year • In office • Full-Time • 12+ years exp • Princeton
JavaScript
TypeScript
Java
Java
Gradle
Hibernate
Maven
Spring Boot
Databases
Databricks
Snowflake
AI/ML
AI Agents
Claude
Claude Code
Frontend
Angular
React.js
Vue.js
DevOps
Amazon CloudWatch
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
API Gateway
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
IAM
Platform Engineering
Rest API
Kubernetes
Apply
$135k – $190k per year • In office • Full-Time • 12+ years exp • New York • Princeton
Databases
Apache Kafka
AI/ML
Hadoop
PyTorch
TensorFlow
DevOps
AWS
Azure
Azure DevOps
CI/CD
GCP
GitLab
GitLab CI
Jenkins
QA
Appium
Cypress
Playwright
Postman
Rest-Assured
Selenium
Apply
$31k – $64k per year (Estimated) • In office • Bengaluru
Management
Confluence
Apply
$80k – $155k per year (Estimated) • In office • Internship • Bachelor's Degree • Austin
Chips/EDA
Altium Designer
Apply
$88k – $183k per year (Estimated) • In office • London
C++
Python
C++
CMake
DevOps
Docker
Podman
HPC
QA
Pytest
Apply
$88k – $183k per year (Estimated) • In office • Cambridge
C++
Python
C++
CMake
DevOps
Docker
Podman
HPC
QA
Pytest
Apply
$88k – $183k per year (Estimated) • In office • Bristol
C++
Python
C++
CMake
DevOps
Docker
Podman
HPC
QA
Pytest
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$120k – $140k per year • In office • Full-Time • Charlotte • Raleigh • Dallas • Boston • New York
SQL
Analytics
ETL/ELT
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$100k – $180k per year • Equity • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Austin
DevOps
VMWare
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.