698,071open jobs
41,308companies
99,448added this week
Browse all
Salary
$76k – $191k per year (Estimated)
Location
In office (Sydney)
Seniority
Senior · 8+ years exp
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

AI FactoryOS Operations  

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system. 

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against. 

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability. It also builds the shared services the estate's own operation depends on, and runs them. Operating the estate every day is what shows how the platform behaves under real load and under failure, and the function works with the engineering teams that build it to turn what it finds into permanent fixes and design improvements. 

Senior Kubernetes Platform Engineer

Role Summary 

Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Kubernetes Platform Engineer runsthe Kubernetes estatethat every product and every tenant runs on: the platform controller layer, the virtual cluster platform tenants are provisioned onto, the Kubernetes environment baselines used across the estate, the GPU integration layer, and the automated tenant onboarding and release pipeline. 

This is a hands-on senior role with deep technical expertise.This role owns the platform lifecycle execution across the fleet: keeping clusters healthy and current across the fleet, keeping tenant workloads running through upgrade and failure, and restoring control planes to service when they degrade.Automation is a first-class part of the role, in the operational tooling, guarded remediation and fleet orchestration that make estate-scale operation possible, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour. 

Key Responsibilities  

  • Responsible for the reliable operation, automation and continuous improvement of the multi-tenant Kubernetes platform that Firmus' products and tenants run on, spanning every site in the estate. 
  • Build the operational tooling, guarded remediation and orchestration that automate operations at fleet scale, and contribute operator and controller requirements, and code where agreed, to AI Infrastructure's platform backlog with the production evidence behind them. 
  • Execute the Kubernetes cluster lifecycle across the fleet, including provisioning, patching, upgrade and decommissioning, running the deployment and upgrade mechanisms built by AI Infrastructure through the agreed staged or canary path, and holding estate version compliance and retirement coordination. 
  • Operate and recover Kubernetes control planes carrying live tenant workload, including etcd state, certificate rotation, failed upgrades and corrupted resources. 
  • Operate the virtual cluster platform and multi-tenant isolation patterns that tenants are provisioned onto, and the automated tenant onboarding and release pipeline that lands new tenants safely and repeatably. 
  • Operate the GPU integration layer for Kubernetes (for example the NVIDIA GPU Operator), including device plugins, GPU scheduling and driver coordination. 
  • Diagnose and resolve scheduling failures, CNI and CSI faults, admission rejections and resource contention from first principles, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks. 
  • Run the Kubernetes baselines in production carrying the admission policy, workload identity and network policy content set by the Senior Platform Security Engineer, and hold the operational acceptance requirements those baselines have to meet before they enter production. 
  • Provide the deepest technical expertise for Kubernetes faults across the estate, diagnosingthe faults that require internals-level knowledgeto root cause, and driving the permanent fix to closure through AI Infrastructure, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results. 
  • Lead technical recovery during major Kubernetes incidents under the incident commander, drive the changes that remove repeat causes through the problem record, and share the after-hours escalation roster for the Kubernetes estate. 

Skills & Experience  

Required Skills 

  • Strong skills in platform and infrastructure engineering, with 8+ years of experience overall and substantial ownership of production Kubernetes platforms in a 24/7 environment. 
  • Deep experience operating Kubernetes at fleet scale, including cluster lifecycle, upgrades and multi-cluster management. 
  • Strong experience with Kubernetes control plane internals, including etcd, the API server, controllers, schedulers, and certificate and credential rotation. 
  • Strong experience writing Kubernetes controllers, operators or admission logic in a production setting. 
  • Experience with multi-tenant or virtual cluster patterns (for example vCluster or equivalent), including tenant isolation at the Kubernetes layer. 
  • Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling and driver coordination (for example the NVIDIA GPU Operator). 
  • Strong skills in infrastructure automation, infrastructure-as-code and GitOps practices (for example OpenTofu or Terraform, Ansible, Argo CD), with change delivered through peer review, automated testing and progressive rollout. 
  • Strongexperience with scripting or programming for operational automation and tooling, such as Go, Python or Bash. 
  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, post-incident review, and the production of runbooks that others can execute successfully. 
  • Solid understanding of Kubernetes security fundamentals, including admission control, workload identity and network policy. 
  • Clear technical judgement and communication skills, with the ability to produce documentation, design notes and escalations that other engineers can act on. 

Preferred Experience 

  • Experience operating Kubernetes for GPU or HPC workloads at scale. 
  • Experience with automated tenant or customer onboarding pipelines in a multi-tenant platform. 
  • Experience with vendor Kubernetes distributions or reference architectures for accelerated computing. 
  • Familiarity with DPU or SmartNIC-based networking as it relates to Kubernetes CNI design. 
  • Experience contributing to open-source Kubernetes ecosystem projects. 
  • A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training. 

Expected Outcomes  

  • Fleet-wide cluster upgrades executed through the staged path with no unplanned tenant-visible outage, and version compliance held across the estate. 
  • Tenant onboarding running as a routine operation rather than a project. 
  • Control plane recovery tested against live-equivalent conditions and executable by an engineer who did not write the procedure. 
  • Kubernetes escalations falling as operational automation and guarded remediation take on the common faults. 
  • Platform defects and operability gaps evidenced into the platform engineering backlog and closed permanently, with repeat causes falling. 

   

Location & Reporting  

Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required. 

On-call: The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function. 

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
698,071 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sydney
$65k – $159k per year (Estimated) • Remote/Hybrid • Full-Time • Auckland
SQL
DevOps
Terraform
CI/CD
AWS
Incident Management
Management
Agile
Apply
In office • PhD
Python
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
MLFlow
Google AI Studio
Machine Learning
DevOps
Azure
CI/CD
Kubernetes
Platform Engineering
Apply
$101k – $221k per year (Estimated) • Remote • Bachelor's Degree
Python
JavaScript
TypeScript
Frontend
React.js
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Apply
$12k – $26k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yekaterinburg
Python
SQL
1C
Databases
MS SQL
Analytics
Power BI
Apply
$173k – $203k per year • Equity • Remote • 5+ years exp • Master's Degree
Python
AI/ML
JAX
Multimodal AI
PyTorch
Time Series Forecasting
Edge AI
Machine Learning
Apply
$71k – $164k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Python
AI/ML
NCCL
InfiniBand
DevOps
Ansible
CI/CD
HPC
Linux
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Least Privilege
SIEM
Apply
$92k – $208k per year (Estimated) • In office • Bachelor's Degree • Sydney
Python
DevOps
Cilium
Kubernetes
Platform Engineering
eBPF
HPC
Cybersecurity
ISO 27001
Least Privilege
Cilium Tetragon
Apply
$92k – $218k per year (Estimated) • In office • Bachelor's Degree • Sydney
Python
DevOps
CI/CD
Kubernetes
HPC
Cybersecurity
Keycloak
Okta
HashiCorp Vault
ISO 27001
Authentik
SOC 2
Zero Trust
Least Privilege
Microsoft Entra ID
PKI
Apply
$155k – $277k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LLM
Management
Jira
Apply
$164k – $315k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
TensorFlow
PyTorch
InfiniBand
Apply
$132k – $236k per year (Estimated) • In office • 5+ years exp • Sydney
Apply
$132k – $236k per year (Estimated) • In office • 5+ years exp • Sydney
Apply
$101k – $189k per year (Estimated) • In office • Full-Time • 6+ years exp • Sydney
DevOps
IAM
Cybersecurity
Keycloak
LDAP
Apply
$68k – $89k per year • Remote/Hybrid • Full-Time • 5+ years exp • Sydney
AI/ML
Claude
Management
Notion
Marketing
HubSpot
Apply
$84k – $132k per year • Remote • Full-Time • 6+ years exp • Sydney
Apply
See all jobs
This is one of many
698,071 more open roles from verified company boards, updated every day.