703,239open jobs
41,516companies
100,102added this week
Browse all
Salary
$93k – $214k per year (Estimated)
Location
In office (Sydney)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

AI FactoryOS Operations  

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system. 

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against. 

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next. 

Role Summary  

Firmus runs large-scale, state-of-the-artAI infrastructure built on the latest generation of GPU rack-scale systems and operatedas one estate to power the next generation of AI innovation. The Service Delivery Manager owns the incident and change management practices this operation runs on: the severity model and major incident command, the change calendar and change records, the runbook programme, production readiness review, and the operational reporting that keeps leadership and customers informed of service health, so that our services remain reliable and secure. 

The role sits at the fusion of ITIL and SRE practices. Incident, problemand change are managed with the rigour of formal service management, anddelivered with the methods of reliability engineering: measured against service level objectives, automated wherever automation makes response faster and safer, and continuously improved from what incidents reveal.  

The role owns the processes, not the technical decisions inside them. This role owns the standard, the record and the discipline that make those decisions consistent, visibleand auditable, and it owns the authorisation path that turns a technically complete release into an authorised production change.  

Key Responsibilities  

  • Own and run the incident management practice for the function: the severity model, major incident command, and the standards every team operates toduring an incident. 
  • Hold the incident record during major incidents, including the communications cadence, stakeholder updatesand the customer commitment, and coordinate blameless post-incident review through to closedactions. 
  • Own the change calendar and change enablement process across the estate, including approvals, scheduling, conflict managementand freeze periods, and maintainthe change record as the definitive account of what changed. 
  • Act as product owner for the runbook programme: prioritisewhich faults are converted into guarded, tested automation, hold authors to a standard the first-response team can execute unaided, and track whether the programmeis reducing escalation volume. 
  • Own the problem record: recurring faults, their root-cause status, and the engineering work required to close them out permanently. 
  • Own operational acceptance as the gateevery new or materially changed service passes through before it is declared supported. Coordinate the domain specialists and testers each acceptance needs, hold the technical sign-off from the accountable engineer and the operational acceptance from the service owner, and refer residual risk to the Head of AI FactoryOSOperations for the declared production risk position. 
  • Report service level and error budget performance across the portfolio, andmake error budget breaches visible as a reliability obligation on the engineering backlog rather than a number in a monthly pack. 
  • Produce operational reporting for leadership covering incident trends, change success rate, service health and the state of the runbook programme, and own customer communication on service health, incidentsand planned change. 
  • Own the collation of access, change and incident evidence for ISO 27001, SOC 2and enterprise customer due diligence, drawing on every team as an evidence source. 
  • Provide continuous cross-region coverage of the incident and change practice alongside peer Service Delivery Managers, and coach engineers onfollowing the practices consistently. 

Skills & Experience  

  • Significant experiencein service delivery, service managementor IT operations management in a large cloud provider, hyperscaleror infrastructure service provider context, including ownership of an incident, changeand problem management practice in a 24/7 environment. 
  • Proven experience running major incident management: severity classification, incident command, stakeholder communicationand post-incident review. 
  • Experience owning a change management or change enablement process for a technical environment, including a change calendar, approvalsand change records. 
  • Experience with service transition and acceptance into production, ensuring new or changed services are supportable before go-live. 
  • Comfortable working closely with technical teams and technical detail, with the credibility to hold engineering teams to an agreed operational standard. 
  • Experience producing operational reporting for leadership, covering incident trends, service level performanceand service health. 
  • Experience operatingunder formal compliance frameworks such as ISO 27001 or SOC 2, including contributing to audit and compliance evidenceproduction. 
  • Strong stakeholder management and communication skills, including customer-facing communication on service health, incidentsand planned change, with the ability to hold a consistent standard across multiple technical teams. 
  • Experience with an ITSM or incident management platform (for example ServiceNow, Jira Service Management or PagerDuty). 

Preferred Experience 

  • Experience in a data centre, cloud, HPC or AI infrastructure environment. 
  • Familiarity with GitOpsor infrastructure-as-code change workflows, sufficient to review a change record without authoring the change. 
  • ITIL certification or an equivalent formal service management qualification. 
  • A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training. 

Location & Reporting  

Location:  Based in Australia or Singapore, with travel to Australian AI Factory sites as required. 

On-call: The function runs 24/7. This role provides major incident command cover across regions alongside peer Service Delivery Managers, on a published roster. 

Reporting to: Reportsto the Head of AI FactoryOSOperations whilethe function is being established, working under broad directionwith a high degree of autonomyand direct access to the decision makers. As the function reaches its planned structure, the role will report to the Service Reliability Manager, with the Head of AI FactoryOSOperations remainingaccountable for the function. The scope, level and remit of the role do not change under either arrangement. 

Employment Basis

Permanent full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
703,239 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sydney
$46k per year • In office
DevOps
Incident Management
Apply
$13k – $28k per year (Estimated) • Remote/Hybrid
Management
Jira
Microsoft Office
Apply
$45k – $113k per year (Estimated) • Equity • Remote • Full-Time • 1+ year exp
Databases
Snowflake
AI/ML
Claude
ChatGPT
dbt
DevOps
GCP
Azure
AWS
Analytics
Fivetran
Management
Notion
Jira
Agile
Apply
$99k – $225k per year • In office • TS/SCI • Full-Time • 5+ years exp • High School Diploma • McLean
Python
JavaScript
C#
DevOps
GCP
Azure DevOps
Azure
CI/CD
Git
AWS
Management
Jira
Agile
ITIL
Apply
$170k – $220k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
DevOps
Incident Management
Web3
Rollup
Management
Notion
Google Workspace
Apply
$80k – $201k per year (Estimated) • In office • Full-Time • 8+ years exp • Sydney
Python
AI/ML
InfiniBand
NVLink
DevOps
Self-Healing
Apply
$71k – $164k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Python
AI/ML
NCCL
InfiniBand
DevOps
Ansible
CI/CD
HPC
Linux
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Least Privilege
SIEM
Apply
$92k – $208k per year (Estimated) • In office • Bachelor's Degree • Sydney
Python
DevOps
Cilium
Kubernetes
Platform Engineering
eBPF
HPC
Cybersecurity
ISO 27001
Least Privilege
Cilium Tetragon
Apply
$92k – $218k per year (Estimated) • In office • Bachelor's Degree • Sydney
Python
DevOps
CI/CD
Kubernetes
HPC
Cybersecurity
Keycloak
Okta
HashiCorp Vault
ISO 27001
Authentik
SOC 2
Zero Trust
Least Privilege
Microsoft Entra ID
PKI
Apply
$76k – $191k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Sydney
Python
JavaScript
Node JS
Node JS
Commander.js
DevOps
Terraform
Ansible
OpenTofu
etcd
CI/CD
GitOps
ArgoCD
Kubernetes
Platform Engineering
HPC
Apply
$60k – $140k per year (Estimated) • In office • Full-Time • Sydney
Apply
$111k – $238k per year (Estimated) • In office • Full-Time • Sydney
JavaScript
TypeScript
SQL
C#
C#
.NET
Frontend
React.js
DevOps
CI/CD
AWS
Docker
Kubernetes
Apply
$200k – $225k per year • Equity • In office • 5+ years exp • Sydney
Databases
Snowflake
Databricks
AI/ML
AI Agents
Mobile
Algolia
DevOps
GCP
Cybersecurity
Okta
Crowdstrike
SentinelOne
Carbon Black
Management
Obsidian
Marketing
Salesforce
Apply
$93k – $209k per year (Estimated) • In office • Full-Time • 5+ years exp • Sydney • Canberra • Adelaide • Perth • Melbourne
Apply
In office • Full-Time • PhD • Melbourne • Sydney
Apex
Apex
Salesforce Data Cloud
AI/ML
AI Agents
Agentforce
Apply
See all jobs
This is one of many
703,239 more open roles from verified company boards, updated every day.