Firmus Technologies
Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
AI FactoryOS Operations
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.
Role Summary
Firmus runs large-scale, state-of-the-artAI infrastructure built on the latest generation of GPU rack-scale systems and operatedas one estate to power the next generation of AI innovation. The Service Delivery Manager owns the incident and change management practices this operation runs on: the severity model and major incident command, the change calendar and change records, the runbook programme, production readiness review, and the operational reporting that keeps leadership and customers informed of service health, so that our services remain reliable and secure.
The role sits at the fusion of ITIL and SRE practices. Incident, problemand change are managed with the rigour of formal service management, anddelivered with the methods of reliability engineering: measured against service level objectives, automated wherever automation makes response faster and safer, and continuously improved from what incidents reveal.
The role owns the processes, not the technical decisions inside them. This role owns the standard, the record and the discipline that make those decisions consistent, visibleand auditable, and it owns the authorisation path that turns a technically complete release into an authorised production change.
Key Responsibilities
- Own and run the incident management practice for the function: the severity model, major incident command, and the standards every team operates toduring an incident.
- Hold the incident record during major incidents, including the communications cadence, stakeholder updatesand the customer commitment, and coordinate blameless post-incident review through to closedactions.
- Own the change calendar and change enablement process across the estate, including approvals, scheduling, conflict managementand freeze periods, and maintainthe change record as the definitive account of what changed.
- Act as product owner for the runbook programme: prioritisewhich faults are converted into guarded, tested automation, hold authors to a standard the first-response team can execute unaided, and track whether the programmeis reducing escalation volume.
- Own the problem record: recurring faults, their root-cause status, and the engineering work required to close them out permanently.
- Own operational acceptance as the gateevery new or materially changed service passes through before it is declared supported. Coordinate the domain specialists and testers each acceptance needs, hold the technical sign-off from the accountable engineer and the operational acceptance from the service owner, and refer residual risk to the Head of AI FactoryOSOperations for the declared production risk position.
- Report service level and error budget performance across the portfolio, andmake error budget breaches visible as a reliability obligation on the engineering backlog rather than a number in a monthly pack.
- Produce operational reporting for leadership covering incident trends, change success rate, service health and the state of the runbook programme, and own customer communication on service health, incidentsand planned change.
- Own the collation of access, change and incident evidence for ISO 27001, SOC 2and enterprise customer due diligence, drawing on every team as an evidence source.
- Provide continuous cross-region coverage of the incident and change practice alongside peer Service Delivery Managers, and coach engineers onfollowing the practices consistently.
Skills & Experience
- Significant experiencein service delivery, service managementor IT operations management in a large cloud provider, hyperscaleror infrastructure service provider context, including ownership of an incident, changeand problem management practice in a 24/7 environment.
- Proven experience running major incident management: severity classification, incident command, stakeholder communicationand post-incident review.
- Experience owning a change management or change enablement process for a technical environment, including a change calendar, approvalsand change records.
- Experience with service transition and acceptance into production, ensuring new or changed services are supportable before go-live.
- Comfortable working closely with technical teams and technical detail, with the credibility to hold engineering teams to an agreed operational standard.
- Experience producing operational reporting for leadership, covering incident trends, service level performanceand service health.
- Experience operatingunder formal compliance frameworks such as ISO 27001 or SOC 2, including contributing to audit and compliance evidenceproduction.
- Strong stakeholder management and communication skills, including customer-facing communication on service health, incidentsand planned change, with the ability to hold a consistent standard across multiple technical teams.
- Experience with an ITSM or incident management platform (for example ServiceNow, Jira Service Management or PagerDuty).
Preferred Experience
- Experience in a data centre, cloud, HPC or AI infrastructure environment.
- Familiarity with GitOpsor infrastructure-as-code change workflows, sufficient to review a change record without authoring the change.
- ITIL certification or an equivalent formal service management qualification.
- A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training.
Location & Reporting
Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
On-call: The function runs 24/7. This role provides major incident command cover across regions alongside peer Service Delivery Managers, on a published roster.
Reporting to: Reportsto the Head of AI FactoryOSOperations whilethe function is being established, working under broad directionwith a high degree of autonomyand direct access to the decision makers. As the function reaches its planned structure, the role will report to the Service Reliability Manager, with the Head of AI FactoryOSOperations remainingaccountable for the function. The scope, level and remit of the role do not change under either arrangement.
Employment Basis
Permanent full-time
Diversity
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

