702,996open jobs
41,505companies
99,778added this week
Browse all
Salary
$80k – $201k per year (Estimated)
Location
In office (Sydney)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

AI FactoryOS Operations  

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system. 

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against. 

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next. 

Role Summary  

Firmus runs large-scale, state-of-the-artAI infrastructure built on the latest generation of GPU rack-scale systems and operatedas one estate to power the next generation of AI innovation. The Senior Platform Reliability Engineer, Fabric and Interconnect, owns the reliability of the fabrics this estate runs on: the GPU-to-GPU interconnect domains, the high-performance network fabrics carrying training and inference traffic, and the DPU-based host networking that binds compute to the rest of the platform. 

This is a hands-on senior role with deep technical expertise. Automation is a first-class part of the role: the team builds and maintains the guarded automation and remediation tooling that turn manual fabric response into a self-healing capability, and the role engages fabric vendors at engineering level, reproducing faults to their standard and holding them to their answers. 

Key Responsibilities  

  • Responsible for the reliable operation, automation and continuous improvement of the estate's GPU interconnect and network fabrics (for example NVLinkand NVSwitchdomains, InfiniBandand Spectrum-X). 
  • Build and maintainthe guarded automation and remediation tooling for fabric faults, contributing to the software-driven remediation of AI clusters, including fault isolation and fabric reconvergence. 
  • Diagnose and tune performance across the interconnect stack, from applicationcollective communication down to link level, working with technologies including NVLink, InfiniBand, RoCE and congestion control tuning. 
  • Operate DPU-based host networking across the fleet, including offload path configuration and driver and firmware compatibility. 
  • Execute firmware upgrade waves, fabric expansions and capacity changes to the supported paths and scaling patterns defined by AI Infrastructure, owning the production window, the staged or canary path, verification against declared success criteria, and rollback execution. 
  • Provide the deepest technical expertisefor fabric and interconnect faults, correlating a collective communication failure to a specific physical link and diagnosingthe faults that require internals-level knowledgeto root cause, and driving the permanent fix to closure through AI Infrastructure. 
  • Lead vendor escalations at engineering level, reproducing faults to the vendor's standard and pushing back credibly when a diagnosis does not explain the observed behaviour. 
  • Lead technical recovery during major fabric incidents, drive the changes that remove repeat causes, share the follow-the-sun on-call roster, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooksand performance results. 

Skills & Experience  

  • Strong skills in high-performance networking and systems engineering, with 8+ years of experience including substantial ownership of production network or interconnect infrastructure in a 24/7 environment. 
  • Deep operational experience with high-performance GPU interconnect fabrics (for example NVLinkand NVSwitch), including domain topology and failure diagnosis. 
  • Extensive experience with high-performance networking fabrics (for example InfiniBand or RoCE-based Ethernet such as NVIDIA Spectrum-X), including routing internals and congestion control tuning. 
  • Experience with DPU or SmartNIC-based host networking, including offload paths and driver and firmware coordination. 
  • Experience planning and executing firmware upgrade waves and capacity expansions on production fabrics with defined rollback. 
  • Strong skills in infrastructure automation and infrastructure-as-code practices, with change delivered through peer review and progressive rollout. 
  • Practical experience with scripting or programming for operational automation and tooling, such as Python, Goor Bash. 
  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, vendor escalation at engineering level, and the production of runbooks that others can execute successfully. 
  • Clear technical judgement and communication skills, with the ability to explain complex fabric failures to engineersand non-specialists. 

Preferred Experience 

  • Experience operating fabrics for large-scale distributed training or inference workloads. 
  • Experience with NVIDIA rack-scale or multi-node GPU systems and their interconnect topology. 
  • Experience in a multi-tenant service provider, cloudor colocation environment. 
  • Knowledge of data centre and hardware fundamentals, including cabling and optics, firmware managementand hardware fault workflows. 
  • Relevant vendor certification. 

Location & Reporting  

Location:  Based in Australia or Singapore, with travel to Australian AI Factory sites as required. 

On-call: The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function. 

Reporting to: Reportsto the Head of AI FactoryOSOperations whilethe function is being established, working under broad directionwith a high degree of autonomyand direct access to the decision makers. As the function reaches its planned structure, the role will report to the Infrastructure Operations Manager, with the Head of AI FactoryOSOperations remainingaccountable for the function. The scope, level and remit of the role do not change under either arrangement. 

Employment Basis

Permanent full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
702,996 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sydney
$175k – $342k per year (Estimated) • In office • 8+ years exp • Los Angeles
Python
C++
C++
Qt
Apply
$99k – $225k per year • In office • TS/SCI • Full-Time • 5+ years exp • High School Diploma • McLean
Python
JavaScript
C#
DevOps
GCP
Azure DevOps
Azure
CI/CD
Git
AWS
Management
Jira
Agile
ITIL
Apply
$105k – $204k per year (Estimated) • In office • Full-Time • 6+ years exp • Associate's Degree • Salem
Python
SQL
Databases
Snowflake
Analytics
Power BI
ETL/ELT
Azure Data Factory
Dimensional Modeling
Apply
$147k – $231k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Houston • Taipei
Python
C++
AI/ML
Edge AI
DevOps
RTOS
CI/CD
Platform Engineering
Windows
Management
Agile
Apply
Data Scientist 5 hours ago
$110k – $176k per year • In office • Full-Time • Bachelor's Degree • New York
Python
SQL
Databases
Databricks
AI/ML
Hadoop
Spark
Machine Learning
Apply
$93k – $214k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Sydney
DevOps
PagerDuty
Incident Management
Error Budget
HPC
Cybersecurity
ISO 27001
SOC 2
Management
Jira
ServiceNow
ITIL
ITSM
Apply
$71k – $164k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Python
AI/ML
NCCL
InfiniBand
DevOps
Ansible
CI/CD
HPC
Linux
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Least Privilege
SIEM
Apply
$92k – $208k per year (Estimated) • In office • Bachelor's Degree • Sydney
Python
DevOps
Cilium
Kubernetes
Platform Engineering
eBPF
HPC
Cybersecurity
ISO 27001
Least Privilege
Cilium Tetragon
Apply
$92k – $218k per year (Estimated) • In office • Bachelor's Degree • Sydney
Python
DevOps
CI/CD
Kubernetes
HPC
Cybersecurity
Keycloak
Okta
HashiCorp Vault
ISO 27001
Authentik
SOC 2
Zero Trust
Least Privilege
Microsoft Entra ID
PKI
Apply
$76k – $191k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Sydney
Python
JavaScript
Node JS
Node JS
Commander.js
DevOps
Terraform
Ansible
OpenTofu
etcd
CI/CD
GitOps
ArgoCD
Kubernetes
Platform Engineering
HPC
Apply
$60k – $140k per year (Estimated) • In office • Full-Time • Sydney
Apply
$111k – $238k per year (Estimated) • In office • Full-Time • Sydney
JavaScript
TypeScript
SQL
C#
C#
.NET
Frontend
React.js
DevOps
CI/CD
AWS
Docker
Kubernetes
Apply
$200k – $225k per year • Equity • In office • 5+ years exp • Sydney
Databases
Snowflake
Databricks
AI/ML
AI Agents
Mobile
Algolia
DevOps
GCP
Cybersecurity
Okta
Crowdstrike
SentinelOne
Carbon Black
Management
Obsidian
Marketing
Salesforce
Apply
$93k – $209k per year (Estimated) • In office • Full-Time • 5+ years exp • Sydney • Canberra • Adelaide • Perth • Melbourne
Apply
In office • Full-Time • PhD • Melbourne • Sydney
Apex
Apex
Salesforce Data Cloud
AI/ML
AI Agents
Agentforce
Apply
See all jobs
This is one of many
702,996 more open roles from verified company boards, updated every day.