388,323open jobs
10,252companies
48,140added this week
Browse all
Salary
$79k – $198k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

ROLE SUMMARY  

Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to join our Engineering and Technology team. You will establishhow we measure, validateand communicate the health of GPU infrastructure used for customer and internal workloads. You will define trusted health signals and service-readiness criteria, and turn them into reusable dashboards, alerts, queries, diagnostic checksand operational guidance. Your work will help commissioning, infrastructure and operations teams bring capacity online safely, identifydegradation early and recover from failures quickly. You will also make knowledge self-service by publishing clear reference implementations, runbooksand AI-ready operational knowledge that other teams can use and extend. 

KEYRESPONSIBILITIES  

  • The Health Standard  

    Define GPU and host health criteria for customer and internal workloads, and the service-readiness gates that repair and capacity workflows depend on. Cover what stops or corrupts AI jobs: GPUs falling off the bus, XID events, ECC and memory faults, NVLink/NVSwitch degradation, thermal and power capping, NCCL and collective failures, silent data corruption, and stragglers running below fleet baseline. 

  • Reference Implementations  

    Publish golden dashboards, alerts, PromQL/LogQL queries and health checks that other teams adopt and extend. Your output is the standard and the examples. The alerts must be specific, low-noise, and with a clear next action. 

     

  • Diagnostics with Real Pass/Fail Criteria  

    DCGM checks, NCCL and bandwidth tests, stress and burn-in, and validation jobs that confirm a server matches expected performance. Others must be able to run them without you. 

  • Fault Isolation  

    Separate a bad GPU from a cooling, host, network or power-limit problem using host and BMC/management-interface telemetry together, including when the same signature appears across many servers. A clean management view means nothing if the host is throwing faults. 

     

  • Detection of the Failures that Don't Crash  

Rising correctable ECC counts, NVLink retries, thermal slowdown, XID patterns, wrong results with no error. Keep the knowledge current: what each signal means, what to do next on the machine, and where the operation team must make the final decision. 

  • Partnership and Escalation

    Work with commissioning on bring-up and acceptance, operations on break-fix, infrastructure on what healthy hardware looks like, and the telemetry owners on making your signals production-grade. Run technical sessions with customer teams on the telemetry they need. Join GPU and host incidents, including debugging on the server and convert every finding into a reusable check, alert or runbook. Give engineering leadership a straight read on fleet health and risk. 

SKILLS AND EXPERIENCE 

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience. 
  • 7+ years in GPU, HPC, AI infrastructure or closely related large-scale systems engineering environments, including ownership of health monitoring or diagnostics used by customers or internal teams. 
  • Experience with production GPU fault diagnosis. You'vefound genuine faults through XID events, ECC/memory errors, NVLinkissues, power or thermal limits and caught at least one before the job died. 
  • Strong Linux and server fundamentals. You can debug on the machine and reason across GPU, CPU, memory, PCIe, power and cooling. You automate in Python or similar. 
  • Hands-on with GPU diagnostics and validation such as DCGM, NCCL/collective tests, stress testing and you've turned them into checks other people run. 
  • Experience analysing infrastructure telemetry, building trusted dashboards and alerts. Savvy with PromQL, LogQL, Grafana or comparable tooling. Able to follow an unexpected signal until thereis a cause and can turn that into an alert or check others will trust. 
  • Uses AI tools as a normal part of analysis and build work. Can structure health knowledge and build AI skills so operation team withAI assistants can use iteffectively to recover from incidents. 
  • Willing to take part in the incident-response on-call rotation. 
  • Willing to travel overseas occasionally when the role requires it. 
  • Clear and effective written and verbal communication in English. 

Highly Desirable Experiences

  • Production experience on large GPU systems with tight GPU-to-GPU interconnect. 
  • Multiple sites or GPU generations, especially preparing health monitoring before a new GPU class entered production. 
  • Background at an AI cloud, GPU cloud or HPC centre running customer workloads. 
  • Experience working with data or ML engineers on GPU/host failure prediction from health signals. 
  • Built a structured operational knowledge so AI assistants can use it effectively during incident recovery. 

Location & Reporting

  • Location: Singapore
  • Report to: Senior Manager, Platform Engineering

Employment Basis

Full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
388,323 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$41k – $87k per year (Estimated) • In office • Contractor • 6+ years exp • Noida
JavaScript
Python
TypeScript
Java
Kotlin
Java
Hibernate
Spring Boot
Spring Cloud
Spring Security
Testcontainers
Kotlin
Mockito
Databases
Apache Kafka
InfluxDB
Kafka
PostgreSQL
RabbitMQ
Redis
TimescaleDB
AI/ML
Computer Vision
Feature Store
LLM
Time Series Forecasting
Frontend
Next.js
React.js
Mobile
JUnit
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
GitLab
GitLab CI
GitOps
Grafana
gRPC
Jaeger
Jenkins
Kubernetes
Loki
OpenTelemetry
Prometheus
Pulumi
Terraform
Cybersecurity
ISO 27001
SOC 2
SonarQube
IoT
MQTT
OPC UA
QA
Gatling
JMeter
k6
Swagger
Apply
Remote/Hybrid • Full-Time • Palo Alto
Python
AI/ML
Reinforcement Learning
Robotics
MuJoCo
Reinforcement Learning
ROS2
Sim-to-Real
Apply
In office • Full-Time • 15+ years exp • Master's Degree • Hangzhou
Python
SQL
Databases
Databricks
Snowflake
AI/ML
A2A
AI Agents
Anthropic
Copilot
Human-in-the-Loop
LLMOps
Model Context Protocol
OpenAI Codex
DevOps
AWS
Azure
FinOps
GCP
GitHub
Apply
$58k – $141k per year (Estimated) • In office • Full-Time • Bachelor's Degree
C++
Go
Python
Rust
DevOps
Debian
Kubernetes
Ubuntu
Apply
$72k – $173k per year (Estimated) • In office • Full-Time • Bachelor's Degree
C++
Go
Python
Rust
DevOps
Debian
Kubernetes
Ubuntu
Apply
$83k – $188k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Sydney
Bash
Python
DevOps
AWS
Azure
GCP
Kubernetes
Platform Engineering
Cybersecurity
Auth0
Calico
Checkov
CIS Benchmarks
Cosign
Falco
HashiCorp Boundary
HashiCorp Vault
HIPAA
ISO 27001
kube-bench
Kyverno
OWASP Top 10
PCI DSS
Snyk
SOC 2
Teleport
Trivy
Cryptography
Vault
Apply
$77k – $218k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Singapore
Go
Node JS
Python
SQL
JavaScript
Databases
Apache Kafka
ClickHouse
Kafka
Frontend
GraphQL
Mobile
JUnit
DevOps
Ansible
ArgoCD
AWS
Azure
CI/CD
Configuration Management
Docker
GCP
GitHub
GitHub Actions
Grafana
Jenkins
Loki
Platform Engineering
Prometheus
Thanos
Kubernetes
Cybersecurity
ISO 27001
SOC 2
Game Dev
Meta Quest SDK
QA
Cypress
k6
Pytest
Apply
$126k – $280k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Singapore
Go
Python
AI/ML
AI Agents
DevOps
CI/CD
Incident Management
Kubernetes
Platform Engineering
SLI/SLO/SLA
Cybersecurity
ISO 27001
SOC 2
Apply
$117k – $254k per year (Estimated) • In office • Contractor • Bachelor's Degree • Singapore
DevOps
CI/CD
GitOps
Kubernetes
Platform Engineering
Cybersecurity
ISO 27001
SOC 2
Apply
$84k – $199k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sydney
Bash
Go
Python
Databases
ElasticSearch
AI/ML
InfiniBand
DevOps
Ansible
ArgoCD
CI/CD
Cilium
etcd
GitHub
GitHub Actions
GitLab
GitLab CI
GitOps
Grafana
Jenkins
kubeadm
Kubernetes
Loki
OpenTelemetry
Platform Engineering
Prometheus
Service Mesh
Terraform
Cybersecurity
Calico
Kyverno
OPA Gatekeeper
Open Policy Agent
Apply
$77k – $193k per year (Estimated) • In office • Full-Time • 3+ years exp • Singapore
Python
DevOps
Amazon EKS
AWS
Azure AKS
CI/CD
CloudFormation
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Jenkins
Kubernetes
Prometheus
Terraform
Azure
GitHub
GitLab
IAM
Apply
$107k – $263k per year (Estimated) • Equity • Remote • 10+ years exp • Singapore
Apply
Platform Engineer 1 day ago
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
DevOps
Amazon EC2
Amazon EKS
Amazon S3
AWS
CI/CD
Docker
Helm
IAM
Kubernetes
Terraform
Apply
$59k – $134k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • Singapore
C++
Apply
$116k – $182k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Houston
C++
Python
DevOps
Git
Apply
See all jobs
This is one of many
388,323 more open roles from verified company boards, updated every day.