368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$78k – $197k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior · 7+ years exp
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. This co-designed approach from model to grid allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Why you’ll love working here

At Firmus, you’ll work at the intersection of sustainability and artificial intelligence in a fast-paced environment powered by next-generation technology. You’ll be helping to transform an entire industry - and you’ll feel it every day.

Our team is made up of true innovators and leaders in their fields, and as an emerging company, you won’t be lost in a crowd. You’ll work closely with the founders, build a strong network, and see the impact of your work first-hand as we democratise AI tools for everyone - more sustainably and more affordably.

We believe great things happen when people from diverse backgrounds come together to do their best work and be their authentic selves. We are proud to be an equal opportunity employer.

ROLE

Firmus Technologies is seeking a Senior Platform Engineer to join our Engineering and Technology team. You will drivethe design and implementation of our MLOpscapability. You will also collaborate with other engineers and make technical decisionson scalingFirmus AI factory platform engineering capabilitiesto planet scale,from IaC, container orchestration, observability, self-service portal to platform security.This role is ideal for a self-starter with passion for building things from first principles. You naturally break down complex problems into their fundamental truths to uncover novel and elegant solutions - rather than relying on conventional patterns. 

KEY RESPONSIBILITIES 

  • Build MLOps capabilities from the ground up, enabling reproducible, scalable, and secure ML workflows across internal and customer-facing environments.  
  • Continuously improve our DevOps platform to ensure reliability, scalability, security, and seamless integration with CI/CD pipelines and infrastructure services. 
  • Design, implement, operate and secure Kubernetes-based production infrastructure for high reliability, performance and security, including clusters supporting NVIDIA GB300 NVL72 systems with NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet.  
  • Develop world-class observability platforms for internal and external customers
  • Integrate Firmus central services with NVIDIA’s software stack, including Mission Control, NETQ, UFM, and NMX. 
  • Lead the enhancement and evangelism of internal platform products that provide cohesive, composable, secure-by-default, and low-friction self-service experiences that accelerates time to market and reduce engineers' cognitive load. 
  • Drive incident response efforts, participate actively in the on-call rotation, and lead detailed Root Cause Analysis (RCA) to continuously improve system reliability, operational maturity, and incident handling processes. 

SKILLS AND EXPERIENCE   

  • Bachelor's degree in computer science or a related technical field. 
  • 7+ years of experience as Platform Engineer, Site Reliability Engineer, DevOps engineer, MLOps Engineer or Observability Engineer. 
  • Demonstrated strong proficiency on the following areas:  
    • Infrastructure-as-Code, configuration management and CI/CD (e.g., Terraform, Ansible, GitHub Actions, Jenkins, ArgoCD). 
    • Containerization technologies (e.g., Docker), Kubernetes networking and cluster management, including upgrades and troubleshooting. 
    • Observability stack design and scaling (e.g., Loki, Grafana, Tempo, Prometheus, Thanos, ClickHouse). 
    • Telemetry solutions using various technology (e.g., Redfish, gNMI, SNMP, eBPF, streaming analytics). 
    • Unified telemetry collection with OpenTelemetry. 
    • Compliance automation (e.g., OPA, Kyverno). 
  • Competent in scripting and programming skills (e.g., Bash, Python, Go). 
  • Systems knowledge on Linux internals, networking stacks, and distributed storage. 
  • Clear and effective English communication, written and spoken. 
  • Bonus: Experience in high-growth startups or regulated industries with robust security and data privacy requirements, including SOC 2 Type 2 and ISO 27001. 
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
Bash
C++
Java
Python
DevOps
CI/CD
Git
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
C++
Python
DevOps
CI/CD
SpaceTech
NASA cFS
Apply
In office • Part-Time • 2+ years exp • Bachelor's Degree • Ness Ziona
C#
Java
AI/ML
Copilot
Cursor
DevOps
CI/CD
GitHub
Apply
$113k – $136k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Westminster
C
C
U-Boot
DevOps
CI/CD
Apply
$60k – $148k per year (Estimated) • In office • Full-Time • Launceston
DevOps
HPC
IoT
OPC UA
Apply
$83k – $193k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Sydney
Apply
$149k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LLM
Management
Jira
Apply
$181k – $331k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
PyTorch
TensorFlow
InfiniBand
Apply
$166k – $356k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
Fine-tuning
Apply
$117k – $251k per year (Estimated) • Remote/Hybrid • Full-Time • Singapore
Apply
$74k – $126k per year (Estimated) • In office • Full-Time • Singapore
Python
Apply
$88k – $191k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Singapore
C++
Java
Kotlin
Python
Mobile
JUnit
DevOps
Git
gRPC
Jenkins
JFrog Artifactory
Shift-Left
Cybersecurity
Shift-Left Security
QA
Pytest
Robot Framework
TestNG
Apply
Senior AI Architect 5 hours ago
$138k – $304k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
LangGraph
OpenAI
RAG
Spark
LangChain
DevOps
Azure
Apply
$64k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Apache Kafka
AI/ML
Amazon SageMaker
Kubeflow
MLFlow
Spark
Vertex AI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
GitLab
GitLab CI
Jenkins
Kubernetes
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.