368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$79k – $198k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies  

Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-efficient AI infrastructure, combining cutting edge technology with a steadfast commitment to sustainability.  

Through ground-breaking research and development, we invented a verticalized AI Factory - a new class of digital infrastructure that replaces traditional data centres. Built on new approaches to liquid cooling, energy management, water use and modular construction methodology, the Firmus AI Factory delivers low-cost AI tokens across Asia-Pacific. 

Firmus AI Cloud  

We provide customers with access to energy savings via our large-scale GPU cloud, Firmus AI Cloud. Rated Silver in The GPU Cloud ClusterMAX™ Rating System, our cloud empowers developers, enterprise, education and government users to train AI models with unmatched efficiency and cost savings. With an ever-growing list of services and applications, we are committed to building a cloud experience for our customers that is market-leading, proprietary and built to scale.  

Why you’ll love working here  

  • A fast-paced and dynamic environment working with next-gen technology. You’ll be operating at the intersection of sustainability and artificial intelligence - helping to transform an industry.  
  • Working with and access to colleagues who are true innovators and leaders in their field.  
  • As an emerging company, we work as a close-knit team. Work with the founders, grow a strong network, and witness the impact you make first-hand as we democratise AI tools for everyone - more sustainably, and more affordably.  
  • We believe that people from diverse backgrounds come together to do their best work, be their authentic selves, and build great things. We are proud to be an equal opportunity employer.  

ROLE SUMMARY 

Firmus Technologies is seeking a skilled Site Reliability Engineer to join our Operations team, supporting the daily operations and maintenance of our AI-accelerated High-Performance Computing (HPC) infrastructure. This role will work closely with Field Service Engineers, HPC and Network Engineering teams, and assist the Global Operations Centre (GOC). This is a unique opportunity to contribute directly to the stability and growth of cutting-edge AI infrastructure.

 

KEY RESPONSIBILITIES 

  • Support in the deployment, configuration, and maintenance of various high-end GPU servers, storage servers, networking equipment and software components in highly secure environments. 
  • Perform hardware diagnostics, systems functionality and firmware updates as required. 
  • Collaborate with engineering teams to assist in tailored customer environments deployment (eg: bare-metal systems, HPC Clusters, Kubernetes, Slurm etc). 
  • Serve as first line of engineering support for onsite operational issues, including troubleshooting hardware, network and software problems, and firmware compliance. 
  • Troubleshoot incidents, escalate critical issues and provide feedback to appropriate teams for improvements. 
  • Participate in an on-call rotation to ensure 24/7 availability and responsiveness to critical issues. 
  • Provide technical support to the GOC Support Specialist team in troubleshooting compute infrastructure related problems. 
  • Document incident details, resolutions, and lessons learned to enhance future problem-solving. 
  • Maintain clear, accurate, and up-to-date documentation to promote effective knowledge sharing across the team. 
  • Communicate effectively with GOC, HPC Engineers, internal teams, stakeholders, and end-users to ensure alignment on issue resolution. 
  • Take part in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning. 

SKILLS AND EXPERIENCE  

  • Bachelor’s degree in computer engineering, computer science, or a related technical field.  
  • 5+ years of experience in field service technical areas. 
  • Strong understanding of server hardware technology, firmware lifecycle, Linux environments and troubleshooting hardware problems, with adherence to physical and system-level security standards. 
  • Experience with scripting languages (eg: Bash, Python) 
  • Familiarity with using configuration management, CICD tools, workload manager and cluster softwares (eg: Slurm, Kubernetes, Nvidia BCM) and Observability tools (eg: Prometheus, Grafana, ELK, etc) 
  • Excellent problem-solving and analytical skills.  
  • Ability to work independently and as part of a team.  
  • Strong communication skills, both written and verbal. 

  

Location & Reporting  

  • Based in: Singapore  
  • Reporting to: Senior Operations Manager 

Employment Basis  

Full-time 

Diversity  

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.  

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.  

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
Bash
C++
Java
Python
DevOps
CI/CD
Git
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
C++
Python
DevOps
CI/CD
SpaceTech
NASA cFS
Apply
$60k – $148k per year (Estimated) • In office • Full-Time • Launceston
DevOps
HPC
IoT
OPC UA
Apply
$83k – $193k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Sydney
Apply
$149k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LLM
Management
Jira
Apply
$181k – $331k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
PyTorch
TensorFlow
InfiniBand
Apply
$166k – $356k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
Fine-tuning
Apply
$117k – $251k per year (Estimated) • Remote/Hybrid • Full-Time • Singapore
Apply
$74k – $126k per year (Estimated) • In office • Full-Time • Singapore
Python
Apply
$88k – $191k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Singapore
C++
Java
Kotlin
Python
Mobile
JUnit
DevOps
Git
gRPC
Jenkins
JFrog Artifactory
Shift-Left
Cybersecurity
Shift-Left Security
QA
Pytest
Robot Framework
TestNG
Apply
Senior AI Architect 4 hours ago
$138k – $304k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
LangGraph
OpenAI
RAG
Spark
LangChain
DevOps
Azure
Apply
$64k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Apache Kafka
AI/ML
Amazon SageMaker
Kubeflow
MLFlow
Spark
Vertex AI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
GitLab
GitLab CI
Jenkins
Kubernetes
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.