428,801open jobs
14,629companies
63,028added this week
Browse all
Salary
$78k – $192k per year (Estimated)
Location
In office (Melbourne)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

   

Why you’ll love working here  

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience,and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.  

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work.  

Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap. 

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.  

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in  

rather than drawing from them.  

Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience. 

ROLE SUMMARY  

Firmus Technologies is seeking a skilled Site Reliability Engineer, AI Infrastructure to join our Operations team, supporting the daily operations and maintenance of our AI-accelerated high-performance computing (HPC) infrastructure. This role will work closely with Field Service Engineers, HPC and Network Engineering teams, and assist the Global Operations Centre (GOC). This is a unique opportunity to contribute directly to the stability and growth of cutting-edge AI infrastructure. 

KEY RESPONSIBILITIES  

  • Support in the deployment, configuration, and maintenance of various high-end GPU servers, storage servers, networking equipment and software components in highly secure environments. 
  • Perform hardware diagnostics, systems functionality and firmware updates as required. 
  • Collaborate with engineering teams to assist in tailored customer environments deployment (eg: bare-metal systems, HPC Clusters, Kubernetes, Slurm etc). 
  • Serve as first line of engineering support for onsite operational issues, including troubleshooting hardware, network and software problems, and firmware compliance. 
  • Troubleshoot incidents, escalate critical issues and provide feedback to appropriate teams for improvements. 
  • Participate in an on-call rotation to ensure 24/7 availability and responsiveness to critical issues. 
  • Provide technical support to the GOC Support Specialist team in troubleshooting compute infrastructure related problems. 
  • Document incident details, resolutions, and lessons learned to enhance future problem-solving. 
  • Maintain clear, accurate, and up-to-date documentation to promote effective knowledge sharing across the team. 
  • Communicate effectively with GOC, HPC Engineers, internal teams, stakeholders, and end-users to ensure alignment on issue resolution. 
  • Take part in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning. 

SKILLS AND EXPERIENCE 

  • Bachelor’s degree in computer engineering, computer science, or a related technical field.  
  • 5+ years of experience in field service technical areas. 
  • Strong understanding of server hardware technology, firmware lifecycle, Linux environments and troubleshooting hardware problems, with adherence to physical and system-level security standards. 
  • Experience with scripting languages (eg: Bash, Python) 
  • Familiarity with using configuration management, CICD tools, workload manager and cluster softwares (eg: Slurm, Kubernetes, Nvidia BCM) and Observability tools (eg: Prometheus, Grafana, ELK, etc) 
  • Excellent problem-solving and analytical skills.  
  • Ability to work independently and as part of a team.  
  • Strong communication skills, both written and verbal. 

Location

This role is based in Brooklyn, Melbourne, Australia. 

Employment Basis

Full-time

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,801 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Melbourne
Lead Test Manager, VP 8 hours ago
$22k – $58k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Python
Java
SQL
Databases
Apache Kafka
DevOps
Rest API
GCP
Azure
CI/CD
AWS
QA
Selenium
Playwright
Appium
Postman
Rest-Assured
Apply
$65k – $156k per year (Estimated) • In office • Full-Time • PhD • Immenstaad am Bodensee
Python
C++
C++
Qt
AI/ML
Reinforcement Learning
Ray
DevOps
CI/CD
Kubernetes
GitHub
GitLab
Management
Confluence
Apply
QA Manager 8 hours ago
$14k – $37k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Pune • Mumbai
Python
SQL
PowerShell
Python
pySpark
Databases
Snowflake
Databricks
Azure SQL Database
AI/ML
Spark
Great Expectations
DevOps
GCP
Azure DevOps
Azure
CI/CD
AWS
Analytics
Power BI
ETL/ELT
Azure Data Factory
QA
TestRail
Apply
$46k – $72k per year • In office • Full-Time • Tokyo
JavaScript
TypeScript
Node JS
Databases
Redis
DevOps
Terraform
Vercel
Datadog
CI/CD
AWS
AWS Lambda
Amazon S3
Amazon ECS
Apply
$21k – $34k per year (net) • Remote • Freelance • 3+ years exp • Moscow
Python
Python
Django
Databases
PostgreSQL
Redis
DevOps
GitHub
Management
Telegram
Apply
$142k – $269k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco
Python
Go
Rust
Bash
Databases
ElasticSearch
AI/ML
InfiniBand
DevOps
Terraform
Ansible
Cilium
GitHub Actions
Loki
OpenTelemetry
etcd
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
Kubernetes
Grafana
Platform Engineering
Service Mesh
kubeadm
GitHub
GitLab
Cybersecurity
Open Policy Agent
Kyverno
OPA Gatekeeper
Calico
Apply
$77k – $179k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Apply
$81k – $215k per year (Estimated) • In office • 10+ years exp • Singapore
AI/ML
LLM
DevOps
HPC
Apply
$86k – $214k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Python
SQL
Python
FastAPI
Celery
Databases
PostgreSQL
Weaviate
Neo4j
Milvus
pgvector
Pinecone
Qdrant
ElasticSearch
OpenSearch
AI/ML
LangGraph
AutoGen
LangChain
DSPy
LlamaIndex
Model Context Protocol
Dagster
Prefect
Embeddings
Multimodal AI
Function Calling
AI Agents
Arize Phoenix
DeepEval
Haystack
Langfuse
LangSmith
Promptfoo
Pydantic AI
Ragas
Semantic Kernel
CrewAI
LLM
RAG
TruLens
Reranking
Hybrid Search
Time Series Forecasting
OpenAI
Human-in-the-Loop
Structured Outputs
Knowledge Graph
LLM Guardrails
Multi-Agent Systems
Tool Use
DevOps
gRPC
OpenTelemetry
WebSockets
CI/CD
Kubernetes
Argo Workflows
Vector
Cybersecurity
Defense in Depth
Apply
$89k – $222k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Python
Go
AI/ML
Fine-tuning
AI Agents
NVLink
DevOps
gRPC
Terraform
Helm
GitHub Actions
OpenTelemetry
Kustomize
Prometheus
GitLab CI
SLURM
CI/CD
GitOps
ArgoCD
Kubernetes
Grafana
SRE
Platform Engineering
GitHub
GitLab
HPC
Cybersecurity
Kyverno
Apply
$55k – $127k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Auckland • Melbourne
Apply
$74k – $160k per year (Estimated) • In office • Full-Time • Melbourne
Apply
Test Delivery Lead 2 days ago
$71k – $185k per year (Estimated) • In office • Full-Time • Melbourne
DevOps
Azure DevOps
Azure
Management
Jira
Apply
Remote/Hybrid • Full-Time • Melbourne
Apply
Lead Civil Engineer 2 days ago
$82k – $194k per year (Estimated) • Remote/Hybrid • Full-Time • Melbourne
Design
AutoCAD
Apply
See all jobs
This is one of many
428,801 more open roles from verified company boards, updated every day.