368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$90k – $227k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Role Summary

The AI Engineer will establish Firmus AI Factory as the foundation for efficient, production-grade distributed training by delivering pre-built training recipes (TorchTitan, Megatron etc.), evaluation benchmarks, and model guidance. You'll work with customers and internal teams to optimize training efficiency, define baselines, and document best practices. Your templates and benchmarks are the anchor point for our hyperscale customers' training workflows and our model arena differentiator.

Key Responsibilities

  • Build production-ready training recipes using TorchTitan and Megatron-LM: model configs, parallelism strategies (FSDP, tensor/pipeline parallelism), checkpointing patterns.
  • Document parameter tuning for different scales (e.g., "to train Llama 7B on 8xH100s, use this config and expect X throughput").
  • Create and validate multi-node NCCL communication patterns on AI Factory K8s/Slurm clusters.
  • Design and build benchmarking suites: accuracy, latency, throughput (tokens/sec), cost per token, energy efficiency, MFU.
  • Implement offline evaluation harnesses for standardized model comparison and leaderboard tracking.
  • Conduct fine-tuning experiments (LoRA, QLoRA) where they improve product outcomes (e.g., ops domain data), document gains.
  • Create training efficiency playbooks and publish benchmark results so customers can optimize workloads.
  • Partner with job scheduling and orchestration engineers on template integration and other AI engineers and software engineers on model optimization trade-offs for inferencing and AI applications.

Skills & Experience

  • 5-7 years of experience in distributed machine learning (PyTorch/JAX, FSDP, DeepSpeed, multi-node training at 10+ GPUs).
  • Expert-level understanding of GPU optimization: utilization, memory patterns, communication bottlenecks (NCCL collectives).
  • Hands-on distributed training at scale: debugged convergence issues, profiled bottlenecks, optimized throughput.
  • Strong benchmarking methodology: design-controlled experiments, measure noise, communicate results rigorously.
  • Familiarity with TorchTitan, Megatron-LM, or similar production training frameworks. 
  • Understanding of model parallelism strategies and trade-offs (FSDP vs. tensor parallelism vs. pipeline parallelism etc.).

Key Competencies

  • Distributed Systems Mastery: can explain NCCL, collective communications, and scaling inefficiency.
  • Benchmarking Rigor: doesn't just run benchmarks; validates assumptions, explains variance, communicates uncertainty.
  • Production Thinking: understands checkpointing, recovery, resource constraints, and cost optimization.
  • Mentorship: can guide engineers on training best practices and debugging distributed training issues.
  • Documentation: creates clear, actionable playbooks that customers can follow.

Success Metrics

  • Benchmark credibility & decision impact increases: benchmarks are trusted and used to drive model/hardware/product decisions.
  • Training efficiency leadership: sustained improvement in benchmarked training efficiency on representative workloads.
  • Shorter time-to-validate new models: model candidates can be evaluated quickly and consistently end-to-end.
  • Template effectiveness improves: recipes reduce misconfigurations and repeated setup failures; fewer training config escalations.
  • Competitive differentiation strengthens: model arena outputs influence customer adoption and internal roadmap priorities.

Location & Reporting

  • Singapore or Australia (Launceston, TAS or Sydney, NSW)
  • Reporting to Head of AI & Applications

Employment Basis

Full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions. 

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure. 

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$217k – $304k per year • Equity • Remote • Full-Time • 8+ years exp
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$60k – $148k per year (Estimated) • In office • Full-Time • Launceston
DevOps
HPC
IoT
OPC UA
Apply
$83k – $193k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Sydney
Apply
$149k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LLM
Management
Jira
Apply
$181k – $331k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
PyTorch
TensorFlow
InfiniBand
Apply
$166k – $356k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
Fine-tuning
Apply
$117k – $251k per year (Estimated) • Remote/Hybrid • Full-Time • Singapore
Apply
$74k – $126k per year (Estimated) • In office • Full-Time • Singapore
Python
Apply
$88k – $191k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Singapore
C++
Java
Kotlin
Python
Mobile
JUnit
DevOps
Git
gRPC
Jenkins
JFrog Artifactory
Shift-Left
Cybersecurity
Shift-Left Security
QA
Pytest
Robot Framework
TestNG
Apply
Senior AI Architect 5 hours ago
$138k – $304k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
LangGraph
OpenAI
RAG
Spark
LangChain
DevOps
Azure
Apply
$64k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Apache Kafka
AI/ML
Amazon SageMaker
Kubeflow
MLFlow
Spark
Vertex AI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
GitLab
GitLab CI
Jenkins
Kubernetes
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.