368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$86k – $218k per year (Estimated)
Location
In office (Sydney)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

Role Summary

As a Senior Software Engineer on the AI and Applications team, you'll own the control plane that powers AI workload submission across Firmus AI Platforms. You'll design and build unified job submission APIs, CLI, and web interfaces for training, inference, and fine-tuning workloads on Kubernetes and Slurm-implementing RBAC, multi-tenant isolation, resource quotas, and intelligent scheduling policies (priority classes, pre-emption, fairness). You'll create template catalog for pre-built training and inference recipes, wire observability pipelines for per-job GPU metrics cost tracking and expose telemetry APIs for platform monitoring. This role requires deep Kubernetes and Slurm expertise, strong distributed systems knowledge, and close collaboration with infra, platform, and LLM engineering teams to deliver a seamless, production-grade job orchestration experience for hyperscaler customers.

Key Responsibilities

  • Design and build unified job submission APIs, CLI, and web UI for all AI workload types (training, inference, fine-tuning) on Kubernetes and Slurm with Firmus AI Factory context (tenant isolation, resource requests, metadata tagging, observability hooks).
  • Implement comprehensive job metadata models and schemas: track job ID, job type, tenant, user, resource requirements, priority class, timestamps, lineage, execution status.
  • Integrate authentication/authorization (RBAC) and resource quotas; enforce multi-tenant isolation at submission time across all job types.
  • Build AI job scheduling and orchestration layer: priority classes, preemption policies, fairness algorithms, resource quota enforcement, and intelligent job routing.
  • Build the AI Factory template catalog: discovery, parameter validation, and manifest generation for training templates, inference serving templates, and fine-tuning recipes.
  • Wire job submissions to observability pipeline: inject labels/annotations (job_id, tenant, user, model_name, job_type) so metrics are tagged per-job.
  • Expose job-level telemetry APIs (GPU metrics, cost accrual, MFU progression for training; latency, throughput, tokenomics for inferencing) for platform telemetry and monitoring.
  • Extend job submission to handle inference workloads: design inference job specifications (model, batch size, latency SLA, cost constraints); integrate with inference serving APIs.
  • Coordinate with platform team on observability dashboard integration, with LLM engineers on template design, and with ModelOps on reliability standards. 

Skills & Experience

  • 5-7 years of backend engineering experience building production APIs and distributed systems (Python, Go, or Java).
  • Deep Kubernetes expertise: understand Job controllers, Pod specs, resource requests/limits, RBAC, network policies, debugging.
  • Hands-on Slurm experience: job submission, resource allocation, job queues, sbatch scripting.
  • Strong distributed systems knowledge: understand scheduling algorithms, fairness, preemption, resource management.
  • Strong data modelling: can design clear schemas for job metadata, handle versioning and migrations, ensure backward compatibility.
  • DevOps mindset: comfortable with observability, logging, tracing, and production troubleshooting.
  • Experience with streaming APIs and real-time webhooks, and system-level integration patterns. 

Key Competencies

  • Job Orchestration & Scheduling: shipped job scheduling or workflow systems at scale; understands job lifecycle, failure modes, and scheduling policies.
  • Multi-Tenancy Design: can architect fair resource allocation, quota enforcement, pre-emption, and data isolation across job types.
  • API Design: RESTful or gRPC APIs that are intuitive and extensible; handles versioning gracefully.
  • Systems Architecture: understands how job submission connects to training, inferencing, observability, cost tracking, and incident response.
  • Cross-Domain Partnership: works closely with infra team, platform team, LLM engineers; clear handoff points and API contracts. 

Success Metrics

  • Unified orchestration adoption increases: teams use the standard job interface rather than bespoke/manual pathways.
  • Scheduling effectiveness & fairness improves: predictable scheduling under contention with reduced noisy-neighbor impact.
  • Orchestration reliability stays high: jobs reliably start, run, and complete across K8s/Slurm/inference integrations.
  • End-to-end workflow automation increases: higher share of workflows complete without human intervention (e.g., train→register→serve).
  • Interface stability & compatibility remains strong: the orchestration API evolves without breaking users. 

Location & Reporting

  • This role can be based in either Sydney, Australia, or Singapore.
  • Reporting to Head of AI & Applications

Employment Basis

Full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions. 

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure. 

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sydney
$20k – $50k per year (Estimated) • In office • Full-Time • Tomsk
C++
Go
Java
Kotlin
Databases
Apache Kafka
ClickHouse
ElasticSearch
DevOps
Jaeger
Kubernetes
Prometheus
SLI/SLO/SLA
Apply
$20k – $50k per year (Estimated) • In office • Full-Time • Perm
C++
Go
Java
Kotlin
Databases
Apache Kafka
ClickHouse
ElasticSearch
DevOps
Jaeger
Kubernetes
Prometheus
SLI/SLO/SLA
Apply
$11k – $25k per year (Estimated) • In office • Full-Time • 2+ years exp • Perm
Go
Java
Kotlin
SQL
Databases
Apache Kafka
PostgreSQL
DevOps
GitLab
QA
Postman
Swagger
Apply
$11k – $25k per year (Estimated) • In office • Full-Time • 2+ years exp • Tomsk
Go
Java
Kotlin
SQL
Databases
Apache Kafka
PostgreSQL
DevOps
GitLab
QA
Postman
Swagger
Apply
$17k – $43k per year (Estimated) • In office • Full-Time • Tomsk
C#
Java
Python
DevOps
CI/CD
Docker
Git
Grafana
Graylog
Kubernetes
Prometheus
Splunk
Apply
$60k – $148k per year (Estimated) • In office • Full-Time • Launceston
DevOps
HPC
IoT
OPC UA
Apply
$83k – $193k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Sydney
Apply
$149k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LLM
Management
Jira
Apply
$181k – $331k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
PyTorch
TensorFlow
InfiniBand
Apply
$166k – $356k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
Fine-tuning
Apply
$67k – $173k per year (Estimated) • In office • Full-Time • 3+ years exp • Sydney
AI/ML
AI Agents
Apply
$95k – $224k per year (Estimated) • In office • Full-Time • Sydney
PowerShell
Python
DevOps
AIOps
Amazon ECS
Ansible
AWS
Grafana
Prometheus
Terraform
VMWare
Cybersecurity
CyberArk
Apply
$96k – $227k per year (Estimated) • In office • Full-Time • Melbourne • Sydney
PowerShell
Python
DevOps
AWS
Azure
CI/CD
GCP
Incident Management
Kubernetes
Platform Engineering
Service Mesh
Terraform
Apply
In office • Full-Time • Sydney
Python
SQL
Databases
Databricks
Snowflake
AI/ML
AI Agents
Copilot
DevOps
AWS
Azure
GCP
GitHub
Apply
$108k – $240k per year (Estimated) • In office • Full-Time • Sydney
C#
JavaScript
Frontend
Next.js
React.js
DevOps
AWS
CI/CD
Rest API
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.