413,239open jobs
14,163companies
59,792added this week
Browse all
Location
In office (Sydney)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Why Firmus?

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap. 

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.

Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.

ROLE SUMMARY

The Senior AI Engineer (Inferencing) will build and improve the AI & Applications team’s inference capability, making models available as reliable, secure, scalable, and high-performance endpoints for internal products, external customers, and future Inference-as-a-service offerings.

The role will establish the engineering foundation for self-hosted model serving in the organization’s AI-factory environment. This includes model onboarding, deployment, endpoint provisioning, runtime selection, performance benchmarking and optimization, observability, capacity management, security, and operational lifecycle management. The objective is to provide users with predictable and efficient access to models while maintaining control over performance, cost, data handling, deployment configuration, and infrastructure utilization.

The role is a key contributor to the Model-to-Grid product and agentic applications roadmap. It will convert model and runtime characteristics into benchmarked, repeatable inference recipes and endpoint profiles that can inform workload scheduling, topology-aware placement, capacity planning, performance recommendations, and operational decision-making. It will also provide the governed and fit-for-purpose model endpoints needed by agentic systems for reasoning, retrieval, tool use, diagnosis, recommendation, and controlled automation.

KEY RESPONSIBILITIES

  • Build, operate, and continuously improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-service offerings.
  • Define and implement standard model-onboarding workflows covering model intake, compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management.
  • Provision and manage secure, scalable inference endpoints for common AI application patterns, including interactive generation, RAG, embeddings, reranking, batch processing, multimodal use cases, tool calling, and agentic workflows.
  • Develop reusable deployment templates, APIs, SDKs, configuration standards, and self-service workflows for users to request, configure, access, monitor, update, and retire model endpoints.
  • Work with leading inference frameworks and toolkits, such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, CUDA, cuDNN, NCCL, and related serving, profiling, and observability tools.
  • Optimize model-serving performance using appropriate techniques, including quantization, compilation, batching, continuous batching, request routing, KV-cache management, prefix caching, speculative decoding, load balancing, model routing, memory optimization, and distributed parallelism.
  • Build and validate reusable inference recipes that specify compatible model versions, framework and runtime versions, precision formats, GPU configurations, topology requirements, scaling approaches, scheduler profiles, benchmark results, and expected performance envelopes.
  • Use quantization and optimization approaches such as NVFP4, FP8, INT8, TensorRT compilation, kernel optimization, efficient attention mechanisms, and memory-management techniques while maintaining agreed model-quality targets.
  • Design distributed inference configurations for large models, including tensor, pipeline, expert, context, and data parallelism where appropriate.
  • Work with the Kubernetes and proprietary scheduler team to define endpoint resource profiles, placement requirements, topology preferences, priority classes, quota models, autoscaling rules, capacity reservations, and workload-management policies.
  • Contribute inference workload characteristics, benchmarks, and performance profiles to the Model-to-Grid product so that endpoint placement, scheduling, capacity planning, and AI-factory operations can make more informed decisions.
  • Build benchmarking and qualification workflows using controlled experiments, reproducible baselines, load tests, latency tests, throughput tests, concurrency tests, scaling tests, performance profiling, regression testing, and internal or industry-standard benchmark methodologies where relevant.
  • Measure and improve key inference indicators, including time-to-first-token, inter-token latency, tokens per second, requests per second, end-to-end latency, concurrency, GPU utilization, memory efficiency, cache hit rate, scaling efficiency, power efficiency, and cost efficiency.
  • Establish automated performance-regression testing and release qualification for model versions, runtime and toolkit upgrades, CUDA and driver changes, Kubernetes releases, scheduler changes, networking and storage changes, and new GPU platforms.
  • Build operational observability for inference services, including endpoint availability, request volume, latency, queueing, errors, GPU utilization, GPU memory use, cache behavior, capacity, cost, power, and service-level objectives.
  • Partner with the agentic applications team to provide fit-for-purpose self-hosted endpoints for agent planning, retrieval, tool use, summarization, diagnosis, recommendation, optimization, and AI-factory operations.
  • Expose governed inference, benchmark, recipe, performance, and capacity information to agentic systems, allowing them to recommend suitable models, identify degradation, diagnose bottlenecks, plan optimization experiments, and validate results.
  • Work with Product, UX, DevOps, Platform, Infrastructure, Security, and Global Operations teams to ensure that inference provisioning, model selection, endpoint configuration, performance visibility, quota management, and troubleshooting are clear, secure, and operationally supportable.

SKILLS AND EXPERIENCE

  • 5+ years of software engineering experience, including 3+ years in AI inference, model serving, ML systems, high-performance computing, distributed systems, or comparable performance-critical environments.
  • Demonstrated experience building, operating, or materially improving production model-serving platforms, inference APIs, GPU-backed services, AI developer platforms, or multi-tenant AI systems.
  • Hands-on experience with one or more modern inference frameworks, such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, Hugging Face Text Generation Inference, or equivalent technologies.
  • Strong understanding of the NVIDIA AI software stack, including CUDA, cuDNN, NCCL, TensorRT, GPU profiling, distributed communication, and GPU performance analysis.
  • Practical understanding of LLM and generative-AI serving behavior, including prompt processing, token generation, batching, context length, concurrency, KV-cache management, prefill and decode performance, request scheduling, model routing, and latency-throughput trade-offs.
  • Experience with model optimization methods, including quantization, compilation, calibration, mixed precision, kernel fusion, memory optimization, caching, speculative decoding, parallelism, and accuracy-performance validation.
  • Strong Python skills and working proficiency in C++ or Go for inference services, APIs, automation, benchmarking, profiling, runtime integrations, and performance-critical development.
  • Experience with distributed inference or training patterns, including tensor, pipeline, expert, context, and data parallelism; collective communication; fault handling; and multi-node scaling.
  • Familiarity with Kubernetes, containers, CI/CD, GitOps, service APIs, autoscaling, workload scheduling, observability, and production multi-tenant platform operations.
  • Understanding of high-performance GPU infrastructure, including GPU topology, NVLink, NVSwitch, PCIe, NUMA, NIC affinity, RDMA, RoCEv2, network fabrics, storage throughput, and their impact on inference performance.
  • Experience with inference benchmarking, performance profiling, reproducibility, load testing, regression testing, and analysis of throughput, latency, utilization, scaling, power, and cost metrics.
  • Familiarity with model-serving use cases such as RAG, embeddings, reranking, multimodal inference, agentic applications, model routing, and tool-calling workflows.
  • Understanding of security and governance for inference services, including identity, authentication, authorization, tenant isolation, quotas, rate limiting, secrets handling, audit logging, abuse prevention, and data protection.

KEY COMPETENCIES

  • Self-hosted inference platform engineering and production ownership.
  • Model serving, endpoint provisioning, lifecycle management, and developer self-service experience.
  • Inference-as-a-service foundations, including multi-tenancy, scalable endpoint operations, usage visibility, quotas, service profiles, and operational supportability.
  • High-performance LLM, generative-AI, embedding, reranking, and multimodal inference optimization.
  • Practical use of modern inference frameworks, model-serving toolkits, GPU profiling tools, and benchmarking methods.
  • Development of validated, repeatable, versioned inference recipes and deployment configurations.
  • Quantitative performance engineering across latency, throughput, GPU utilization, memory, scaling, power, energy, cost, and reliability.
  • Model-to-Grid thinking: connecting endpoint workload characteristics to benchmarking, scheduler policies, topology-aware placement, capacity, power, thermal state, and AI-factory operations.
  • Ability to provide reliable, governed, and cost-efficient inference services for agentic applications and autonomous operational workflows.
  • Cross-functional collaboration with AI applications, Model-to-Grid, scheduling, DevOps, Platform, SDI, Security, UX, product, and global operations teams.

SUCCESS METRICS

  • Reliable, secure, and scalable self-hosted inference endpoints are available for priority internal applications, external products, Model-to-Grid capabilities, and agentic systems.
  • Reduction in the time required to onboard, validate, optimize, deploy, provision, update, and retire a supported model endpoint.
  • Adoption of standardized inference deployment workflows supported model catalogues, endpoint templates, APIs, SDKs, recipes, and self-service capabilities.
  • Demonstrated foundations for future Inference-as-a-service offerings, including defined service profiles, endpoint lifecycle controls, tenant isolation, quota management, usage measurement, observability, support processes, and release governance.
  • Improvement in inference performance for priority workloads, measured by latency, time-to-first-token, inter-token latency, tokens per second, requests per second, concurrency, GPU utilization, memory efficiency, and scaling efficiency.
  • Reduction in cost per request, cost per token, energy per inference task, and avoidable resource over-provisioning, while maintaining agreed quality, availability, and reliability standards.
  • Number and adoption of benchmark-validated, documented, versioned inference recipes across supported models, frameworks, toolkits, precision formats, endpoint types, GPU configurations, and deployment topologies.
  • Effective use of benchmarking and performance-validation practices to prevent regressions across model updates, runtime upgrades, CUDA or driver changes, infrastructure changes, scheduler releases, and new GPU platforms.
  • Percentage of production endpoints with appropriate authentication, authorization, tenant isolation, quota controls, rate limits, metering, observability, audit logging, documentation, support runbooks, and rollback procedures.
  • Contribution of inference workload profiles and performance data to Model-to-Grid scheduling, capacity planning, placement quality, operational visibility, power efficiency, and user time-to-results.
  • Reliability and performance of endpoints supporting agentic applications, measured through agent task latency, response quality, tool-call completion, workflow success rate, safe automation outcomes, and reduced reliance on unmanaged external model services.

LOCATION

Singapore or Australia (Launceston, Hobart, Syndey, Melbourne)

Employment Basis

Permanent full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionise the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
413,239 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sydney
Remote/Hybrid • Full-Time • 2+ years exp
Python
Go
JavaScript
TypeScript
SQL
Apex
Apex
MuleSoft
Frontend
React.js
DevOps
CI/CD
Apply
In office • Internship • Singapore
Python
AI/ML
AI Agents
Apply
Remote/Hybrid • 5+ years exp
Python
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
pgvector
AI/ML
Claude
Claude Code
Embeddings
Prompt Engineering
Function Calling
LLM
RAG
Semantic Search
Anthropic
Structured Outputs
Semantic Search
Tool Use
Frontend
React.js
DevOps
Azure
Apply
In office • 5+ years exp
Python
AI/ML
Copilot
Claude
ChatGPT
Prompt Engineering
AI Agents
Midjourney
Instructor
TensorFlow
PyTorch
Perplexity
DevOps
GitHub
Apply
$38k – $44k per year • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree
Python
JavaScript
SQL
C#
Databases
Google BigQuery
BigQuery
Analytics
Microsoft Excel
Management
Google Workspace
Apply
$86k – $195k per year (Estimated) • In office • 5+ years exp • Sydney
Python
Go
AI/ML
CUDA Toolkit
Function Calling
AI Agents
LLM
RAG
CUDA
Human-in-the-Loop
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
CI/CD
Kubernetes
Vector
Cybersecurity
ISO 27001
SOC 2
OWASP ASVS
Least Privilege
Threat Modeling
Apply
$95k – $219k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sydney
Python
Go
Rust
Bash
Databases
ElasticSearch
AI/ML
InfiniBand
DevOps
Terraform
Ansible
Cilium
GitHub Actions
Loki
OpenTelemetry
etcd
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
Kubernetes
Grafana
Platform Engineering
Service Mesh
kubeadm
GitHub
GitLab
Cybersecurity
Open Policy Agent
Kyverno
OPA Gatekeeper
Calico
Apply
$99k – $217k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Snowflake
ClickHouse
Apache Iceberg
Delta Lake
Apache Kafka
Trino
Kafka
AI/ML
Spark
Airflow
dbt
AI Agents
LLM
Time Series Forecasting
DevOps
Helm
Kubernetes
Vector
Cybersecurity
ISO 27001
SOC 2
Apply
$40k – $96k per year (Estimated) • In office • 2+ years exp • Singapore
DevOps
Incident Management
Apply
$89k – $201k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Sydney
Python
Bash
DevOps
GCP
Azure
AWS
Kubernetes
Platform Engineering
Cybersecurity
Snyk
Trivy
HashiCorp Vault
Falco
ISO 27001
Checkov
kube-bench
CIS Benchmarks
OWASP Top 10
PCI DSS
SOC 2
HIPAA
Cosign
Kyverno
Calico
Auth0
Teleport
HashiCorp Boundary
Cryptography
Vault
Apply
$64k – $180k per year (Estimated) • In office • Full-Time • 3+ years exp • Sydney
Apply
$85k – $229k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Sydney
Apply
In office • Full-Time • Sydney
DevOps
Splunk
Azure
Cybersecurity
Microsoft Sentinel
ISO 27001
Apply
Sales Executive 10 hours ago
$68k – $137k per year (Estimated) • Remote • Full-Time • Bachelor's Degree • Sydney
Cybersecurity
Okta
Marketing
Salesforce
Apply
$111k – $284k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Sydney
Apply
See all jobs
This is one of many
413,239 more open roles from verified company boards, updated every day.