428,413open jobs
14,457companies
63,723added this week
Browse all
Salary
$86k – $214k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Why Firmus?

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap. 

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.

Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.

Role Summary

The Senior AI Engineer (Agents & Applications) will design, build, and operate production-grade agentic systems that coordinate, optimize, and automate decision-making across the  design-build-operate lifecycle of AI factories. The role is a core contributor to the AI & Applications team’s Model-to-Grid product, connecting models, inference endpoints, benchmark intelligence, validated workload recipes, job-scheduler decisions, infrastructure telemetry, AI-factory operations, and grid-related constraints into safe, explainable, and measurable workflows.

The role will build more than conversational co-pilots. It will create agentic applications that ingest and reason over time-series telemetry, logs, traces, events, scheduler state, benchmark results, configuration data, operational documentation, incident records, and multimodal sources where appropriate. These applications will help engineers, operators, and customers move from observation to diagnosis, recommendation, planning, simulation, controlled execution, verification, and continuous improvement.

The engineer will define and implement the underlying agent architecture and engineering framework: orchestration, state and memory management, retrieval, tool use, specialized sub-agents, evaluation, safety controls, human approvals, observability, and deployment. The role will use fit-for-purpose self-hosted and external model endpoints, with close integration to the team’s inference platform.

Key Responsibilities

  • Design, build, and operate agentic applications supporting AI-factory planning, commissioning, validation, workload onboarding, benchmark analysis, model and recipe optimisation, scheduling, operations, maintenance, incident response, and continuous improvement.
  • Define reference architectures for single-agent, multi-agent, workflow-based, eventdriven, and human-in-the-loop agentic systems.
  • Build orchestration workflows using appropriate agent frameworks and libraries, such as LangGraph, LangChain, LlamaIndex, Microsoft AutoGen, Semantic Kernel, CrewAI, PydanticAI, Haystack, DSPy, or equivalent custom-built frameworks.
  • Select the appropriate architecture for each use case rather than applying multi-agent 

    patterns by default:

    • Deterministic workflow and state-machine architectures for repeatable, high-confidence operational processes.
    • Planner-executor architectures for decomposing complex investigation, planning, and remediation tasks.
    • Supervisor-worker or manager-worker architectures for coordinating specialist domain agents.
    • Router architectures for selecting the right model, tool, knowledge source, workflow, or specialist agent.
    • Reflection, critic, verifier, or judge patterns for quality assurance, validation, and safety checks.
    • Event-driven architectures for responding to telemetry anomalies, workload failures, scheduler events, benchmark regressions, and operational alerts.
    • Human-in-the-loop architectures for high-impact recommendations, privileged actions, or changes to production environments.
  • Build specialist agents for relevant Model-to-Grid and AI-factory domains, such as:
    • Benchmark-analysis and performance-diagnosis agents.
    • Workload recipe and runtime-configuration recommendation agents.
    • Inference-endpoint selection, capacity, and optimization agents.
    • Kubernetes and job-scheduler diagnostic agents.
    • GPU-topology, network, RDMA, storage, and utilization-analysis agents.
    • AI-factory health, capacity, maintenance, and operational-triage agents.
    • Documentation, knowledge, incident-review, and runbook-execution assistants.
    • Thermal domain specific monitoring and optimization agents.
    • Power domain specific monitoring and optimization agents.
    • Grid-integration specific monitoring and optimization agents.
  • Develop the intelligent coordination layer for Model-to-Grid, enabling agents to reason across model characteristics, inference and training configuration, validated recipes, GPU resources, topology, scheduling policies, network and storage performance, capacity, power, thermal conditions, health signals, and operational constraints.
  • Build Retrieval-Augmented Generation (RAG) pipelines using a combination of vector retrieval, hybrid search, metadata filtering, reranking, structured-data queries, graphbased retrieval where valuable, source attribution, and permission-aware access controls.
  • Use tools such as pgvector, OpenSearch, Elasticsearch, Milvus, Weaviate, Pinecone, Qdrant, Neo4j, or equivalent data and retrieval platforms as appropriate to the product architecture and deployment environment.
  • Design knowledge-ingestion pipelines for documentation, runbooks, ticketing systems, configuration repositories, benchmark reports, experiment records, cluster state, telemetry catalogues, incident reports, and approved internal knowledge sources.
  • Build data and context pipelines that combine unstructured knowledge with structured operational data, including metrics, logs, traces, events, time-series databases, scheduler queues, job states, resource inventories, and configuration-management data.
  • Integrate agents with governed tools and APIs, including Kubernetes, proprietary scheduler services, observability platforms, benchmark services, inference endpoints, configuration repositories, CI/CD pipelines, ticketing systems, workflow engines, 

    databases, and operational tooling.

  • Define tool contracts using structured input and output schemas, typed interfaces, validation, retries, idempotency controls, rate limits, timeouts, circuit breakers, approval requirements, and detailed audit logging.
  • Build robust agent harnesses that provide context assembly, model routing, prompt and policy versioning, structured output handling, memory management, state persistence, retries, failure handling, task recovery, escalation, and end-to-end tracing.
  • Implement short-term task memory, long-term user or operational memory where permitted, episodic memory for prior investigations or incidents, and semantic memory based on approved knowledge stores; apply retention, access-control, and datagovernance requirements to each.
  • Implement model-routing and fallback strategies across self-hosted inference endpoints and approved external models, selecting models according to task complexity, latency, cost, context-window requirement, tool-use capability, privacy needs, and reliability targets.
  • Partner with the Self-Hosted Inference Platform & Optimization team to ensure that agentic applications have suitable endpoint profiles for planning, reasoning, embeddings, reranking, summarization, classification, tool use, multimodal analysis, and high-throughput operational workflows.
  • Build agent workflows for detection, diagnosis, recommendation, planning, action simulation, controlled execution, verification, and learning loops.
  • Develop offline replay, simulation, shadow-mode, and what-if evaluation capabilities to validate recommendations before allowing actions in production-especially for scheduler policies, workload placement, runtime changes, capacity decisions, and operational remediation.
  • Design human approval and policy enforcement workflows that clearly present an agent’s evidence, recommendation, expected impact, proposed action, confidence, risk classification, authorization scope, and rollback option.
  • Work with the Security Engineer to implement defense-in-depth controls against direct and indirect prompt injection, insecure output handling, excessive agency, unsafe tool use, data leakage, cross-tenant exposure, privilege escalation, credential misuse, unauthorized actions, and insufficient auditability.
  • Use policy engines, guardrail frameworks, structured output validation, content and tool filters, permission checks, sandboxing, and allowlisted action patterns to ensure that agent behavior remains bounded and trustworthy.
  • Implement agent observability using tracing, metrics, logs, prompt and model version tracking, tool-call records, evaluation results, token and cost tracking, user feedback, incident evidence, and action audit trails.
  • Build and maintain evaluation frameworks using automated tests, curated test sets, simulation, replay, benchmark tasks, regression suites, model-based evaluators, human review, and operational-outcome measures.
  • Evaluate agents on task completion, factuality, groundedness, retrieval quality, diagnostic accuracy, recommendation quality, tool-selection accuracy, tool-execution correctness, policy compliance, latency, cost, safety, and user or operator satisfaction.
  • Partner with the Kubernetes and custom scheduler team to consume and explain queue state, placement rationale, topology information, capacity signals, workload lifecycle events, policy outcomes, and performance data.
  • Partner with the Model-to-Grid product, inference, Platform, SDI, Security, UX, and global operations teams to turn agent capabilities into clear product workflows, production releases, runbooks, and measurable user and operational outcomes.

Skills & Experience

  • 5+ years of software engineering experience, including 3+ years building AI/ML applications, distributed systems, automation platforms, data products, or production workflow systems.
  • Demonstrated experience delivering LLM-powered, agentic, retrieval-augmented, decision-support, or operational-automation applications into production.
  • Strong Python expertise, including experience with FastAPI or comparable API frameworks, asynchronous programming, event-driven services, distributed task execution, data pipelines, and API integrations.
  • Hands-on experience with one or more agent frameworks, such as LangGraph, LangChain, LlamaIndex, Microsoft AutoGen, Semantic Kernel, CrewAI, PydanticAI, Haystack, DSPy, or an equivalent in-house agent framework.
  • Demonstrated ability to build custom agent orchestration when frameworks are insufficient, including state machines, graph-based execution, durable workflow execution, task queues, tool routers, model routers, planners, evaluators, and humanapproval flows.
  • Experience with workflow and orchestration technologies such as Temporal, Dagster, Prefect, Airflow, Argo Workflows, Kubernetes Jobs, Celery, or equivalent event-driven or durable-execution platforms.
  • Strong understanding of multi-agent architectures, including supervisor-worker, planner-executor, router, reflection, critic-verifier, debate, hierarchical, blackboard, and event-driven coordination patterns.
  • Practical experience with RAG architectures, including chunking, embeddings, vector search, hybrid retrieval, metadata filters, reranking, structured-data retrieval, SQL generation controls, knowledge graphs, provenance, and evaluation.
  • Experience with vector databases, search platforms, or graph databases such as pgvector, OpenSearch, Elasticsearch, Milvus, Weaviate, Pinecone, Qdrant, Neo4j, or equivalent technologies.
  • Familiarity with self-hosted and managed LLM inference, including model routing, OpenAI-compatible APIs, embeddings, reranking, tool calling, structured outputs, streaming, rate limits, latency, capacity, and cost management.
  • Understanding of model protocols and integration approaches such as REST, gRPC, WebSockets, event streams, OpenAPI, JSON Schema, and Model Context Protocol (MCP) or comparable tool-integration patterns.
  • Experience integrating AI systems with time-series data, logs, traces, observability platforms, ticketing systems, configuration repositories, CI/CD systems, databases, cloud services, and enterprise APIs.
  • Familiarity with cloud-native AI platforms, including Kubernetes, containers, workload scheduling, model serving, GPU resources, observability, multi-tenancy, and operational runbooks.
  • Understanding of agent security and responsible-AI controls, including prompt injection, indirect prompt injection, data and tenant isolation, tool authorisation, workload identity, permission boundaries, auditability, output validation, sandboxing, and humanin-the-loop safeguards.
  • Experience with agent evaluation and observability tools or patterns, such as OpenTelemetry, OpenInference, LangSmith, Langfuse, Arize Phoenix, Weave, TruLens, Ragas, DeepEval, promptfoo, custom test harnesses, or equivalent tooling.

Location & Reporting

  • Location:Singapore
  • Reports to:Head of AI & Applications

Employment Basis

Permanent full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionise the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,413 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$118k – $163k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
Python
JavaScript
TypeScript
Node JS
Python
Django
Databases
PostgreSQL
Frontend
Vue.js
Angular
React.js
DevOps
Terraform
AWS
Docker
Kubernetes
QA
Swagger
Apply
$14k – $33k per year (Estimated) • Remote/Hybrid • 4+ years exp • Almaty
Java
Kotlin
SQL
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
Claude
ChatGPT
DevOps
GitLab CI
CI/CD
Jenkins
Git
Docker
Shift-Left
GitLab
Cybersecurity
Shift-Left Security
QA
Selenium
Playwright
Rest-Assured
Apply
Remote • 3+ years exp • Tbilisi
Go
JavaScript
TypeScript
Databases
PostgreSQL
Redis
RabbitMQ
Apache Kafka
Kafka
Frontend
Vue.js
React.js
DevOps
Git
AWS
Docker
Kubernetes
Apply
$88k – $202k per year (Estimated) • In office • Full-Time • 3+ years exp • Singapore
SQL
Databases
Databricks
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Platform Engineering
GitHub
IAM
Cybersecurity
GDPR
Analytics
ETL/ELT
Apply
$102k – $234k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Singapore
SQL
Databases
Snowflake
Databricks
Delta Lake
Google BigQuery
BigQuery
AI/ML
LLM
Tokenization
DevOps
GCP
Azure
AWS
IAM
Cybersecurity
PCI DSS
GDPR
HIPAA
Management
SharePoint
Apply
$81k – $215k per year (Estimated) • In office • 10+ years exp • Singapore
AI/ML
LLM
DevOps
HPC
Apply
$89k – $222k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Python
Go
AI/ML
Fine-tuning
AI Agents
NVLink
DevOps
gRPC
Terraform
Helm
GitHub Actions
OpenTelemetry
Kustomize
Prometheus
GitLab CI
SLURM
CI/CD
GitOps
ArgoCD
Kubernetes
Grafana
SRE
Platform Engineering
GitHub
GitLab
HPC
Cybersecurity
Kyverno
Apply
$93k – $233k per year (Estimated) • In office • Full-Time • 5+ years exp • Sydney
Python
Go
C++
AI/ML
vLLM
CUDA Toolkit
Triton Inference Server
Embeddings
Quantization
Multimodal AI
Function Calling
AI Agents
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
RAG
Reranking
NVIDIA NIM
CUDA
Triton
Hugging Face
NCCL
NVLink
cuDNN
Speculative Decoding
KV Cache
Agentic Workflows
Tool Use
DevOps
CI/CD
GitOps
Kubernetes
Platform Engineering
Apply
$93k – $234k per year (Estimated) • In office • 5+ years exp • Sydney
Python
Go
AI/ML
CUDA Toolkit
Function Calling
AI Agents
LLM
RAG
CUDA
Human-in-the-Loop
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
CI/CD
Kubernetes
Vector
Cybersecurity
ISO 27001
SOC 2
OWASP ASVS
Least Privilege
Threat Modeling
Apply
$78k – $192k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sydney
Python
Go
Rust
Bash
Databases
ElasticSearch
AI/ML
InfiniBand
DevOps
Terraform
Ansible
Cilium
GitHub Actions
Loki
OpenTelemetry
etcd
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
Kubernetes
Grafana
Platform Engineering
Service Mesh
kubeadm
GitHub
GitLab
Cybersecurity
Open Policy Agent
Kyverno
OPA Gatekeeper
Calico
Apply
$126k – $311k per year (Estimated) • In office • 10+ years exp • Singapore
Management
Stripe
Marketing
LinkedIn
Apply
In office • Singapore
Management
Outlook
Apply
$33k – $67k per year (Estimated) • In office • Full-Time • Singapore
Apply
In office • Full-Time • Singapore • Hong Kong
JavaScript
Frontend
Parcel
Apply
$84k – $197k per year (Estimated) • In office • Full-Time • 2+ years exp • Singapore
DevOps
SLI/SLO/SLA
Apply
See all jobs
This is one of many
428,413 more open roles from verified company boards, updated every day.