428,639open jobs
14,645companies
62,546added this week
Browse all
Salary
$142k – $269k per year (Estimated)
Location
In office (San Francisco)
Seniority
Senior · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.

Firmus Technologies

Firmus Technologies is a global leaderpioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operatea new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

Role Summary

The Senior Kubernetes Engineer, AI Infrastructure owns the technical design and delivery of the backend infrastructure that powers the Firmus Kubernetes platform. This is a hands-on principal-level individual contributor role, responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities across GPU-accelerated bare-metal environments.

They solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign-off for platform designs. They work across AI Platforms, Solutions Architecture & Delivery, networking, security, and operations to create a secure, resilient, multi-tenant platform that can be deployed and operated consistently at AI-factory scale.

Key Responsibilities

  • Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design.
  • Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably.
  • Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3.
  • Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR-IOV, BGP, InfiniBand, or RoCE where required for high-performance AI workloads.
  • Define persistent-storage and data-service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster-recovery mechanisms appropriate for stateful platform and AI workloads.
  • Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement for multi-node accelerated workloads.
  • Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices.
  • Build platform security into the architecture through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management.
  • Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.
  • Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross-team technical decisions while remaining directly involved in implementation.

Skills & Experience

  • 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level.
  • Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control-plane failure modes.
  • Demonstrated experience designing, building, and operating highly available, large scale and multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
  • Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
  • Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behaviour, performance analysis, and low-level troubleshooting.
  • Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.
  • Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
  • Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
  • Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance.
  • Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements.
  • Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.
  • CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.
  • Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
  • Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.

Location & Reporting

  • San Francisco Bay Area
  • Reporting to Head of AI Platform

Employment Basis

Full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,639 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Lead Golang Engineer 10 hours ago
Remote/Hybrid • 6+ years exp
Go
Java
Rust
SQL
C++
Databases
RabbitMQ
Apache Kafka
DevOps
gRPC
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
Remote/Hybrid • Bachelor's Degree
PowerShell
Bash
Apply
Remote/Hybrid • 5+ years exp • Bachelor's Degree
Python
DevOps
Azure DevOps
Azure
CI/CD
Docker
Kubernetes
Incident Management
Apply
Network Architect 7 hours ago
Remote/Hybrid • 10+ years exp • Bachelor's Degree
Python
DevOps
Azure
AWS
Cybersecurity
Zscaler
ISO 27001
NIST CSF
Zero Trust
Least Privilege
Apply
$87k – $188k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Yokneam • Tel Aviv
Python
MATLAB
Apply
$78k – $192k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Melbourne
Python
Bash
DevOps
Prometheus
SLURM
CI/CD
Kubernetes
Grafana
HPC
Apply
$81k – $215k per year (Estimated) • In office • 10+ years exp • Singapore
AI/ML
LLM
DevOps
HPC
Apply
$86k – $214k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Python
SQL
Python
FastAPI
Celery
Databases
PostgreSQL
Weaviate
Neo4j
Milvus
pgvector
Pinecone
Qdrant
ElasticSearch
OpenSearch
AI/ML
LangGraph
AutoGen
LangChain
DSPy
LlamaIndex
Model Context Protocol
Dagster
Prefect
Embeddings
Multimodal AI
Function Calling
AI Agents
Arize Phoenix
DeepEval
Haystack
Langfuse
LangSmith
Promptfoo
Pydantic AI
Ragas
Semantic Kernel
CrewAI
LLM
RAG
TruLens
Reranking
Hybrid Search
Time Series Forecasting
OpenAI
Human-in-the-Loop
Structured Outputs
Knowledge Graph
LLM Guardrails
Multi-Agent Systems
Tool Use
DevOps
gRPC
OpenTelemetry
WebSockets
CI/CD
Kubernetes
Argo Workflows
Vector
Cybersecurity
Defense in Depth
Apply
$89k – $222k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Python
Go
AI/ML
Fine-tuning
AI Agents
NVLink
DevOps
gRPC
Terraform
Helm
GitHub Actions
OpenTelemetry
Kustomize
Prometheus
GitLab CI
SLURM
CI/CD
GitOps
ArgoCD
Kubernetes
Grafana
SRE
Platform Engineering
GitHub
GitLab
HPC
Cybersecurity
Kyverno
Apply
$93k – $233k per year (Estimated) • In office • Full-Time • 5+ years exp • Sydney
Python
Go
C++
AI/ML
vLLM
CUDA Toolkit
Triton Inference Server
Embeddings
Quantization
Multimodal AI
Function Calling
AI Agents
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
RAG
Reranking
NVIDIA NIM
CUDA
Triton
Hugging Face
NCCL
NVLink
cuDNN
Speculative Decoding
KV Cache
Agentic Workflows
Tool Use
DevOps
CI/CD
GitOps
Kubernetes
Platform Engineering
Apply
$76k – $92k per year • Remote • Internship • Bachelor's Degree • San Francisco
SQL
Analytics
SSIS
SSAS
Management
Microsoft Project
Apply
$70k – $82k per year • Remote • Internship • San Francisco
Analytics
Microsoft Excel
Apply
$70k – $82k per year • In office • Internship • San Francisco
Apply
$79k – $95k per year • Remote • Full-Time • San Francisco
Apply
$213k – $374k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Chicago • New York • Atlanta • San Francisco
AI/ML
AI Agents
Agentforce
Agentic Workflows
Marketing
Salesforce
Apply
See all jobs
This is one of many
428,639 more open roles from verified company boards, updated every day.