747,513open jobs
44,904companies
108,446added this week
Browse all
Salary
$156k – $312k per year
Location
In office (Singapore)
Seniority
Staff
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 24, 2026. Inferact scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world.

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

About the Role

We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.

You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.

  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.

  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.

  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.

  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.

  • Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.

  • Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Preferred qualifications:

  • Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.

  • Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.

  • Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.

  • Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.

  • Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.

  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.

Bonus points if you have:

  • Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.

  • Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.

  • Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.

  • Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.

  • Carried real operational responsibility for infrastructure used by many engineers or researchers.

Logistics

  • Location: This role is based in Singapore.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
747,513 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Singapore
≈ $87k – $217k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Singapore
C++
AI/ML
Machine Learning
Apply
≈ $87k – $217k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Singapore
C++
AI/ML
Machine Learning
Apply
≈ $86k – $215k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Singapore
Python
Go
Java
C++
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
LLM
RAG
Context Engineering
Tool Use
DevOps
Linux
Apply
≈ $87k – $217k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Singapore
Python
Go
Java
SQL
Databases
ClickHouse
Presto
Apache Kafka
Trino
StarRocks
AI/ML
Copilot
Spark
Prompt Engineering
Function Calling
AI Agents
Flink
LLM
RAG
LLM Guardrails
Tool Use
DevOps
Platform Engineering
Apply
≈ $86k – $215k per year (Estimated) • Hybrid • 6+ years exp • Singapore
Python
TypeScript
ABAP
ABAP
SAP BTP
AI/ML
Claude
Model Context Protocol
AI Agents
RAG
Multi-Agent Systems
DevOps
Terraform
Apply
≈ $42k – $102k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Python
Go
Databases
Databricks
Amazon Aurora
AI/ML
MLFlow
AWS Bedrock
LLM
Ray
DevOps
Terraform
GitHub Actions
CI/CD
AWS
Kubernetes
Apply
AI Tech Lead 1 day ago
≈ $28k – $60k per year (Estimated) • Hybrid • Moscow
Python
Python
FastAPI
Databases
Chroma
Milvus
RabbitMQ
Apache Kafka
AI/ML
Copilot
LangGraph
LangChain
vLLM
Fine-tuning
Function Calling
Chain-of-Thought
AI Agents
NLP
SGLang
Langflow
Langfuse
LangSmith
LLM
RAG
LLMOps
Chips/EDA
PoC Library
Management
n8n
Apply
AI Engineer 1 day ago
Remote (India) • Full-Time • Pune
Python
Python
Flask
FastAPI
Databases
Supabase
AI/ML
LangGraph
AutoGen
LangChain
Claude
DSPy
LlamaIndex
LoRA
Model Context Protocol
vLLM
Fine-tuning
Prompt Engineering
Multimodal AI
Chain-of-Thought
AI Agents
OpenAI SDK
PEFT
Transformers
TensorFlow
PyTorch
CrewAI
Gemini
LLM
RAG
LLMOps
GPT-4
Agentic Workflows
Multi-Agent Systems
Vercel AI SDK
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Management
n8n
Apply
$11k – $20k per year • Remote (EAEU) • Moscow
Python
JavaScript
PHP
TypeScript
SQL
PHP
Bitrix
DevOps
Rest API
Git
Apply
Remote (Belgium) • Full-Time • 2+ years exp • Belgium
Python
JavaScript
Python
Django
Celery
Ruff
Django REST Framework
flake8
Databases
MySQL
PostgreSQL
Redis
AI/ML
Machine Learning
Frontend
Bootstrap
JQuery
DevOps
Azure
CI/CD
Git
AWS
Docker
Linux
Cybersecurity
CVE
CWE
CVSS
Management
Agile
Scrum
QA
Selenium
Playwright
Apply
$156k – $312k per year • In office • Full-Time • Singapore
AI/ML
vLLM
Quantization
JAX
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
TPU
MLIR
XLA
KV Cache
Machine Learning
Apply
$156k – $312k per year • In office • Full-Time • Singapore
AI/ML
vLLM
Quantization
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Triton
ROCm
MLIR
KV Cache
Machine Learning
Apply
$156k – $312k per year • In office • Full-Time • Singapore
Rust
C++
AI/ML
vLLM
InfiniBand
NVLink
Apply
$200k – $400k per year • In office • Full-Time • Singapore
Python
C++
C++
LLVM
AI/ML
vLLM
CUDA Toolkit
Quantization
CUDA
Triton
TPU
MLIR
XLA
Apply
$156k – $312k per year • In office • Full-Time • Singapore
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
≈ $86k – $215k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Singapore
Python
Go
Java
C++
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
LLM
RAG
Context Engineering
Tool Use
DevOps
Linux
Apply
≈ $104k – $208k per year (Estimated) • In office • 2+ years exp • Master's Degree • Singapore
Python
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Fine-tuning
Multimodal AI
Function Calling
AI Agents
LLM
RAG
DPO
SFT
Post-training
Pre-training
Context Engineering
Tool Use
Reward Modeling
Apply
≈ $54k – $115k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Singapore
Management
Outlook
Apply
≈ $66k – $145k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Singapore
Analytics
Microsoft Excel
Apply
In office • Bachelor's Degree • Singapore
Robotics
Digital Twin
IoT
OPC UA
Apply
See all jobs
This is one of many
747,513 more open roles from verified company boards, updated every day.