601,883open jobs
28,504companies
85,665added this week
Browse all
Salary
$158k – $316k per year
Location
In office (Singapore)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.

  • Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.

  • Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.

  • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.

  • Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.

Preferred qualifications:

  • Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.

  • Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.

  • Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.

  • Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.

Bonus points if you have:

  • Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.

  • Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.

  • Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.

Logistics

  • Location: This role is based in Singapore.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
601,883 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
Remote/Hybrid • 3+ years exp • Hong Kong
Python
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Qwen
DeepSeek
vLLM
SGLang
Langfuse
LiteLLM
Ollama
Portkey
TGI
Llama
Mistral
LLM
RAG
Kong AI Gateway
MiniMax
OpenAI
Text-to-Speech
LLM Evaluation
LLM Guardrails
KV Cache
Frontend
GraphQL
DevOps
OpenTelemetry
WebRTC
Kong
Prometheus
CI/CD
Kubernetes
Grafana
Graylog
IAM
Analytics
ETL/ELT
Apply
Remote/Hybrid • 3+ years exp • Seoul
Python
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Qwen
DeepSeek
vLLM
SGLang
Langfuse
LiteLLM
Ollama
Portkey
TGI
Llama
Mistral
LLM
RAG
Kong AI Gateway
MiniMax
OpenAI
Text-to-Speech
LLM Evaluation
LLM Guardrails
KV Cache
Frontend
GraphQL
DevOps
OpenTelemetry
WebRTC
Prometheus
CI/CD
Kubernetes
Grafana
Graylog
IAM
Analytics
ETL/ELT
Apply
$37k – $85k per year (Estimated) • Remote/Hybrid • 3+ years exp • Tokyo
Python
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Qwen
DeepSeek
vLLM
SGLang
Langfuse
LiteLLM
Ollama
Portkey
TGI
Llama
Mistral
LLM
RAG
Kong AI Gateway
MiniMax
OpenAI
Text-to-Speech
LLM Evaluation
LLM Guardrails
KV Cache
Frontend
GraphQL
DevOps
OpenTelemetry
WebRTC
Prometheus
CI/CD
Kubernetes
Grafana
Graylog
IAM
Analytics
ETL/ELT
Apply
Remote/Hybrid • 3+ years exp • Singapore
Python
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Qwen
DeepSeek
vLLM
SGLang
Langfuse
LiteLLM
Ollama
Portkey
TGI
Llama
Mistral
LLM
RAG
Kong AI Gateway
MiniMax
OpenAI
Text-to-Speech
LLM Evaluation
LLM Guardrails
KV Cache
Frontend
GraphQL
DevOps
OpenTelemetry
WebRTC
Prometheus
CI/CD
Kubernetes
Grafana
Graylog
IAM
Analytics
ETL/ELT
Apply
Remote/Hybrid • 3+ years exp • Bogotá
Python
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Qwen
DeepSeek
vLLM
SGLang
Langfuse
LiteLLM
Ollama
Portkey
TGI
Llama
Mistral
LLM
RAG
BERT
Kong AI Gateway
MiniMax
OpenAI
Text-to-Speech
LLM Evaluation
LLM Guardrails
KV Cache
Frontend
GraphQL
DevOps
OpenTelemetry
WebRTC
Prometheus
CI/CD
Kubernetes
Grafana
Graylog
IAM
Analytics
ETL/ELT
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 3 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
$68k – $154k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Singapore
Apply
GTM Engineer 1 day ago
$100k – $150k per year • In office • Full-Time • 2+ years exp • Singapore
AI/ML
Reinforcement Learning
Post-training
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
AI Agents
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
LLM
Post-training
DevOps
Docker
Apply
See all jobs
This is one of many
601,883 more open roles from verified company boards, updated every day.