594,528open jobs
27,125companies
84,688added this week
Browse all
Salary
$158k – $316k per year
Location
In office (Singapore)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.

  • Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling.

  • Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads.

  • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.

  • Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.

Preferred qualifications:

  • Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.

  • Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.

  • Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.

  • Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.

Bonus points if you have:

  • Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.

  • Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.

  • Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.

Logistics

  • Location: This role is based in Singapore.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
594,528 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
In office • Contractor • 1+ year exp • Bachelor's Degree • Hanoi
Python
AI/ML
Copilot
Cursor
Claude Code
Model Context Protocol
vLLM
Multimodal AI
Function Calling
AI Agents
Gemini
LLM
RAG
Reranking
Triton
OpenAI
Anthropic
OCR
Structured Outputs
LLM Guardrails
Speculative Decoding
KV Cache
Prompt Caching
Tool Use
DevOps
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Apply
$29k – $70k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Beijing
Python
AI/ML
JAX
Multimodal AI
Computer Vision
TensorFlow
PyTorch
DevOps
HPC
Apply
$31k – $62k per year (Estimated) • Remote/Hybrid • Moscow
Python
Python
SQLAlchemy
FastAPI
Celery
Pydantic
Alembic
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
LangGraph
LangChain
AI Agents
Transformers
LLM
RAG
KV Cache
DevOps
Rest API
WebSockets
CI/CD
Git
Docker
Apply
Remote/Hybrid • Full-Time • Master's Degree • Germany
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Reinforcement Learning
JAX
Computer Vision
AI Agents
TensorFlow
PyTorch
Hugging Face
Edge AI
World Models
Physical AI
DevOps
CI/CD
Robotics
ROS
Imitation Learning
Reinforcement Learning
Apply
Operations Engineer 10 hours ago
$30k – $63k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
PowerShell
AI/ML
Copilot
XLA
DevOps
Azure
Windows Server
Self-Healing
AIOps
SLI/SLO/SLA
Cybersecurity
Microsoft Entra ID
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 2 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
$68k – $154k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Singapore
Apply
GTM Engineer 9 hours ago
$100k – $150k per year • In office • Full-Time • 2+ years exp • Singapore
AI/ML
Reinforcement Learning
Post-training
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
AI Agents
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
LLM
Post-training
DevOps
Docker
Apply
See all jobs
This is one of many
594,528 more open roles from verified company boards, updated every day.