579,193open jobs
24,952companies
79,803added this week
Browse all
Salary
$200k – $400k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.

You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.

  • Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.

  • Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.

  • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.

  • Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.

Preferred qualifications:

  • Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.

  • Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.

  • Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.

  • Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.

Bonus points if you have:

  • Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.

  • Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.

  • Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
579,193 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
In office • Contractor • 1+ year exp • Bachelor's Degree • Hanoi
Python
AI/ML
Copilot
Cursor
Claude Code
Model Context Protocol
vLLM
Multimodal AI
Function Calling
AI Agents
Gemini
LLM
RAG
Reranking
Triton
OpenAI
Anthropic
OCR
Structured Outputs
LLM Guardrails
Speculative Decoding
KV Cache
Prompt Caching
Tool Use
DevOps
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Apply
$35k – $88k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Hanoi
Python
AI/ML
Copilot
Cursor
Claude Code
LoRA
Model Context Protocol
vLLM
Fine-tuning
RLHF
Multimodal AI
Function Calling
AI Agents
PEFT
Gemini
LLM
RAG
Reranking
Triton
OpenAI
Anthropic
DPO
SFT
Post-training
OCR
Structured Outputs
Context Engineering
LLM Guardrails
Speculative Decoding
KV Cache
Prompt Caching
Tool Use
DevOps
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Vector
Apply
$140k – $210k per year • Equity • Remote • Full-Time • 7+ years exp
Python
Go
TypeScript
Databases
Redis
DynamoDB
AI/ML
vLLM
Triton
Amazon SageMaker
TorchServe
Feature Store
DevOps
Terraform
OpenTelemetry
CloudFormation
Prometheus
CI/CD
AWS
Docker
Kubernetes
Grafana
Amazon EKS
AWS Lambda
Vector
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Master's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Multimodal AI
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
CUTLASS
Management
Agile
Apply
$31k – $62k per year (Estimated) • Remote/Hybrid • Moscow
Python
Python
SQLAlchemy
FastAPI
Celery
Pydantic
Alembic
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
LangGraph
LangChain
AI Agents
Transformers
LLM
RAG
KV Cache
DevOps
Rest API
WebSockets
CI/CD
Git
Docker
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 2 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
$115k – $180k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • Seattle • Raleigh • New York
Apply
$97k – $124k per year • In office • Full-Time • 1+ year exp • San Francisco
Design
Canva
Management
Slack
Google Workspace
Apply
$200k – $240k per year • Remote/Hybrid • San Francisco
AI/ML
AI Agents
LLM
RAG
Apply
$152k – $301k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Dallas • Austin • San Francisco • Fort Worth • Los Angeles
Marketing
Salesforce
Apply
$285k – $335k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco
DevOps
VMWare
containerd
Kubernetes
KVM
QEMU
Hyper-V
Apply
See all jobs
This is one of many
579,193 more open roles from verified company boards, updated every day.