601,883open jobs
28,504companies
85,665added this week
Browse all
Salary
$163k – $336k per year (Estimated)
Location
Remote/Hybrid (United States)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.

  • Deep understanding of transformer architectures and their variants.

  • Strong programming skills in Python with experience in PyTorch internals.

  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).

  • Ability to read and implement model architectures and inference techniques from research papers.

  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.

Preferred qualifications:

  • Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.

  • Familiarity with RL frameworks and algorithms for LLMs.

  • Experience with multimodal inference (audio/image/video/text).

  • Contributions to open-source ML or system infrastructure projects.

Bonus points if you have:

  • Implemented core features in vLLM or other inference engine projects.

  • Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).

  • Written widely-shared technical blogs or side projects on vLLM or LLM inference.

Logistics

  • Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.

  • Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
601,883 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$34k – $73k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Pune
Python
JavaScript
Java
COBOL
Java
Spring Boot
Hibernate
Databases
Db2
AI/ML
Copilot
Claude
AI Agents
RAG
Devin
Frontend
React.js
DevOps
CI/CD
Jenkins
Management
Agile
Scrum
Apply
$43k – $110k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Bengaluru
AI/ML
Multimodal AI
AI Agents
PEFT
Transformers
LLM
SFT
Management
Agile
Apply
$63k – $134k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • São Paulo
Python
Databases
Snowflake
Apply
$14k – $33k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Pune
Python
Java
Kotlin
Java
Spring Boot
Spring MVC
Spring Data JPA
Spring Security
Gradle
Testcontainers
Kotlin
Mockito
Databases
Apache Kafka
Mobile
JUnit
DevOps
gRPC
Splunk
GCP
OpenShift
GitHub Actions
Prometheus
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Bitbucket
GitHub
Amazon S3
Management
Agile
Apply
$33k – $71k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Pune
SQL
AI/ML
AI Agents
Edge AI
Analytics
Tableau
Management
Confluence
Jira
Agile
Scrum
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 3 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$200k – $400k per year • Remote • Full-Time
Python
Rust
AI/ML
vLLM
Ray
DevOps
Terraform
GCP
Helm
SLURM
Azure
AWS
Kubernetes
Apply
See all jobs
This is one of many
601,883 more open roles from verified company boards, updated every day.