594,528open jobs
27,125companies
84,688added this week
Browse all
Salary
$176k – $332k per year (Estimated)
Location
Remote (United States)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

About the Role

This is a globally remote opportunity. We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems. This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.

You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure. You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after. The work you do will directly impact how the world runs AI inference.

Potential focus areas include:

  • Inference Runtime: Push the boundaries of LLM and diffusion model serving. Work at the core of vLLM to optimize how models execute across diverse hardware and architectures.

  • Kernel Engineering: Write the low-level kernels and optimizations that make vLLM the fastest inference engine in the world, running on hundreds of accelerator types.

  • Performance & Scale: Build the distributed systems that power inference at global scale-design foundational layers enabling vLLM to serve models across thousands of accelerators with minimal latency.

  • Cloud Orchestration: Build the operational backbone for cluster management, deployment automation, and production monitoring that enables teams worldwide to serve AI models without friction.

What We're Looking For

We're looking for engineerswho thrive with autonomy. You should be able to take a vague problem statement and turn it into shipped code with minimal supervision. You communicate proactively, over-communicate context across time zones, and know when to ask for help versus when to push forward independently.

Core Requirements:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar

  • Demonstrated ability to work autonomously and drive projects to completion without close supervision

  • Excellent asynchronous communication skills and ability to collaborate effectively across time zones

  • Strong track record of shipping high-impact work in complex technical environments

  • Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure

Technical Depth (strong in at least two):

  • CUDA kernels or equivalent (Triton, TileLang, Pallas) with deep understanding of GPU architecture

  • High-performance distributed systems in Rust, Go, or C++

  • Python with PyTorch internals and LLM inference systems (vLLM, TensorRT-LLM, SGLang)

  • Kubernetes, container orchestration, and infrastructure-as-code at scale

  • Transformer architectures, KV-cache memory management, and model serving

Preferred Qualifications:

  • Contributions to vLLM or other major open-source ML/systems projects

  • Experience with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)

  • Knowledge of quantization techniques, ML-specific kernel optimization, or compiler technologies

  • Track record of improving system reliability and performance at scale

  • Written widely-shared technical blogs or impactful side projects in the ML infrastructure space

Logistics

  • Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.

  • Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
594,528 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$184k – $288k per year • In office • Full-Time • 6+ years exp • Master's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Multimodal AI
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
CUTLASS
Management
Agile
Apply
AI Engineer 1 day ago
$22k – $59k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Python
FastAPI
Databases
PostgreSQL
Weaviate
Chroma
Milvus
pgvector
Pinecone
AI/ML
LangGraph
AutoGen
LangChain
Claude
LlamaIndex
vLLM
Fine-tuning
Prompt Engineering
Multimodal AI
AI Agents
TensorRT
TensorRT-LLM
TGI
Llama
Mistral
TensorFlow
PyTorch
CrewAI
Gemini
LLM
RAG
Hallucination
Agentic Workflows
Multi-Agent Systems
DevOps
Rest API
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Vector
Apply
$17k – $42k per year (Estimated) • Remote • Full-Time • Moscow
Python
Bash
Databases
OpenSearch
AI/ML
Model Context Protocol
vLLM
Ollama
LLM
DevOps
Terraform
Docker Compose
Helm
Prometheus
Yandex Cloud
GitLab CI
CI/CD
GitOps
ArgoCD
Docker
Kubernetes
Grafana
Harbor
GitLab
Cybersecurity
SBOM
Apply
R&D Engineer 1 day ago
$101k – $211k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Python
AI/ML
CUDA Toolkit
YOLO
Fine-tuning
Quantization
Knowledge Distillation
Computer Vision
PyTorch
CUDA
Edge AI
Model Distillation
Robotics
Localization
Apply
$140k – $210k per year • Equity • Remote • Full-Time • 7+ years exp
Python
Go
TypeScript
Databases
Redis
DynamoDB
AI/ML
vLLM
Triton
Amazon SageMaker
TorchServe
Feature Store
DevOps
Terraform
OpenTelemetry
CloudFormation
Prometheus
CI/CD
AWS
Docker
Kubernetes
Grafana
Amazon EKS
AWS Lambda
Vector
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 2 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
See all jobs
This is one of many
594,528 more open roles from verified company boards, updated every day.