vLLM Jobs - Remote & On-site

vLLM jobs curated from vetted startups and product companies, updated daily. Many vLLM roles offer remote or hybrid schedules and include transparent compensation to help you decide. Browse and apply today.

AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

Chubb

chubb.com
Wij bieden in de Benelux een scala aan verzekeringsproducten voor particulieren, families en bedrijven. Wij hebben onder andere verzekeringsproducten op het gebied van Ongevallen & Welzijn, Financial Lines, Transport, Brand- en Bedrijfsschade, Aansprakelijkheid, Reizen en Overlijdensrisico.
chubb.com • London • Dublin • Philadelphia • Taipei • Bengaluru • 1001-5000 employees
London • Dublin • Philadelphia • Taipei • Bengaluru • 1001-5000 employees
Open 125 daysVerified live · 25 min ago 5 months ago

AI Engineer (Fluent Portuguese & English)

$89k – $184k per year (Estimated) • Senior • 5+ years expRemote/Hybrid (London, United Kingdom) • Senior • Portuguese: C1 • English: C1 (Optional)
Python
AI/ML
Fine-tuning
LLM
LoRA
PEFT
Prompt Engineering
QLoRA
Quantization
RAG
Semantic Search
Triton
vLLM
Transformers
Semantic Search
Text-to-Speech
Speech Recognition
TGI
CUDA
CUDA Toolkit
Multimodal AI
OCR
DevOps
CI/CD
Docker
Kubernetes
WebSockets
Apply
Report

Satori Analytics

satorianalytics.com
About Us Data and AI changes everything! Who we are At Satori Analytics, we are strong believers in the transformative power of data and AI.
satorianalytics.com • HQ: Athens, Greece • Data Science • Artificial Intelligence • Data & Analytics
HQ: Athens, Greece • Data Science • Artificial Intelligence • Data & Analytics
Open 125 daysVerified live · 1 day ago 5 months ago

Senior MLOps Engineer

Senior • 5+ years expRemote/Hybrid (Athens, Greece) • Full-Time • Senior
Python
Databases
Apache Kafka
Databricks
Delta Lake
AI/ML
BentoML
Kubeflow
LLM
MLFlow
RAG
Triton Inference Server
Vertex AI
Triton
Amazon SageMaker
TorchServe
Feast
LangSmith
Quantization
vLLM
Feature Store
TGI
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
Platform Engineering
Prometheus
Pulumi
Terraform
Vector
Marketing
Salesforce
Apply
Report

Modal

modal.com
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.
modal.com • HQ: New York, United States • MLOps • LLM & Generative AI • Machine Learning • Information Technology • Hardware • Cloud Computing • AI Infrastructure • Artificial Intelligence • Est. 2021
HQ: New York, United States • MLOps • LLM & Generative AI • Machine Learning • Information Technology • Hardware • Cloud Computing • AI Infrastructure • Artificial Intelligence • Est. 2021
Open 132 daysTop 25% pay 5 months ago

Member of Technical Staff - ML Performance

$200k – $350k per year • Staff+ • 5+ years expIn office (New York, San Francisco, United States) • Full-Time • Staff
CUDA
CUDA Toolkit
Diffusion Models
TensorRT
vLLM
Analytics
Seaborn
Matplotlib
Apply
Open 180 daysVerified live · 1 day ago 6 months ago

Forward Deployed Engineer - ML

$52k – $154k per year (Estimated) • Junior • 2+ years expIn office (Stockholm, Sweden) • Full-Time • Junior
LLM
RLHF
SGLang
TRL
vLLM
Transformers
SFT
Analytics
Seaborn
Matplotlib
Apply
Open 188 daysTop 25% pay 7 months ago

Forward Deployed Engineer - ML

$180k – $250k per year • Junior • 2+ years expIn office (New York, San Francisco, United States) • Full-Time • Junior
LLM
RLHF
SGLang
TRL
vLLM
Transformers
SFT
Analytics
Seaborn
Matplotlib
Apply
Report

Binance

binance.com
Binance is the world's largest cryptocurrency exchange by trading volume, founded in 2017 by Changpeng Zhao and Yi He. The platform offers spot, margin and derivatives trading across hundreds of digital assets, alongside staking, savings products, payments, an institutional custody arm and a self-custodial Web3 wallet. The group also created BNB Chain, one of the most used smart contract networks, and now operates under a licensed regional structure after a 2023 settlement with United States authorities that installed new leadership and compliance oversight.
binance.com • HQ: Abu Dhabi, United Arab Emirates • Mining & Staking • Crypto Payments • Web3 & dApps • Blockchain & Crypto • Cryptocurrencies • Crypto Exchanges • 1001-5000 employees • Est. 2017
HQ: Abu Dhabi, United Arab Emirates • Mining & Staking • Crypto Payments • Web3 & dApps • Blockchain & Crypto • Cryptocurrencies • Crypto Exchanges • 1001-5000 employees • Est. 2017
Verified live · 2 hours ago 5 months ago

Pioneer Talent Program - Applied Data Scientist

Remote/Hybrid • Bachelor's Degree
Python
AI/ML
AI Agents
Chain-of-Thought
Hallucination
LLM
Prompt Engineering
TensorRT
TensorRT-LLM
vLLM
Function Calling
Model Context Protocol
Post-training
Apply
Verified live · 2 hours ago 5 months ago

Binance Accelerator Program - Data Scientist (LLM & Trading)

$61k – $175k per year (Estimated) • Junior • Remote (Taiwan, Hong Kong, Singapore) • Bachelor's Degree • Junior
AI Agents
DeepSpeed
Fine-tuning
Llama
LLaMA-Factory
LLM
PyTorch
Reinforcement Learning
RLHF
TensorFlow
Transformers
vLLM
Hugging Face
Post-training
SFT
Apply
Open 1451 dayVerified live · 2 hours ago 4 years ago

LLM Data Scientist/Algorithm Engineer (Fully Remote)

$64k – $190k per year (Estimated) • Junior • 2+ years exp • Remote (Japan, Taiwan, Hong Kong, Singapore, Australia) • Master's Degree • Full-Time • Junior • Chinese: B2
AutoGen
Chain-of-Thought
CrewAI
Hallucination
LangGraph
LLM
NLP
Prompt Engineering
Quantization
RAG
SGLang
vLLM
LangChain
DPO
SFT
AI Agents
Apply
Report

Ottimate

ottimate.com
ottimate (formerly plate iq) is the leading ap automation ai. ottimate is ap automation ai that provides a smarter way for ap managers, approvers, controllers, and cfos to work through the entire invoice lifecycle. with mature deep learning capab...
ottimate.com • Kansas City • Bengaluru • Chennai • Tampa • Chihuahua • Artificial Intelligence • Financial Services • FinTech • Est. 2014
Kansas City • Bengaluru • Chennai • Tampa • Chihuahua • Artificial Intelligence • Financial Services • FinTech • Est. 2014
Open 133 daysVerified live · 1 day ago 5 months ago

Director of AI Engineering

$167k – $305k per year (Estimated) • Executive • Remote (United States) • Full-Time • Architect
Python
Python
Celery
Databases
PostgreSQL
AI/ML
Anthropic SDK
Claude
Claude Code
Copilot
Cursor
Fine-tuning
Function Calling
LLM
LoRA
OpenAI SDK
QLoRA
RAG
Reranking
vLLM
PEFT
Anthropic
OpenAI
TGI
AI Agents
Model Context Protocol
OCR
DevOps
Platform Engineering
AWS
Analytics
A/B Testing
Apply
Report

Zencore

zencore.dev
Zencore is a leading cloud consulting and services firm. Above all else we value experience, hard work and the ability to keep things in perspective and have fun as we ensure our customers are delighted with the work we do.
zencore.dev • London • Mexico City • New York • Copenhagen • Berlin • Blockchain & Crypto • Data & Analytics • Information Technology • Est. 2013
London • Mexico City • New York • Copenhagen • Berlin • Blockchain & Crypto • Data & Analytics • Information Technology • Est. 2013
Open 136 days 5 months ago

Principal Architect, AI/ML

$139k – $249k per year (Estimated) • Lead • Remote (United States) • Master's Degree • Full-Time • Architect
AI Agents
Claude
Fine-tuning
Gemini
Google ADK
JAX
LangChain
Langfuse
LangGraph
LangSmith
Llama
LLM
LoRA
Mistral
PyTorch
Quantization
TensorRT
TensorRT-LLM
Vertex AI
vLLM
PEFT
LLM Guardrails
TPU
DevOps
AWS
Azure
GCP
Google GKE
Kubernetes
Cybersecurity
GDPR
Apply
Open 201 day 7 months ago

Principal Architect, AI/ML

$80k – $190k per year (Estimated) • Lead • Remote (United Kingdom) • Master's Degree • Full-Time • Architect
AI Agents
Claude
Fine-tuning
Gemini
Google ADK
JAX
LangChain
Langfuse
LangGraph
LangSmith
Llama
LLM
LoRA
Mistral
PyTorch
Quantization
TensorRT
TensorRT-LLM
Vertex AI
vLLM
PEFT
LLM Guardrails
TPU
DevOps
AWS
Azure
GCP
Google GKE
Kubernetes
Cybersecurity
GDPR
Apply
Report

H Company

hcompany.ai
H Company is an artificial intelligence research company headquartered in Paris, France, and founded in 2023 by Charles Kantor with a founding team drawn largely from Google DeepMind. The company develops agentic AI systems, including Runner H, an agent that carries out multi-step tasks in a browser on the user's behalf. It raised one of the largest seed rounds in European history and positions itself as a French alternative in the agent tooling market.
hcompany.ai • HQ: Paris, France • Business Process Automation (BPA) • Software • Multimodal AI • LLM & Generative AI • AI Agents • Artificial Intelligence • 201-500 employees • Est. 2023
HQ: Paris, France • Business Process Automation (BPA) • Software • Multimodal AI • LLM & Generative AI • AI Agents • Artificial Intelligence • 201-500 employees • Est. 2023
Open 138 daysVerified live · 1 day ago 5 months ago

Research Engineer, Model Inference & Serving - Paris

$77k – $183k per year (Estimated)Remote/Hybrid (Paris, France) • Full-Time
C++
Python
Rust
C++
PyTorch C++
AI/ML
AI Agents
JAX
Multimodal AI
PyTorch
SGLang
vLLM
Edge AI
CUDA Toolkit
llama.cpp
LLM
MLX ML
ONNX
Quantization
TensorRT
TensorRT-LLM
CUDA
Triton
DevOps
Kubernetes
Apply
Open 142 daysVerified live · 1 day ago 5 months ago

Research Engineer, Model Inference & Serving - London

$104k – $219k per year (Estimated)Remote/Hybrid (London, United Kingdom) • Full-Time
C++
Python
Rust
C++
PyTorch C++
AI/ML
AI Agents
JAX
Multimodal AI
PyTorch
SGLang
vLLM
Edge AI
CUDA Toolkit
llama.cpp
LLM
MLX ML
ONNX
Quantization
TensorRT
TensorRT-LLM
CUDA
Triton
DevOps
Kubernetes
Apply
Report

Cyncly

cyncly.com
Cyncly was created in September of 2022 as the new brand to unite Compusoft, 2020 and their affiliate companies after the two companies merged in 2021. The combined group created a global software powerhouse with more than 2,300 employees and 70,000+ customers across 100+ countries. Our company brings the best together, providing specialized visualization, sales, manufacturing and content solutions for customers wanting to bring spaces to life and bring life to spaces.
cyncly.com • New York • Pune • London • Lisbon • Bengaluru • Software • Manufacturing • Artificial Intelligence • 1001-5000 employees
New York • Pune • London • Lisbon • Bengaluru • Software • Manufacturing • Artificial Intelligence • 1001-5000 employees
Open 145 daysVerified live · 1 day ago 5 months ago

Sr. AI/ML Engineer - SDLC

$28k – $71k per year (Estimated) • Senior • 5+ years expIn office (Kochi, India) • Bachelor's Degree • Senior
Java
Python
Databases
ArangoDB
Neo4j
AI/ML
AI Agents
AutoGen
DeepSeek
Fine-tuning
LangChain
LangGraph
Llama
LLM
Ollama
Prompt Engineering
Qwen
RAG
vLLM
GraphRAG
Hugging Face
Knowledge Graph
OpenAI
DevOps
CI/CD
Vector
Web3
Abstract
Apply
Report

Inworld AI

inworld.ai
Inworld AI is a company founded in 2021 that builds a runtime for artificial intelligence characters and agents in games and interactive media. Its engine handles dialogue, memory, emotion and safety so studios can drop believable non-player characters into a title without building model infrastructure. The company works with game developers and consumer platforms and has partnered with Microsoft and Nvidia.
inworld.ai • HQ: Mountain View, United States • Speech & Audio AI • LLM & Generative AI • Artificial Intelligence • Est. 2021
HQ: Mountain View, United States • Speech & Audio AI • LLM & Generative AI • Artificial Intelligence • Est. 2021
Open 145 daysVerified live · 1 day ago 5 months ago

Staff / Principal Machine Learning Engineer, Serving - Switzerland

$140k – $334k per year (Estimated) • Lead • 10+ years exp • Remote (Switzerland) • Relocation • Visa sponsorship • PhD • Full-Time • Principal
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
Open 145 daysTop 25% pay 5 months ago

Staff / Principal Machine Learning Engineer, Serving - UK

$190k – $271k per year • Lead • 10+ years expIn office • Relocation • Visa sponsorship • PhD • Full-Time • Principal
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
Open 145 daysVerified live · 1 day ago 5 months ago

Senior / Lead Machine Learning Engineer, Serving - Serbia

Lead • 10+ years expIn office • Relocation • PhD • Full-Time • Staff
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
Report

SatSure

satsure.co
SatSure builds decision analytics from satellite imagery and other geospatial data. Its products serve agricultural lending, insurance, infrastructure and climate risk. Banks and government bodies use its crop and land intelligence.
satsure.co • HQ: Bengaluru, India • Data & Analytics • Decision Intelligence • Artificial Intelligence • Est. 2017
HQ: Bengaluru, India • Data & Analytics • Decision Intelligence • Artificial Intelligence • Est. 2017
Open 151 day 5 months ago

Senior Cloud/DevOps Engineer (SDE - 3)

$25k – $63k per year (Estimated) • Senior • 5+ years expRemote/Hybrid (Bengaluru, India) • Bachelor's Degree • Full-Time • Senior
Python
Python
Dask
AI/ML
Airflow
KServe
LLM
Ray
Ray Serve
vLLM
DevOps
Amazon EC2
Amazon EKS
Ansible
ArgoCD
AWS
Azure
Azure AKS
Error Budget
FinOps
GCP
GitOps
Google GKE
Grafana
Helm
Istio
Karpenter
Kubernetes
Loki
Mimir
OpenTelemetry
Prometheus
Service Mesh
SLI/SLO/SLA
Terraform
Amazon CloudWatch
Amazon S3
IAM
Bitbucket
CI/CD
Datadog
Envoy
Jenkins
Cybersecurity
CIS Benchmarks
ISO 27001
Keycloak
Apply
Report

GENIEE

geniee.co.jp
GENIEE is a Tokyo advertising technology company founded in 2010 that builds supply side and demand side advertising platforms. Its systems help publishers in Japan and Southeast Asia maximise revenue from programmatic advertising, and it has acquired sales and marketing software businesses. The company operates across several Asian markets from its Tokyo base.
geniee.co.jp • HQ: Tokyo, Japan • Semiconductors • Document Management • Embedded Systems • Est. 2010
HQ: Tokyo, Japan • Semiconductors • Document Management • Embedded Systems • Est. 2010
Open 152 daysVerified live · 1 day ago 6 months ago

【JAPAN AI】Research Engineer, LLM modeling / English

$47k – $137k per year (Estimated) • Equity • Remote/Hybrid (Tokyo, Japan) • Master's Degree • Full-Time • Japanese: C1
Python
JavaScript
TypeScript
AI/ML
AI Agents
ChatGPT
Claude
Cursor
Fine-tuning
JAX
Knowledge Distillation
LLM
Multimodal AI
NLP
Prompt Engineering
PyTorch
Quantization
Reinforcement Learning
RLHF
Synthetic Data
Transformers
DPO
Edge AI
Google AI Studio
Function Calling
Claude Code
vLLM
Weights & Biases
Devin
Frontend
Next.js
React.js
DevOps
Docker
GCP
Kubernetes
GitHub
Management
Confluence
Google Workspace
Slack
Apply
Open 152 daysVerified live · 1 day ago 6 months ago

【JAPAN AI】Research Engineer, LLM modeling / Japanese

$45k – $131k per year (Estimated)In office (Tokyo, Japan) • Full-Time • Japanese: C1
Python
TypeScript
JavaScript
AI/ML
ChatGPT
Claude
Claude Code
Cursor
JAX
LLM
NLP
PyTorch
RAG
RLHF
Transformers
vLLM
Weights & Biases
Devin
DPO
Google AI Studio
AI Agents
Frontend
Next.js
React.js
DevOps
Docker
GCP
Kubernetes
GitHub
Management
Confluence
Google Workspace
Slack
Apply
Report

Sola

sola.ai
Sola is an agentic process automation platform that makes it easy for companies to automate data entry, scraping, and processing flows using LLMs and computer vision – delivering results much faster and with less effort than traditional RPA tools.
sola.ai • New York • Data Science • Computer Vision • Artificial Intelligence
New York • Data Science • Computer Vision • Artificial Intelligence
Open 158 daysTop 25% pay 6 months ago

Software Engineer, ML Platform

$150k – $350k per year • Middle • 3+ years expIn office (New York, United States) • Relocation • Full-Time • Middle
Python
TypeScript
AI/ML
AI Agents
DSPy
Fine-tuning
PyTorch
Anthropic
Hugging Face
TensorRT
vLLM
DevOps
AWS
Kubernetes
Apply
Report

Transluce

transluce.org
Infrastructure for understanding AI Infrastructure for understanding AI. Transluce is a non-profit research lab building the public tech stack for scalable oversight of AI.
transluce.org • San Francisco • 11-50 employees
San Francisco • 11-50 employees
Open 502 daysTop 25% pay 1 year ago

AI Systems Engineer

$350k – $600k per year • In office (San Francisco, United States) • Full-Time
Python
AI/ML
LLM
vLLM
AI Agents
Apply
Report

Makersite

makersite.io
Makersite is an award-winning data software company that specializes in providing sustainability data and product lifecycle intelligence solutions. The company empowers engineers, procurement teams, and sustainability experts to transform products and supply chains by making sustainable product decisions at scale. Makersite integrates artificial intelligence, data, and applications to offer features like automated lifecycle assessments, supply chain risk management, and AI-enabled ecodesign.
makersite.io • Stuttgart • Supply Chain • Science & Engineering • Environmental Science • Est. 2018
Stuttgart • Supply Chain • Science & Engineering • Environmental Science • Est. 2018
Open 166 daysVerified live · 8 hours ago 6 months ago

Senior Data Scientist (m/f/x)

$93k – $116k per year • Senior • 5+ years exp • Remote (Germany) • Master's Degree • Senior
Python
Python
FastAPI
Databases
PostgreSQL
Chroma
pgvector
Pinecone
Weaviate
AI/ML
AI Agents
CrewAI
LangGraph
RAG
AWS Bedrock
Claude
Fine-tuning
LangChain
MLFlow
Prompt Engineering
Vertex AI
vLLM
Weights & Biases
Hugging Face
LLMOps
OpenAI
TGI
Frontend
GraphQL
DevOps
CI/CD
Vector
AWS
Azure
Docker
GCP
GitHub Actions
Kubernetes
GitHub
QA
Rest-Assured
Apply
Report

Zyphra

zyphra.com
COMPANY RESEARCH CLOUD Zyphra Cloud Login Two sides. Two sides.
zyphra.com • HQ: San Francisco, United States • AI Infrastructure • Artificial Intelligence • LLM & Generative AI
HQ: San Francisco, United States • AI Infrastructure • Artificial Intelligence • LLM & Generative AI
Open 167 daysVerified live · 2 days ago 6 months ago

Platform Engineer

$125k – $280k per year (Estimated)In office (San Francisco, United States) • Relocation • Full-Time
Ray
SGLang
vLLM
Triton
DevOps
Ansible
AWS
CI/CD
Docker
GCP
Kubernetes
SLURM
Terraform
Chaos Engineering
Apply
Report

MLabs

mlabs.city
MLabs is a global software engineering consultancy specializing in high-assurance software engineering, formal verification, and mission-critical systems. Founded in 2018, the company focuses on functional programming languages like Haskell and Rust to build secure, high-performance infrastructure for blockchain, Web3, fintech, and AI projects.
mlabs.city • HQ: London, United Kingdom • Information Technology • Blockchain & Crypto • Software
HQ: London, United Kingdom • Information Technology • Blockchain & Crypto • Software
Verified live · 1 day ago 6 months ago

Staff Software Engineer - Backend & AI Infra

$228k – $379k per year (Estimated) • Equity • Staff+ • Remote (United States) • Staff
Go
Node JS
Python
TypeScript
JavaScript
Databases
ClickHouse
Google BigQuery
PostgreSQL
Redis
TimescaleDB
AI/ML
LLM
Model Context Protocol
Anthropic
TensorRT
TensorRT-LLM
vLLM
AI Agents
TGI
DevOps
Amazon EKS
AWS
CI/CD
Kubernetes
WebSockets
Apply
Report

Domyn

domyn.com
Domyn is an artificial intelligence company headquartered in Milan, Italy, and founded in 2016 under the name iGenius. The company builds large language models and decision intelligence tools aimed at regulated industries, including sovereign models trained for specific languages and jurisdictions. It works with banks, insurers, and public sector bodies in Europe that need AI deployed inside their own infrastructure rather than on a public model provider.
domyn.com • HQ: Milan, Italy • Financial Consulting • Financial Services • Business Intelligence • Data & Analytics • Decision Intelligence • LLM & Generative AI • Artificial Intelligence • Est. 2016
HQ: Milan, Italy • Financial Consulting • Financial Services • Business Intelligence • Data & Analytics • Decision Intelligence • LLM & Generative AI • Artificial Intelligence • Est. 2016
Open 185 daysVerified live · 1 day ago 7 months ago

Senior DevOps Engineer

$123k – $239k per year (Estimated) • Senior • 6+ years expIn office (New York, United States) • Full-Time • Senior
Python
Databases
PostgreSQL
GraphDB
AI/ML
AI Agents
Kubeflow
Ray
vLLM
DevOps
AIOps
Amazon EKS
Ansible
ArgoCD
AWS
Azure
Azure AKS
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Google GKE
Grafana
Jenkins
Kubernetes
Prometheus
Terraform
Vector
GitHub
GitLab
IAM
Apply
Report

Parspec

parspec.io
Parspec is a technology company that leverages AI to help sales agents and distributors by simplifying the process of discovering and sourcing the best available construction products and materials.
parspec.io • Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Open 174 daysVerified live · 1 day ago 6 months ago

AI Ops Engineer

$30k – $75k per year (Estimated) • Senior • 5+ years expIn office (Bengaluru, India) • Bachelor's Degree • Full-Time • Senior
Python
Python
Asyncio
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
Weaviate
PostgreSQL
Qdrant
AI/ML
AWS Bedrock
Kubeflow
LiteLLM
LLM
MLFlow
Portkey
Ray
vLLM
AWQ
CUDA
CUDA Toolkit
Embeddings
GPTQ
Hallucination
Langfuse
LoRA
Prompt Engineering
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
PEFT
Amazon SageMaker
AWS Trainium
LLM Guardrails
LLMOps
NCCL
NVLink
TGI
Multimodal AI
AI Agents
DevOps
AIOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
GitHub Actions
Grafana
Kubernetes
OpenTelemetry
Prometheus
Terraform
Vector
Karpenter
Platform Engineering
Amazon EventBridge
Amazon S3
API Gateway
IAM
AWS Step Functions
GitHub
Cybersecurity
Least Privilege
Apply
Report

Percepta

percepta.com
Headquartered in Dearborn, Michigan, Percepta is a global customer experience (CX) and BPO services provider specializing in the automotive sector. Originally established as a joint venture between Ford Motor Company and TeleTech (now TTEC), the company delivers customer care, technical support, warranty administration, and dealer relations services. Additionally, it operates contact centers and support networks worldwide to help automotive brands optimize customer loyalty and streamline operations.
percepta.com • HQ: Dearborn, United States • Professional Services
HQ: Dearborn, United States • Professional Services
Open 220 daysVerified live · 6 hours ago 8 months ago

Research Engineer / Scientist – Reinforcement Learning (RL)

$195k – $427k per year (Estimated)In office (New York, Boston, United States) • Master's Degree • Full-Time
Python
AI/ML
LLM
Ray
Reinforcement Learning
SGLang
vLLM
Anthropic
Post-training
AI Agents
DevOps
Amazon EKS
AWS
Kubernetes
Apply
Report

Inference

inference.ai
Models. Agents. GPUs. Whatever you need — we've got it. Ghost agent VMs, Maestro model routing, Engine wholesale GPUs, and Academy, on one platform.
inference.ai • San Francisco
San Francisco
Open 221 dayTop 25% pay 8 months ago

Senior Software Engineer - Model Performance

$220k – $320k per year • Senior • 2+ years expIn office (San Francisco, United States) • Full-Time • Senior
C++
Python
C#
C++
PyTorch C++
C#
.NET
AI/ML
CUDA Toolkit
Knowledge Distillation
LLM
LoRA
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
CUDA
GPT-5
Embeddings
DevOps
Docker
Kubernetes
GitHub
Apply
Report
vLLM Jobs - Remote & On-site
Frequently asked questions
vLLM: How many jobs are available now?
There are 562 vLLM job openings listed on Alion right now; listings are updated daily.
vLLM: What is the typical salary for vLLM roles?
The typical vLLM role on Alion averages $241,236 USD annually.
vLLM: Are remote, relocation, or visa sponsorship options available?
vLLM listings include remote, hybrid, and on-site roles, and some companies may provide relocation or visa sponsorship-check individual postings.
vLLM: How do I apply for vLLM jobs on Alion?
Sign in to Alion, upload your resume, and apply directly from the vLLM job pages; applications are sent to the hiring company.