TensorRT-LLM Jobs - Remote & On-site

TensorRT-LLM jobs on Alion link machine learning engineers and infrastructure builders with vetted companies working on high-performance LLM inference; listings are updated daily and often include remote or hybrid roles. Browse and apply today.

AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

Yottalabs

yottalabs.com
yottalabs.com • Hong Kong • Singapore
Hong Kong • Singapore
Verified live · 10 hours ago 3 months ago

Research Engineer - AI Systems

$136k – $298k per year (Estimated)Hong KongSingapore • Remote (United States, Canada, Hong Kong, Singapore) • Full-Time
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
AWS Trainium
ROCm
Mixture of Experts
RLHF
DevOps
AWS
Apply
Report

Pure Storage

purestorage.com
Pure Storage is an enterprise technology company that specializes in all-flash data storage systems, data management software, and integrated cloud services. The firm focuses on replacing legacy disk-based infrastructure with modern, high-performance flash technology to help organizations simplify their operations and enhance data accessibility. Through a software-led platform and subscription-based service model, the company provides resilient, scalable solutions designed to support diverse workloads ranging from artificial intelligence to cloud-native applications.
purestorage.com • HQ: Santa Clara, United States • Information Technology • Data Centers • Hardware • Storage • 1001-5000 employees • Est. 2009
HQ: Santa Clara, United States • Information Technology • Data Centers • Hardware • Storage • 1001-5000 employees • Est. 2009
Verified live · 2 hours ago 3 months ago

Senior LLM AI Engineer

Senior • In office (Prague, Czech Republic) • Senior
Python
SQL
AI/ML
Claude
Fine-tuning
Gemini
LLM
ONNX
PyTorch
RLHF
Semantic Search
TensorRT
TensorRT-LLM
Transformers
vLLM
DPO
Hugging Face
Semantic Search
AI Agents
Model Context Protocol
Apply
Report

Commonai

commonai.org
Home About Programmes News Careers About Us CommonAI provides a new backbone for AI infrastructure. We believe that collaborative engineering is the means to accelerate AI innovation.
commonai.org • Cambridge • Clinical Research
Cambridge • Clinical Research
Verified live · 1 day ago 3 months ago

Senior Software Engineer (vLLM)

$99k – $179k per year (Estimated) • Senior • In office (Cambridge, United States) • Bachelor's Degree • Full-Time • Senior
C++
Python
Rust
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
LLM
PyTorch
Ray
TensorRT
TensorRT-LLM
vLLM
Hugging Face
TGI
TPU
DevOps
CI/CD
Apply
Report

Amgen

amgen.com
Amgen is a global biopharmaceutical pioneer headquartered in Thousand Oaks, California, that specializes in discovering, developing, and manufacturing innovative biologic therapies. The company focuses on treating serious illnesses with high unmet medical needs across key areas including oncology, cardiovascular disease, inflammation, rare diseases, and nephrology. Leveraging advanced human genetics, molecular engineering, and biosimilar development, it serves millions of patients worldwide through established blockbuster treatments and cutting-edge pipelines.
amgen.com • HQ: Thousand Oaks, United States • Biotechnology • Health Care • Pharmaceuticals • 1001-5000 employees • Est. 1980
HQ: Thousand Oaks, United States • Biotechnology • Health Care • Pharmaceuticals • 1001-5000 employees • Est. 1980
Verified live · 1 day ago 3 months ago

Principal Machine Learning Engineer - Forecasting

$32k – $76k per year (Estimated) • Lead • 12+ years expIn office (Hyderabad, India) • Full-Time • Principal
Python
SQL
Python
FastAPI
AI/ML
AI Agents
Dagster
JAX
LLM
MLFlow
NLP
Prefect
PyTorch
Ray
Scikit-learn
Spark
TensorFlow
Human-in-the-Loop
LLM Guardrails
Function Calling
SGLang
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
RAG
DevOps
CI/CD
Analytics
A/B Testing
Apply
Report

Illumio

illumio.com
Illumio is a pioneer in cybersecurity that specializes in Zero Trust Segmentation and breach containment across hybrid, cloud, and enterprise environments. Its platform provides real-time visibility into application traffic, helping organizations isolate critical workloads and prevent lateral movement of threats like ransomware. By adopting a default untrusted model, Illumio enables major enterprises to maintain operational resilience and stop cyber incidents from turning into widespread disasters.
illumio.com • San Francisco • Sunnyvale • London • Melbourne • Dubai • Network Security • Cybersecurity • Cloud Security
San Francisco • Sunnyvale • London • Melbourne • Dubai • Network Security • Cybersecurity • Cloud Security
Top 25% payVerified live · 10 hours ago 3 months ago

Sr. Machine Learning Engineer

$190k – $220k per year • Senior • 5+ years expIn office (Sunnyvale, United States) • Full-Time • Senior
Go
Java
Python
Databases
Apache Kafka
AI/ML
AI Agents
AutoGen
CrewAI
Fine-tuning
Flink
Langfuse
Prompt Engineering
RAG
Spark
Model Context Protocol
LLM
TensorRT
TensorRT-LLM
vLLM
DevOps
AWS
Azure
GCP
Kubernetes
AIOps
Terraform
Vector
Cybersecurity
Zero Trust
Apply
Report

T-Systems Iberia

t-systems.es
T-Systems Iberia is the Spanish and Portuguese division of T-Systems, the global enterprise IT services and digital transformation arm of German telecommunications giant Deutsche Telekom. Headquartered in Barcelona, the firm serves major public sector institutions and enterprise clients across Spain and Southern Europe.
t-systems.es • HQ: Barcelona, Spain • IT Consulting • Information Technology • IT Outsourcing • 5000+ employees
HQ: Barcelona, Spain • IT Consulting • Information Technology • IT Outsourcing • 5000+ employees
Verified live · 1 day ago 3 months ago

Lead Senior Backend Engineer (m/f/d)

$50k – $133k per year (Estimated) • Lead • 5+ years expRemote/Hybrid (Granada, Spain) • Bachelor's Degree • Full-Time • Senior
C++
Go
Java
Python
Rust
SQL
Python
Django
FastAPI
Flask
AI/ML
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
gRPC
Kubernetes
Rest API
WebSockets
API Gateway
GitOps
Grafana
OpenTelemetry
Prometheus
Vector
Apply
Report

Anyscale

anyscale.com
Anyscale is an artificial intelligence infrastructure company headquartered in San Francisco, California, and founded in 2019 by the creators of Ray at the UC Berkeley RISELab. The company offers a managed platform for running Ray, the open source framework used to scale model training, batch inference, and reinforcement learning across clusters. It sells to machine learning engineering teams that need to move distributed AI workloads from laptops to production without rebuilding their stack.
anyscale.com • HQ: San Francisco, United States • Hardware • Information Technology • MLOps • Artificial Intelligence • AI Infrastructure • Cloud Computing • Machine Learning • 201-500 employees • Est. 2019
HQ: San Francisco, United States • Hardware • Information Technology • MLOps • Artificial Intelligence • AI Infrastructure • Cloud Computing • Machine Learning • 201-500 employees • Est. 2019
Top 25% payVerified live · 7 hours ago 4 months ago

Distributed LLM Inference Engineer

$170k – $245k per year • Equity • In office (San Francisco, Palo Alto, United States) • Full-Time
LLM
PyTorch
Ray
vLLM
OpenAI
CUDA Toolkit
TensorFlow
TensorRT
TensorRT-LLM
CUDA
Triton
Apply
Report

Fathom

fathom.ai
Fathom captures, transcribes, and summarizes Zoom, Google Meet, and Microsoft Teams calls. Free for individuals, with AI-powered CRM updates for teams.
fathom.ai • Productivity Software • LLM & Generative AI • Artificial Intelligence • Est. 2020
Productivity Software • LLM & Generative AI • Artificial Intelligence • Est. 2020
Open 117 daysVerified live · 10 hours ago 4 months ago

AI Engineer - Model Performance

$138k – $302k per year (Estimated) • Remote (United States) • Master's Degree • Full-Time
Python
AI/ML
Axolotl
CUDA Toolkit
Fine-tuning
LLM
LoRA
Multimodal AI
Prompt Engineering
QLoRA
Quantization
Ray
Ray Serve
SGLang
TensorRT
TensorRT-LLM
torchtune
vLLM
PEFT
PyTorch
CUDA
DPO
SFT
Speech Recognition
DevOps
GitHub
Management
Notion
Slack
Marketing
HubSpot
Apply
Report

Binance

binance.com
Binance is the world's largest cryptocurrency exchange by trading volume, founded in 2017 by Changpeng Zhao and Yi He. The platform offers spot, margin and derivatives trading across hundreds of digital assets, alongside staking, savings products, payments, an institutional custody arm and a self-custodial Web3 wallet. The group also created BNB Chain, one of the most used smart contract networks, and now operates under a licensed regional structure after a 2023 settlement with United States authorities that installed new leadership and compliance oversight.
binance.com • HQ: Abu Dhabi, United Arab Emirates • Mining & Staking • Crypto Payments • Web3 & dApps • Blockchain & Crypto • Cryptocurrencies • Crypto Exchanges • 1001-5000 employees • Est. 2017
HQ: Abu Dhabi, United Arab Emirates • Mining & Staking • Crypto Payments • Web3 & dApps • Blockchain & Crypto • Cryptocurrencies • Crypto Exchanges • 1001-5000 employees • Est. 2017
Verified live · 2 hours ago 5 months ago

Pioneer Talent Program - Applied Data Scientist

Remote/Hybrid • Bachelor's Degree
Python
AI/ML
AI Agents
Chain-of-Thought
Hallucination
LLM
Prompt Engineering
TensorRT
TensorRT-LLM
vLLM
Function Calling
Model Context Protocol
Post-training
Apply
Report

Zencore

zencore.dev
Zencore is a leading cloud consulting and services firm. Above all else we value experience, hard work and the ability to keep things in perspective and have fun as we ensure our customers are delighted with the work we do.
zencore.dev • London • Mexico City • New York • Copenhagen • Berlin • Blockchain & Crypto • Data & Analytics • Information Technology • Est. 2013
London • Mexico City • New York • Copenhagen • Berlin • Blockchain & Crypto • Data & Analytics • Information Technology • Est. 2013
Open 136 days 5 months ago

Principal Architect, AI/ML

$139k – $249k per year (Estimated) • Lead • Remote (United States) • Master's Degree • Full-Time • Architect
AI Agents
Claude
Fine-tuning
Gemini
Google ADK
JAX
LangChain
Langfuse
LangGraph
LangSmith
Llama
LLM
LoRA
Mistral
PyTorch
Quantization
TensorRT
TensorRT-LLM
Vertex AI
vLLM
PEFT
LLM Guardrails
TPU
DevOps
AWS
Azure
GCP
Google GKE
Kubernetes
Cybersecurity
GDPR
Apply
Open 201 day 7 months ago

Principal Architect, AI/ML

$80k – $190k per year (Estimated) • Lead • Remote (United Kingdom) • Master's Degree • Full-Time • Architect
AI Agents
Claude
Fine-tuning
Gemini
Google ADK
JAX
LangChain
Langfuse
LangGraph
LangSmith
Llama
LLM
LoRA
Mistral
PyTorch
Quantization
TensorRT
TensorRT-LLM
Vertex AI
vLLM
PEFT
LLM Guardrails
TPU
DevOps
AWS
Azure
GCP
Google GKE
Kubernetes
Cybersecurity
GDPR
Apply
Report

H Company

hcompany.ai
H Company is an artificial intelligence research company headquartered in Paris, France, and founded in 2023 by Charles Kantor with a founding team drawn largely from Google DeepMind. The company develops agentic AI systems, including Runner H, an agent that carries out multi-step tasks in a browser on the user's behalf. It raised one of the largest seed rounds in European history and positions itself as a French alternative in the agent tooling market.
hcompany.ai • HQ: Paris, France • Business Process Automation (BPA) • Software • Multimodal AI • LLM & Generative AI • AI Agents • Artificial Intelligence • 201-500 employees • Est. 2023
HQ: Paris, France • Business Process Automation (BPA) • Software • Multimodal AI • LLM & Generative AI • AI Agents • Artificial Intelligence • 201-500 employees • Est. 2023
Open 138 daysVerified live · 1 day ago 5 months ago

Research Engineer, Model Inference & Serving - Paris

$77k – $183k per year (Estimated)Remote/Hybrid (Paris, France) • Full-Time
C++
Python
Rust
C++
PyTorch C++
AI/ML
AI Agents
JAX
Multimodal AI
PyTorch
SGLang
vLLM
Edge AI
CUDA Toolkit
llama.cpp
LLM
MLX ML
ONNX
Quantization
TensorRT
TensorRT-LLM
CUDA
Triton
DevOps
Kubernetes
Apply
Open 142 daysVerified live · 1 day ago 5 months ago

Research Engineer, Model Inference & Serving - London

$104k – $219k per year (Estimated)Remote/Hybrid (London, United Kingdom) • Full-Time
C++
Python
Rust
C++
PyTorch C++
AI/ML
AI Agents
JAX
Multimodal AI
PyTorch
SGLang
vLLM
Edge AI
CUDA Toolkit
llama.cpp
LLM
MLX ML
ONNX
Quantization
TensorRT
TensorRT-LLM
CUDA
Triton
DevOps
Kubernetes
Apply
Report

Liquid AI

liquid.ai
Liquid AI is an artificial intelligence company headquartered in Boston, Massachusetts, and founded in 2023 as a spin-off from the MIT Computer Science and Artificial Intelligence Laboratory. The company builds Liquid Foundation Models, an architecture derived from liquid neural networks that aims to match transformer quality at a fraction of the memory and compute. It targets on-device and edge deployment where models must run on phones, vehicles, and embedded hardware rather than in a data center.
liquid.ai • HQ: Boston, United States • AI Infrastructure • Hardware • Machine Learning • Edge AI • LLM & Generative AI • Artificial Intelligence • Est. 2023
HQ: Boston, United States • AI Infrastructure • Hardware • Machine Learning • Edge AI • LLM & Generative AI • Artificial Intelligence • Est. 2023
Open 139 daysVerified live · 2 hours ago 5 months ago

Solutions Architect

$183k – $335k per year (Estimated) • Staff+ • Remote/Hybrid (San Francisco, Boston, United States) • Full-Time • Architect
Fine-tuning
Jupyter Notebook
AWQ
GGUF
llama.cpp
LLM
Quantization
TensorRT
TensorRT-LLM
vLLM
Apply
Report

Integrant

integrant.com
Extend your team and scale your solutions with Integrant. We provide custom software development, AI solutions, and domain expertise for regulated industries.
integrant.com • HQ: Cairo, Egypt • Artificial Intelligence • Information Technology • Software • Est. 1992
HQ: Cairo, Egypt • Artificial Intelligence • Information Technology • Software • Est. 1992
Verified live · 1 day ago 5 months ago

Lead AI Platform

Lead • 8+ years expIn office (Cairo, Egypt) • Bachelor's Degree • Full-Time • Staff • English (Optional)
Python
AI/ML
CUDA Toolkit
LLM
MLFlow
ONNX
PyTorch
Quantization
TensorRT
TensorRT-LLM
Triton Inference Server
CUDA
Triton
cuDNN
InfiniBand
NCCL
NVLink
Weights & Biases
NVIDIA NIM
Megatron-LM
NVIDIA NeMo
DevOps
Kubernetes
SLURM
HPC
CI/CD
Apply
Report

MLabs

mlabs.city
MLabs is a global software engineering consultancy specializing in high-assurance software engineering, formal verification, and mission-critical systems. Founded in 2018, the company focuses on functional programming languages like Haskell and Rust to build secure, high-performance infrastructure for blockchain, Web3, fintech, and AI projects.
mlabs.city • HQ: London, United Kingdom • Information Technology • Blockchain & Crypto • Software
HQ: London, United Kingdom • Information Technology • Blockchain & Crypto • Software
Verified live · 1 day ago 6 months ago

Staff Software Engineer - Backend & AI Infra

$228k – $379k per year (Estimated) • Equity • Staff+ • Remote (United States) • Staff
Go
Node JS
Python
TypeScript
JavaScript
Databases
ClickHouse
Google BigQuery
PostgreSQL
Redis
TimescaleDB
AI/ML
LLM
Model Context Protocol
Anthropic
TensorRT
TensorRT-LLM
vLLM
AI Agents
TGI
DevOps
Amazon EKS
AWS
CI/CD
Kubernetes
WebSockets
Apply
Report

Parspec

parspec.io
Parspec is a technology company that leverages AI to help sales agents and distributors by simplifying the process of discovering and sourcing the best available construction products and materials.
parspec.io • Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Open 174 daysVerified live · 1 day ago 6 months ago

AI Ops Engineer

$30k – $75k per year (Estimated) • Senior • 5+ years expIn office (Bengaluru, India) • Bachelor's Degree • Full-Time • Senior
Python
Python
Asyncio
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
Weaviate
PostgreSQL
Qdrant
AI/ML
AWS Bedrock
Kubeflow
LiteLLM
LLM
MLFlow
Portkey
Ray
vLLM
AWQ
CUDA
CUDA Toolkit
Embeddings
GPTQ
Hallucination
Langfuse
LoRA
Prompt Engineering
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
PEFT
Amazon SageMaker
AWS Trainium
LLM Guardrails
LLMOps
NCCL
NVLink
TGI
Multimodal AI
AI Agents
DevOps
AIOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
GitHub Actions
Grafana
Kubernetes
OpenTelemetry
Prometheus
Terraform
Vector
Karpenter
Platform Engineering
Amazon EventBridge
Amazon S3
API Gateway
IAM
AWS Step Functions
GitHub
Cybersecurity
Least Privilege
Apply
Report

Elastix

elastix.ai
Elastix AI delivers scalable, energy-efficient AI inference through machine learning, system software, and reconfigurable hardware.
elastix.ai • Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Open 193 daysVerified live · 1 hour ago 7 months ago

AI Software Engineer

$138k – $259k per year (Estimated) • Equity • Middle • 3+ years expIn office (Seattle, United States) • Bachelor's Degree • Full-Time • Middle
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
CUDA
DevOps
Docker
Kubernetes
Apply
Report

Inference

inference.ai
Models. Agents. GPUs. Whatever you need — we've got it. Ghost agent VMs, Maestro model routing, Engine wholesale GPUs, and Academy, on one platform.
inference.ai • San Francisco
San Francisco
Open 221 dayTop 25% pay 8 months ago

Senior Software Engineer - Model Performance

$220k – $320k per year • Senior • 2+ years expIn office (San Francisco, United States) • Full-Time • Senior
C++
Python
C#
C++
PyTorch C++
C#
.NET
AI/ML
CUDA Toolkit
Knowledge Distillation
LLM
LoRA
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
CUDA
GPT-5
Embeddings
DevOps
Docker
Kubernetes
GitHub
Apply
Report

Cohere

cohere.com
Cohere is a Canadian AI company founded in 2019 and headquartered in Toronto. It builds secure, enterprise-focused large language models and AI tools for businesses and regulated industries. The company is known for emphasizing privacy, private deployment, and "sovereign AI" rather than consumer chat products
cohere.com • HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
Verified live · 3 hours ago 9 months ago

Senior ML Systems Engineer, Frameworks & Tooling

$98k – $200k per year (Estimated) • Senior • LondonNew YorkParisSan FranciscoToronto • Remote (United States, France, United Kingdom, Canada) • Full-Time • Senior
Cohere SDK
CUDA Toolkit
JAX
LLM
Ray
CUDA
FSDP
NCCL
DeepSpeed
PyTorch
TensorRT
TensorRT-LLM
vLLM
xFormers
Megatron-LM
DevOps
Docker
Kubernetes
SLURM
HPC
Apply
Report

Specter

specter.com
Specter tracks growth signals across millions of private companies to help investors find opportunities early. Founded in 2021, it combines web, hiring, app and social data into company momentum scores. Venture funds and corporate development teams use it for sourcing.
specter.com • HQ: Singapore • Est. 2021
HQ: Singapore • Est. 2021
Open 331 dayVerified live · 9 hours ago 11 months ago

Software Engineer - ML Infrastructure

$158k – $346k per year (Estimated)In office (San Francisco, United States) • Full-Time
C++
Python
Rust
C++
PyTorch C++
TensorFlow C++
Databases
LanceDB
Qdrant
AI/ML
AI Agents
Computer Vision
LLM
Multimodal AI
PyTorch
Ray
SGLang
Spark
TensorFlow
TensorRT
TensorRT-LLM
VLM
Robotics
Sensor Fusion
Analytics
A/B Testing
Apply
Open 331 dayVerified live · 9 hours ago 11 months ago

ML Research Engineer

$165k – $319k per year (Estimated) • Senior • 5+ years expIn office (San Francisco, United States) • Full-Time • Senior
C++
Rust
C++
PyTorch C++
AI/ML
AI Agents
Computer Vision
CUDA
CUDA Toolkit
Fine-tuning
LLM
Multimodal AI
ONNX
PyTorch
Quantization
RAG
TensorRT
TensorRT-LLM
Apply
Report

OpenAI

openai.com
OpenAI is an American artificial intelligence research and deployment company founded in 2015 with the mission of ensuring that artificial general intelligence benefits all of humanity. It develops the GPT family of large language models and turns them into consumer and developer products, including the ChatGPT assistant, the Sora video model, the Codex coding agent and a commercial API used by millions of developers. Structured as a public benefit corporation controlled by a non-profit foundation, the company is backed by Microsoft and SoftBank and operates from San Francisco.
openai.com • HQ: San Francisco, United States • Multimodal AI • AI Agents • Artificial Intelligence • Machine Learning • LLM & Generative AI • 1001-5000 employees • Est. 2015
HQ: San Francisco, United States • Multimodal AI • AI Agents • Artificial Intelligence • Machine Learning • LLM & Generative AI • 1001-5000 employees • Est. 2015
Open 466 daysTop 25% pay 1 year ago

Software Engineer, Inference - Multi Modal

$295k – $555k per year • In office (San Francisco, United States) • Full-Time
LLM
Multimodal AI
TensorRT
TensorRT-LLM
vLLM
Whisper
OpenAI
Apply
Report
TensorRT-LLM Jobs - Remote & On-site
Frequently asked questions
TensorRT-LLM: How many jobs are available now?
There are 142 TensorRT-LLM jobs listed on Alion right now; listings are refreshed daily.
TensorRT-LLM: What is the typical salary?
The average listed base salary for TensorRT-LLM roles on Alion is about $289,947 USD, based on current listings.
TensorRT-LLM: Are remote, relocation, or visa sponsorship options available?
Many TensorRT-LLM roles include remote or hybrid flexibility, and some companies may offer relocation or visa support-review each posting for details.
TensorRT-LLM: How do I apply on Alion?
Click a TensorRT-LLM job to view the employer's application link, apply through that link or via Alion's apply flow, and sign in to track your applications.