Updated Aug 31, 2026

TensorRT-LLM Jobs - Remote & On-site

TensorRT-LLM jobs on Alion link machine learning engineers and infrastructure builders with vetted companies working on high-performance LLM inference; listings are updated daily and often include remote or hybrid roles. Browse and apply today.

Open positions
139
Companies
80
Salary range
$200K – $360K
Average salary
$290K
AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

Mirantis

mirantis.com
Mirantis is a B2B open-source cloud computing, container management, and AI infrastructure company headquartered in Campbell, California. Originally known as a core contributor to OpenStack, Mirantis now provides open-cloud software and managed services centered on Kubernetes, multi-cloud platforms, and enterprise AI workloads.
mirantis.com • Calgary • Warsaw • Brussels • Prague • Poznań • DevOps • Information Technology • Cloud Computing • 1001-5000 employees
Calgary • Warsaw • Brussels • Prague • Poznań • DevOps • Information Technology • Cloud Computing • 1001-5000 employees
Verified live · 1 hour ago 3 days ago

Director, Presales Solution Architecture - NeoCloud

$185k – $352k per year (Estimated) • Executive • Remote/Hybrid (United States) • Full-Time • Senior • English
CUDA
CUDA Toolkit
Fine-tuning
LLM
PyTorch
TensorRT
TensorRT-LLM
Triton
DevOps
Kubernetes
KubeVirt
Platform Engineering
SLURM
AWS
Apply
Verified live · 1 hour ago 10 days ago

Senior Manager, Sales Engineering — AI / GPU Cloud (NeoCloud)

$162k – $314k per year (Estimated) • Senior • Remote/Hybrid (San Jose, United States) • Full-Time • Senior
CUDA
CUDA Toolkit
Fine-tuning
LLM
PyTorch
TensorRT
TensorRT-LLM
Triton
InfiniBand
NCCL
NVIDIA NeMo
NVLink
DevOps
Kubernetes
KubeVirt
Platform Engineering
SLURM
HPC
AWS
Apply
Open 95 daysVerified live · 1 hour ago 3 months ago

Product Manager - AI Inference & Model Serving

$145k – $248k per year (Estimated) • Equity • Senior • 7+ years expAustin • Remote (United States) • Full-Time • Senior
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Triton
DevOps
AWS
Kubernetes
Platform Engineering
Chips/EDA
PoC Library
Apply
Report

JPMorganChase

jpmorganchase.com
JPMorgan Chase & Co. is a leading global financial services firm and the largest banking institution in the United States by assets. Headquartered in New York City, the company offers a comprehensive range of financial solutions, including investment banking, asset management, treasury services, and commercial banking. Through its widely recognized consumer division, Chase, it delivers retail banking, credit card, and mortgage services to tens of millions of households across the globe.
jpmorganchase.com • HQ: London, United Kingdom • Blockchain & Crypto • Clinical Research • Financial Services • 1001-5000 employees
HQ: London, United Kingdom • Blockchain & Crypto • Clinical Research • Financial Services • 1001-5000 employees
Verified live · 29 min ago 1 day ago

Lead Software Engineer - Python / Go & AI/ML

Lead • In office • PhD • Staff
Python
AI/ML
AWQ
GPTQ
LLM
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
DevOps
AWS
Chaos Engineering
Kubernetes
Apply
Verified live · 29 min ago 5 days ago

Principal Software Engineer - LLM Optimization

$160k – $343k per year (Estimated) • Lead • In office (Jersey City, United States) • PhD • Principal
AWQ
GPTQ
LLM
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
Human-in-the-Loop
LLM Guardrails
AI Agents
DevOps
Amazon EKS
AWS
Chaos Engineering
Kubernetes
Apply
Report

Baseten

baseten.co
Baseten is an AI infrastructure platform designed to help developers and machine learning teams deploy, serve, and scale open-source and custom AI models. The platform provides performant, low-latency inference infrastructure alongside developer tools like Truss, an open-source model packaging framework. Headquartered in San Francisco, California, Baseten enables companies to run state-of-the-art models in production seamlessly without managing underlying cloud infrastructure.
baseten.co • HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
Top 25% payVerified live · 3 hours ago 2 days ago

Technical Program Manager, Model Performance

$165k – $330k per year • Remote/Hybrid (San Francisco, New York, United States, Toronto, Montreal, Canada) • Full-Time
Cursor
DeepSeek
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Management
Notion
Apply
Top 25% payVerified live · 3 hours ago 13 days ago

Forward Deployed Engineer (Training)

$200k – $400k per year • Junior • 1+ year expRemote/Hybrid (San Francisco, New York, United States) • Full-Time • Junior
Cursor
JAX
LLM
PyTorch
Ray
SGLang
TensorRT
TensorRT-LLM
vLLM
InfiniBand
Post-training
SFT
DevOps
Kubernetes
SLURM
Chips/EDA
PoC Library
Apply
Top 25% payVerified live · 3 hours ago 2 months ago

Software Engineer - Baseten Inference Stack

$180k – $360k per year • Remote/Hybrid (San Francisco, New York, United States, Toronto, Montreal, Canada) • Bachelor's Degree • Full-Time
Cursor
LLM
Function Calling
SGLang
TensorRT
TensorRT-LLM
vLLM
DevOps
Kubernetes
Platform Engineering
CI/CD
Management
Notion
Apply
Report

Orcrist Technologies

orcrist.org
Orcrist Technologies offers pioneering AI and data analytics solutions in the private and public sectors, turning sensors into strategy. We are a Berlin-based data defense technology company building AI-powered software for real-time situational awareness and sensor fusion. Our mission is to give decision-makers the clarity they need-when it matters most.
orcrist.org • Berlin • Dresden • Defense AI • Data & Analytics • Artificial Intelligence
Berlin • Dresden • Defense AI • Data & Analytics • Artificial Intelligence
Verified live · 5 hours ago 2 months ago

Infrastructure/Systems Engineer

$69k – $139k per year (Estimated) • Senior • 5+ years expRemote/Hybrid (Berlin, Germany) • Internship • Senior • German: B1
Bash
Python
AI/ML
KServe
vLLM
Triton
InfiniBand
CUDA Toolkit
LLM
Quantization
TensorRT
TensorRT-LLM
CUDA
NCCL
NVLink
DevOps
Ansible
Kubernetes
Terraform
Ubuntu
CI/CD
Git
Red Hat
Cybersecurity
ISO 27001
Apply
Report

Beacon AI

beaconai.com
Beacon AI is an aviation technology company that builds AI-powered pilot assistance systems and flight safety software for commercial, private, and defense aviation. Often described as an "R2-D2 for pilots", the platform acts as an intelligent onboard teammate - integrating directly with cockpit infrastructure to assist aviators with real-time checklists, route optimization, fleet monitoring, and situational awareness to reduce human error and workload.
beaconai.com • HQ: San Francisco, United States • Artificial Intelligence • Space & Aerospace • Avionics & Flight Control
HQ: San Francisco, United States • Artificial Intelligence • Space & Aerospace • Avionics & Flight Control
Top 25% payVerified live · 1 day ago 4 days ago

Software Engineer, Cloud Infrastructure (Multiple Seniority Levels)

$135k – $260k per year • Middle • 4+ years expIn office • Visa sponsorship • Full-Time • Middle
Python
Databases
Amazon Aurora
DynamoDB
OpenSearch
pgvector
Pinecone
PostgreSQL
Redis
AI/ML
LLM
AWS Bedrock
Embeddings
Function Calling
LangChain
Quantization
RAG
TensorRT
TensorRT-LLM
Time Series Forecasting
Triton
Triton Inference Server
Amazon SageMaker
Human-in-the-Loop
LLM Evaluation
LLM Guardrails
DevOps
AWS
Kubernetes
Amazon EKS
AWS CDK
AWS Lambda
CI/CD
GitHub Actions
OpenTelemetry
Terraform
Vector
Amazon CloudWatch
Amazon ECS
Amazon EventBridge
Amazon S3
AWS Step Functions
GitHub
IAM
Cybersecurity
Least Privilege
Analytics
A/B Testing
Apply
Top 25% payVerified live · 1 day ago 1 month ago

Software Engineer, Artificial Intelligence/LLM (Multiple Seniority Levels)

$135k – $260k per year • In office • Visa sponsorship • Full-Time
Python
TypeScript
Databases
Amazon Aurora
DynamoDB
OpenSearch
pgvector
Pinecone
PostgreSQL
Weaviate
AI/ML
AWS Bedrock
Embeddings
Function Calling
Hallucination
LangChain
LLM
RAG
Time Series Forecasting
Anthropic
Human-in-the-Loop
LLM Guardrails
OpenAI
Multimodal AI
TensorRT
TensorRT-LLM
Triton
DevOps
AWS
Vector
Amazon S3
CI/CD
Analytics
A/B Testing
Apply
Report

Reducto

reducto.ai
Reducto is a San Francisco company founded in 2023 that converts complex documents into structured input for language models. Its pipeline handles tables, charts, forms and scanned pages with vision models, producing accurate chunks for retrieval systems where generic parsers fail. It is used by enterprises building document-grounded AI in finance, healthcare and insurance.
reducto.ai • HQ: San Francisco, United States • Information Security • LLM & Generative AI • Cybersecurity • 51-200 employees • Est. 2023
HQ: San Francisco, United States • Information Security • LLM & Generative AI • Cybersecurity • 51-200 employees • Est. 2023
Top 25% payVerified live · 6 min ago 5 days ago

LLM/ML Engineer (Inference)

$200k – $300k per year • Equity 0.1–1% • Middle • 3+ years expIn office (San Francisco, United States) • Full-Time • Middle
Python
AI/ML
LLM
PyTorch
TensorRT
TensorRT-LLM
vLLM
TGI
CUDA
CUDA Toolkit
Triton
Apply
Report

Prime Intellect

primeintellect.com
Prime Intellect is an artificial intelligence infrastructure company headquartered in San Francisco, California, and founded in 2023. The company provides a decentralized platform for training, evaluating, and deploying large-scale AI models, featuring tools for reinforcement learning, agent development, and a global compute marketplace. It operates globally by aggregating computing resources from various providers to enable researchers and developers to build open-source models and autonomous agents.
primeintellect.com • HQ: San Francisco, United States • Cloud Computing • Information Technology • LLM & Generative AI • Reinforcement Learning • AI Agents • Artificial Intelligence • Hardware • AI Infrastructure • Est. 2023
HQ: San Francisco, United States • Cloud Computing • Information Technology • LLM & Generative AI • Reinforcement Learning • AI Agents • Artificial Intelligence • Hardware • AI Infrastructure • Est. 2023
Top 25% payVerified live · 2 days ago 1 month ago

Member of Technical Staff - Training Platform

$150k – $300k per year • Staff+ • In office (San Francisco, United States) • Relocation • Visa sponsorship • Full-Time • Staff
Python
TypeScript
JavaScript
Python
FastAPI
SQLAlchemy
Databases
Databricks
AI/ML
Fine-tuning
LangChain
LLM
LoRA
OpenRouter
Perplexity
QLoRA
RLHF
SGLang
TensorRT
TensorRT-LLM
Together AI
vLLM
PEFT
NCCL
NVLink
OpenAI
Post-training
SFT
Function Calling
Frontend
Next.js
React.js
shadcn/ui
Tailwind CSS
tRPC
Radix UI
DevOps
Ansible
Cloudflare
Datadog
GCP
GitOps
Google Cloud Run
Google GKE
Grafana
Helm
KEDA
kubectl
Kubernetes
Loki
OpenTelemetry
Prometheus
Rest API
Terraform
Management
Zapier
Apply
Top 25% payVerified live · 2 days ago 1 month ago

Member of Technical Staff - Inference

$150k – $300k per year • Equity • Staff+ • Remote/Hybrid (San Francisco, United States) • Relocation • Visa sponsorship • PhD • Full-Time • Staff
Python
C++
Rust
C++
PyTorch C++
Protobuf
Databases
Databricks
Apache Kafka
Redis
AI/ML
CUDA
CUDA Toolkit
LangChain
LLM
OpenRouter
Perplexity
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Together AI
vLLM
InfiniBand
NCCL
OpenAI
Post-training
SFT
Function Calling
Triton
DevOps
AWS
CI/CD
Cloudflare
Datadog
GCP
Kubernetes
Service Mesh
SLI/SLO/SLA
Ansible
Grafana
gRPC
OpenTelemetry
Prometheus
Terraform
Management
Zapier
Apply
Report

xAI

x.ai
xAI is an American artificial intelligence company founded by Elon Musk in 2023 with the stated goal of building models that help humans understand the universe. It develops the Grok family of large language models, distributes them through a consumer assistant, a developer API and deep integration with the X social platform, and adds image and video generation through Grok Imagine. The company runs its own Colossus supercomputer clusters in Memphis, Tennessee, is headquartered in Palo Alto, California, and merged with X Corp in 2025 to combine model development with a large consumer distribution channel.
x.ai • HQ: Palo Alto, United States • Multimodal AI • AI Agents • Machine Learning • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2023
HQ: Palo Alto, United States • Multimodal AI • AI Agents • Machine Learning • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2023
Open 695 daysTop 25% pay 2 years ago

Software Engineer - Training/Inference (C++)

$440k per year • In office (Palo Alto, United States)
C++
Rust
AI/ML
Grok
Knowledge Distillation
LLM
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
Triton
DevOps
CI/CD
Apply
Report

ServiceNow

servicenow.com
ServiceNow is a leading cloud-based enterprise platform that connects data, workflows, and artificial intelligence to automate business operations. Originally focused on IT service management, the platform now streamlines processes across human resources, customer service, security, and supply chain management within a single digital ecosystem. Through its built-in low-code tools and AI capabilities, it helps global organizations optimize workforce productivity, eliminate manual tasks, and scale modern enterprise workflows.
servicenow.com • HQ: Santa Clara, United States • Financial Services • Data & Analytics • Artificial Intelligence • Software • Information Technology • Cloud Computing • IT Management • Business Process Automation (BPA) • 1001-5000 employees • Est. 2004
HQ: Santa Clara, United States • Financial Services • Data & Analytics • Artificial Intelligence • Software • Information Technology • Cloud Computing • IT Management • Business Process Automation (BPA) • 1001-5000 employees • Est. 2004
$213k – $430k per year (Estimated) • Staff+ • 10+ years expIn office (Mountain View, United States) • Bachelor's Degree • Full-Time • Staff
C++
Python
C++
PyTorch C++
AI/ML
LLM
NLP
PyTorch
TensorRT
TensorRT-LLM
vLLM
Hugging Face
AI Agents
Analytics
ETL/ELT
Management
ServiceNow
Apply
Top 25% payVerified live · 2 hours ago 1 month ago

Senior Machine Learning Engineer, Agentic Systems - Moveworks

$161k – $274k per year • Equity • Senior • 5+ years expIn office (Mountain View, United States) • Bachelor's Degree • Full-Time • Senior
C++
Python
C++
PyTorch C++
AI/ML
LLM
NLP
PyTorch
TensorRT
TensorRT-LLM
vLLM
Hugging Face
AI Agents
Analytics
ETL/ELT
Management
ServiceNow
Apply
Open 131 dayVerified live · 2 hours ago 4 months ago

Senior Machine Learning Engineer, Agentic Systems - Moveworks

$183k – $354k per year (Estimated) • Senior • 5+ years expIn office (Mountain View, United States) • Bachelor's Degree • Full-Time • Senior • English
C++
Python
C++
PyTorch C++
AI/ML
LLM
NLP
PyTorch
TensorRT
TensorRT-LLM
vLLM
Hugging Face
AI Agents
Analytics
ETL/ELT
Management
ServiceNow
Apply
Report

DoorDash

doordash.com
DoorDash is an American technology company headquartered in San Francisco that operates a leading global local commerce and on-demand delivery platform. The company connects consumers with local merchants, offering last-mile logistics for restaurant meals, groceries, convenience goods, and retail items delivered by independent couriers. Beyond its core consumer marketplace, DoorDash provides digital ordering software, merchant advertising tools, and white-label fulfillment services for businesses across multiple countries.
doordash.com • HQ: San Francisco, United States • Commerce • Marketplaces • Transportation & Logistics • Delivery
HQ: San Francisco, United States • Commerce • Marketplaces • Transportation & Logistics • Delivery
$4k – $14k per year • Equity • Senior • 6+ years expIn office (San Francisco, United States) • Bachelor's Degree • Senior
Python
AI/ML
Claude
Claude Code
Cursor
DeepSeek
Embeddings
Fine-tuning
LLM
LoRA
Qwen
Reinforcement Learning
RLHF
PEFT
DPO
LLM Guardrails
OpenAI Codex
Post-training
SFT
AI Agents
AWQ
GPTQ
Quantization
RAG
SGLang
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
DevOps
AWS
GCP
Kubernetes
Vector
Apply
Report

Cloudflare

cloudflare.com
Cloudflare is a global web infrastructure and cybersecurity company that provides content delivery network (CDN) services, DDoS mitigation, and edge computing solutions. The platform acts as a protective shield and performance booster between website visitors and origin servers, safeguarding online applications from cyber threats while optimizing speed. Headquartered in San Francisco, it secures and accelerates millions of internet properties and handles a significant portion of all global web traffic.
cloudflare.com • HQ: San Francisco, United States • Information Technology • Cloud Computing • Email Security • Network Security • Cybersecurity • CDN • 1001-5000 employees • Est. 2009
HQ: San Francisco, United States • Information Technology • Cloud Computing • Email Security • Network Security • Cybersecurity • CDN • 1001-5000 employees • Est. 2009
Verified live · 3 hours ago 1 month ago

Senior Machine Learning Engineer

$170k – $329k per year (Estimated) • Senior • Remote/Hybrid • Internship • Senior
Python
AI/ML
Quantization
Embeddings
JAX
llama.cpp
LLM
Multimodal AI
ONNX
PyTorch
SGLang
TensorFlow
TensorRT
TensorRT-LLM
vLLM
Triton
Galileo
RAG
DevOps
Cloudflare
Apply
Report

RapidClaims

rapidclaims.ai
RapidClaims is a revenue cycle management company leveraging advanced AI and LLMs to enable autonomous medical coding and workflow automation for modernizing and scaling healthcare billing operations.
rapidclaims.ai • Bengaluru • Billing & Invoicing • Est. 2023
Bengaluru • Billing & Invoicing • Est. 2023
$31k – $78k per year (Estimated) • Senior • 5+ years expIn office (Bengaluru, India) • Senior
Python
Databases
ArangoDB
Neo4j
AI/ML
Braintrust
Fine-tuning
Function Calling
Hybrid Search
Langfuse
LangSmith
Llama
LLM
LoRA
Prompt Engineering
PyTorch
QLoRA
Qwen
RAG
Reranking
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
DPO
Hugging Face
Human-in-the-Loop
Knowledge Graph
SFT
Structured Outputs
AI Agents
Arize Phoenix
Model Context Protocol
DevOps
Vector
Apply
Report

Furiosa

furiosa.ai
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.
furiosa.ai • Seoul • Hwaseong • Santa Clara • San Jose • Data Centers • LLM & Generative AI • Artificial Intelligence
Seoul • Hwaseong • Santa Clara • San Jose • Data Centers • LLM & Generative AI • Artificial Intelligence
Verified live · 2 hours ago 10 days ago

Solutions Architect - US

$158k – $288k per year (Estimated) • Staff+ • In office (Santa Clara, United States) • Architect
Python
C++
Rust
C++
PyTorch C++
TensorFlow C++
AI/ML
AutoGen
LangChain
LangGraph
LlamaIndex
LLM
PyTorch
SGLang
TensorFlow
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
vLLM
Model Context Protocol
Quantization
Apply
Verified live · 2 hours ago 10 days ago

Senior Solution Architect

Staff+ • In office (Seoul, South Korea) • Architect
Python
C++
Rust
AI/ML
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
DevOps
Docker
Kubernetes
Chips/EDA
PoC Library
Apply
Verified live · 2 hours ago 12 days ago

Senior Software Engineer, Inference Engine (Platform Software)

Senior • 3+ years expIn office (Seoul, South Korea) • Bachelor's Degree • Senior
C++
Rust
AI/ML
LLM
Multimodal AI
SGLang
vLLM
CUDA
CUDA Toolkit
TensorRT
TensorRT-LLM
Triton
Apply
Report

NVIDIA

nvidia.com
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.
nvidia.com • HQ: Santa Clara, United States • AI Infrastructure • Semiconductors • AI Agents • Big Data • LLM & Generative AI • Machine Learning • Hardware • Data & Analytics • Information Technology • Artificial Intelligence • 5000+ employees • Est. 1993
HQ: Santa Clara, United States • AI Infrastructure • Semiconductors • AI Agents • Big Data • LLM & Generative AI • Machine Learning • Hardware • Data & Analytics • Information Technology • Artificial Intelligence • 5000+ employees • Est. 1993
Top 25% pay 8 days ago

Engineering Manager, Local AI Agents

$224k – $357k per year • Lead • 10+ years expRemote/Hybrid (Santa Clara, Redmond, United States) • Bachelor's Degree • Full-Time • Staff
C++
Python
C++
PyTorch C++
AI/ML
AI Agents
CUDA
CUDA Toolkit
LangChain
llama.cpp
LLM
LocalAI
Ollama
ONNX
Perplexity
PyTorch
Quantization
Spark
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
DevOps
CI/CD
Apply
$168k – $259k per year • Senior • 8+ years expIn office (Santa Clara, Redmond, United States) • Bachelor's Degree • Full-Time • Staff
C++
Python
C++
PyTorch C++
AI/ML
AI Agents
CUDA
CUDA Toolkit
LangChain
llama.cpp
LLM
LocalAI
Ollama
ONNX
Perplexity
PyTorch
Quantization
Spark
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
Apply
$184k – $288k per year • Senior • 5+ years expIn office (Santa Clara, United States) • Master's Degree • Full-Time • Senior
C++
Python
C++
PyTorch C++
AI/ML
LLM
PyTorch
Ray
Reinforcement Learning
RLHF
AI Agents
DPO
FSDP
GRPO
Post-training
PPO
DeepSpeed
DeepSpeed-Chat
Multimodal AI
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
VLM
InfiniBand
Megatron-LM
Mixture of Experts
NCCL
NVIDIA NeMo
NVLink
Pre-training
DevOps
Kubernetes
Apply
Report

Cast AI

cast.ai
CAST AI is an AI-driven cloud automation and Kubernetes cost optimization platform built to help enterprises manage, scale, and secure their cloud infrastructure. Founded in 2019 and headquartered in Miami, Florida, the company automates cloud operations across major hyper-scalers including Amazon Web Services (AWS), Google Cloud Platform (GCP), and Microsoft Azure.
cast.ai • HQ: Miami, United States • Information Technology • Virtualization • DevOps
HQ: Miami, United States • Information Technology • Virtualization • DevOps
Open 382 daysVerified live · 25 min ago 1 year ago

Senior ML Engineer | Kimchi (LLM Inference Optimization)

Equity • Senior • Remote (EU) • Visa sponsorship • Internship • Senior
Python
Databases
ClickHouse
PostgreSQL
AI/ML
CUDA
CUDA Toolkit
LLM
Perplexity
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
TPOT
vLLM
Hugging Face
DevOps
Akamai
ArgoCD
AWS
Azure
GCP
GitLab CI
Grafana
Kubernetes
Loki
Prometheus
GitLab
Apply
Report

Netskope

netskope.com
Netskope provides a security service edge platform delivered from its own global network. Its products cover cloud access security, secure web gateway and zero trust private access. Enterprises use it to protect users working outside the corporate perimeter.
netskope.com • HQ: Santa Clara, United States • Artificial Intelligence • Information Technology • Cybersecurity • 501-1000 employees • Est. 2012
HQ: Santa Clara, United States • Artificial Intelligence • Information Technology • Cybersecurity • 501-1000 employees • Est. 2012
$148k – $316k per year (Estimated) • Equity • Lead • 10+ years expIn office (Santa Clara, United States) • Master's Degree • Principal
C++
Python
AI/ML
AWQ
Fine-tuning
GGUF
GPTQ
llama.cpp
LLM
LoRA
MLX ML
ONNX
QLoRA
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
Edge AI
LLM Evaluation
AI Agents
Claude
Claude Code
OpenAI Codex
Cybersecurity
Zero Trust
Apply
Verified live · 1 hour ago 2 months ago

Machine Learning Engineer, AI Labs

$169k – $343k per year (Estimated) • Equity • Staff+ • 10+ years expIn office (Santa Clara, United States) • Staff
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Edge AI
Cybersecurity
Zero Trust
Apply
Verified live · 1 hour ago 1 month ago

Machine Learning Engineer, AI Labs

Staff+ • 10+ years expIn office (Taipei, Taiwan) • Staff
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Edge AI
Cybersecurity
Zero Trust
Apply
Report

d-Matrix

d-matrix.ai
d-Matrix is a semiconductor technology company based in Santa Clara, California, and was founded in 2019. The company develops specialized AI inference hardware and software platforms, including its Corsair platform and 3DIMC architecture, designed to accelerate generative AI workloads in data centers. It operates as a venture-backed enterprise serving global hyperscalers and enterprises with a focus on reducing latency, power consumption, and total cost of ownership for large language models.
d-matrix.ai • HQ: Santa Clara, United States • Data Centers • Information Technology • LLM & Generative AI • Artificial Intelligence • AI Infrastructure • Hardware • Semiconductors • Est. 2019
HQ: Santa Clara, United States • Data Centers • Information Technology • LLM & Generative AI • Artificial Intelligence • AI Infrastructure • Hardware • Semiconductors • Est. 2019
Top 25% payVerified live · 1 day ago 13 days ago

Senior Staff LLM Inference Engineer

$195k – $285k per year • Staff+ • 10+ years expRemote/Hybrid (Santa Clara, United States) • Bachelor's Degree • Full-Time • Staff
C++
Python
AI/ML
CUDA
CUDA Toolkit
LLM
ONNX
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
vLLM
JAX
Mixture of Experts
Apply
$198k – $370k per year (Estimated) • Staff+ • 10+ years expSanta Clara • Remote (United States) • Visa sponsorship • Bachelor's Degree • Staff
C++
Python
AI/ML
CUDA Toolkit
LLM
ONNX
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
JAX
Mixture of Experts
Apply
Top 25% payVerified live · 1 day ago 1 month ago

Principal System Software Engineer, AI Inference Execution

$195k – $285k per year • Lead • 12+ years expIn office (Santa Clara, United States) • Bachelor's Degree • Full-Time • Principal
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
LLM
NLP
ONNX
PyTorch
Ray
SGLang
TensorFlow
TensorRT
TensorRT-LLM
vLLM
NCCL
DevOps
Kubernetes
Apply
Report

Sarvam AI

sarvam.ai
Sarvam AI is a leading Indian artificial intelligence company focused on building full-stack sovereign generative AI infrastructure, foundational large language models (LLMs), and speech technologies tailored for India’s diverse languages and enterprise requirements.
sarvam.ai • HQ: Bengaluru, India • Natural Language Processing • LLM & Generative AI • Artificial Intelligence
HQ: Bengaluru, India • Natural Language Processing • LLM & Generative AI • Artificial Intelligence
Verified live · 1 hour ago 13 days ago

Senior Backend Engineer, Vision

$26k – $69k per year (Estimated) • Senior • 6+ years expIn office (Bengaluru, India) • Full-Time • Senior
Go
Python
Databases
PostgreSQL
Redis
AI/ML
Gemini
LLM
OCR
Ray
Ray Serve
SGLang
TensorRT
TensorRT-LLM
vLLM
DevOps
Kubernetes
OpenTelemetry
Azure
Apply
Verified live · 1 hour ago 22 days ago

Performance Engineer, Inference

$22k – $66k per year (Estimated) • Junior • 2+ years expIn office (Bengaluru, Chennai, India) • Full-Time • Junior
C++
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
SGLang
TensorRT
TensorRT-LLM
TPOT
vLLM
NCCL
Text-to-Speech
Mixture of Experts
Speech Recognition
DevOps
SLI/SLO/SLA
Apply
Report

Rubrik

rubrik.com
Rubrik is a Security and AI Operations Company focused on data protection and cyber resilience. We empower organizations to enhance their security posture while accelerating their enterprise AI initiatives. Our unique approach ensures comprehensive protection of critical data and facilitates trusted AI deployments at scale.
rubrik.com • Bengaluru • Chicago • New York • London • Mumbai • Peripherals • AI Agents • Government • Information Technology • Cybersecurity • 51-200 employees • Est. 2015
Bengaluru • Chicago • New York • London • Mumbai • Peripherals • AI Agents • Government • Information Technology • Cybersecurity • 51-200 employees • Est. 2015
Verified live · 10 min ago 2 months ago

Senior Machine Learning Engineer

$130k – $271k per year (Estimated) • Senior • 2+ years expIn office (Palo Alto, United States) • Bachelor's Degree • Full-Time • Senior
Python
AI/ML
AI Agents
Fine-tuning
Google ADK
Knowledge Distillation
LLM
LoRA
PyTorch
Quantization
RLHF
SGLang
Synthetic Data
TensorRT
TensorRT-LLM
Vertex AI
vLLM
PEFT
DPO
GRPO
Post-training
SFT
LiteLLM
LLM Guardrails
Model Context Protocol
DevOps
Azure
Apply
Report

Socure

socure.com
Socure is a leading digital identity verification and fraud prevention platform that uses artificial intelligence and machine learning to verify consumer identities in real time. The company analyzes thousands of predictive data points - including biometrics, device intelligence, and email or phone attributes - to accurately authenticate individuals during online onboarding. By providing high accuracy rates across demographic groups, it enables financial institutions, digital healthcare providers, and online platforms to reduce fraud while streamlining user conversion.
socure.com • HQ: New York, United States • Artificial Intelligence • Fraud Detection • Identity Management • 501-1000 employees • Est. 2012
HQ: New York, United States • Artificial Intelligence • Fraud Detection • Identity Management • 501-1000 employees • Est. 2012
Verified live · 1 day ago 17 days ago

Software Engineer II — Agentic AI Foundations

$97k – $126k per year • Middle • 2+ years expRemote/Hybrid (Toronto, Canada) • Bachelor's Degree • Full-Time • Middle
Go
Java
Python
AI/ML
CUDA Toolkit
LLM
Ollama
TensorRT
TensorRT-LLM
Triton Inference Server
vLLM
AI Agents
Function Calling
Human-in-the-Loop
LLM Guardrails
Fine-tuning
Knowledge Graph
Analytics
A/B Testing
Apply
Report

Salary range

Seniority Jobs 25% Median 75%
Middle 8 $190K $230K $275K
Senior 12 $259K $274K $306K
Staff 15 $260K $285K $300K
All levels 54 $260K $288K $330K

Based only on the listings that state pay. Gross annual amounts, converted to USD so roles in different currencies stay comparable.

Fully remote 11%
15
Hybrid 24%
34
On site 65%
90
Visa sponsorship 8%
11
Equity 14%
19

The market right now

Alion currently lists 139 open TensorRT-LLM jobs from 80 companies. 13 of them were posted or refreshed in the last seven days. 66 of the employers have posted something in the last three months, which is the pool worth watching if you are starting a search now. Every listing links straight to the employer, so you apply on their own board rather than through an intermediary.

What these roles pay

54 of these TensorRT-LLM jobs state pay directly. Across them the middle half of the market sits between $260K and $330K a year, with a median of $288K. By level, the median runs middle at $230K, senior at $274K, and staff at $285K. The step from middle to staff is worth about 1.2x on median pay. Figures are gross annual amounts converted to US dollars, so roles in different currencies stay comparable.

Where the work is

The largest concentrations of these roles are Singapore (7), South Korea (6), United States (5), France (3), and Taiwan (2). 11% of the listings are fully remote and a further 24% are hybrid, so a large part of this market is open to you regardless of where you live.

Who is hiring

The employers with the most open TensorRT-LLM jobs right now are NVIDIA (11), Baseten (7), Furiosa (5), Together AI (5), Fireworks AI (4), and Aion (4). Each company page on Alion carries its size, funding stage, tech stack and every other position it has open, so you can judge the employer before you spend an evening on the application.

What employers ask for

Reading across the current listings, the tools that come up most often are TensorRT (139), TensorRT-LLM (139), LLM (139), vLLM (127), Python (103), SGLang (91), PyTorch (76), and Kubernetes (75). The counts are how many of these openings name each one, which is a better guide to what is actually being hired for than a generic skills list.

Relocation, visas and equity

Of the current TensorRT-LLM jobs, 11 state visa sponsorship, 4 offer a relocation package, and 19 include equity. These are filters on the list above, so you can narrow it to the ones that make a move possible for you.

TensorRT-LLM Jobs - Remote & On-site
Frequently asked questions
How many TensorRT-LLM jobs are open right now?
There are 139 open TensorRT-LLM jobs on Alion from 80 companies. The list was last refreshed on August 31, 2026.
What do TensorRT-LLM jobs pay?
The median is $288K a year, and the middle half of the market falls between $260K and $330K. This is based on the 54 listings that state pay.
Are any of these roles remote?
15 of the 139 listings (11%) are fully remote, and you can filter the list down to them in one click.
Which companies are hiring?
The most active employers right now are NVIDIA, Baseten, Furiosa, Together AI, Fireworks AI, and Aion. Each has its own page on Alion with the rest of its open roles.
Can I get visa sponsorship or relocation?
Yes - of the current listings, 11 state visa sponsorship and 4 offer a relocation package. Both are filters on the jobs page.
How do I apply?
Open any listing and apply on the employer's own board through the link on the page. No account is required to browse or to apply.