TensorRT-LLM Jobs - Remote & On-site

TensorRT-LLM jobs on Alion link machine learning engineers and infrastructure builders with vetted companies working on high-performance LLM inference; listings are updated daily and often include remote or hybrid roles. Browse and apply today.

AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

Socure

socure.com
Socure is a leading digital identity verification and fraud prevention platform that uses artificial intelligence and machine learning to verify consumer identities in real time. The company analyzes thousands of predictive data points - including biometrics, device intelligence, and email or phone attributes - to accurately authenticate individuals during online onboarding. By providing high accuracy rates across demographic groups, it enables financial institutions, digital healthcare providers, and online platforms to reduce fraud while streamlining user conversion.
socure.com • HQ: New York, United States • Artificial Intelligence • Fraud Detection • Identity Management • 501-1000 employees • Est. 2012
HQ: New York, United States • Artificial Intelligence • Fraud Detection • Identity Management • 501-1000 employees • Est. 2012
Verified live · 1 hour ago 17 days ago

Software Engineer II — Agentic AI Foundations

$97k – $126k per year • Middle • 2+ years expRemote/Hybrid (Toronto, Canada) • Bachelor's Degree • Full-Time • Middle
Go
Java
Python
AI/ML
CUDA Toolkit
LLM
Ollama
TensorRT
TensorRT-LLM
Triton Inference Server
vLLM
AI Agents
Function Calling
Human-in-the-Loop
LLM Guardrails
Fine-tuning
Knowledge Graph
Analytics
A/B Testing
Apply
Report

Nuancelabs

nuancelabs.ai
Face-to-face AI interaction that feels human. We are building a human foundation model with emotional intelligence — it reads tone, expression, and even hesitation, and responds in real time.
nuancelabs.ai • Seattle • Software • Conversational AI • Artificial Intelligence • 51-200 employees • Est. 2025
Seattle • Software • Conversational AI • Artificial Intelligence • 51-200 employees • Est. 2025
Top 25% payVerified live · 1 day ago 3 months ago

Member of Technical Staff — Model Optimization and Inference (New Grad)

$200k – $300k per year • Junior • In office (Seattle, United States) • Visa sponsorship • Bachelor's Degree • Internship • Staff
Python
AI/ML
Accelerate
AWQ
CUDA
CUDA Toolkit
Diffusion Models
GPTQ
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
vLLM
Apply
Top 25% payVerified live · 1 day ago 3 months ago

Member of Technical Staff — Model Optimization and Inference (Experienced)

$250k – $350k per year • Staff+ • In office (Seattle, United States) • Visa sponsorship • Staff
Python
AI/ML
Accelerate
AWQ
CUDA
CUDA Toolkit
Diffusion Models
GPTQ
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
vLLM
Post-training
Apply
Report

Intel Corporation

intel.com
Intel is an American semiconductor company founded in 1968 by Robert Noyce and Gordon Moore and headquartered in Santa Clara, California. It created the x86 instruction set that still underpins most personal computers and servers, and designs and manufactures processors, chipsets, discrete graphics, networking silicon and AI accelerators. Unlike most of its competitors the company owns its fabrication plants, and it is investing heavily in Intel Foundry to manufacture chips for external customers while rebuilding its process technology leadership.
intel.com • HQ: Santa Clara, United States • Cloud Computing • Semiconductor Manufacturing • Data & Analytics • Artificial Intelligence • Hardware • Computer Components • Semiconductors • 5000+ employees • Est. 1968
HQ: Santa Clara, United States • Cloud Computing • Semiconductor Manufacturing • Data & Analytics • Artificial Intelligence • Hardware • Computer Components • Semiconductors • 5000+ employees • Est. 1968
Verified live · 1 hour ago 18 days ago

Software Engineer — Distributed LLM Inference Systems

Junior • In office (Shanghai, China) • Master's Degree • Full-Time
C++
Python
C++
PyTorch C++
AI/ML
LLM
PyTorch
AI Agents
SGLang
TensorRT
TensorRT-LLM
vLLM
Function Calling
Apply
Verified live · 1 hour ago 2 months ago

Senior Software Developer - Network and Collectives

$91k – $257k per year (Estimated) • Senior • 5+ years expRemote/Hybrid (Haifa, Israel) • Full-Time • Senior
C++
C
C
MPI
AI/ML
NCCL
DeepSeek
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
InfiniBand
DevOps
HPC
Apply
Report

Together AI

together.ai
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.
together.ai • HQ: San Francisco, United States • Hardware • Events & Ticketing • Corporate Events • AI Infrastructure • Artificial Intelligence • 501-1000 employees • Est. 2022
HQ: San Francisco, United States • Hardware • Events & Ticketing • Corporate Events • AI Infrastructure • Artificial Intelligence • 501-1000 employees • Est. 2022
$100k – $283k per year (Estimated) • Senior • 5+ years expSingapore • Remote (Singapore) • Full-Time • Senior
Python
AI/ML
Fine-tuning
LLM
LoRA
Quantization
RLHF
SGLang
TensorRT
TensorRT-LLM
Together AI
vLLM
PEFT
DPO
GRPO
Post-training
SFT
Apply
Top 25% payVerified live · 1 hour ago 2 months ago

Research Engineer, Post-Training Inference

$200k – $290k per year • Junior • 2+ years expIn office (San Francisco, United States) • Full-Time • Junior
Python
AI/ML
CUDA
CUDA Toolkit
Fine-tuning
LLM
LoRA
Mamba
Reinforcement Learning
SGLang
TensorRT
TensorRT-LLM
Together AI
Triton
vLLM
PEFT
Post-training
DevOps
Kubernetes
Apply
Open 103 daysVerified live · 1 hour ago 4 months ago

Staff Machine Learning Engineer, Voice AI

$220k – $280k per year • Staff+ • 8+ years expIn office (San Francisco, United States) • Bachelor's Degree • Full-Time • Staff
Python
AI/ML
CUDA
CUDA Toolkit
Fine-tuning
LLM
PyTorch
SGLang
TensorRT
TensorRT-LLM
Together AI
Tokenization
vLLM
Whisper
Deepgram
Text-to-Speech
Speech Recognition
DevOps
Platform Engineering
Apply
Report

Ubisoft

ubisoft.com
Ubisoft is a leading global video game publisher and developer headquartered in France, renowned for creating immersive interactive entertainment across consoles, PC, and mobile platforms. The company maintains an extensive network of development studios worldwide responsible for iconic franchises such as Assassin’s Creed, Far Cry, Tom Clancy’s Rainbow Six, Just Dance, and Rayman.
ubisoft.com • HQ: Saint-Mandé, France • Gaming • Console Games • PC Games • Game Development • 501-1000 employees
HQ: Saint-Mandé, France • Gaming • Console Games • PC Games • Game Development • 501-1000 employees
Verified live · 3 hours ago 19 days ago

Senior Devops - F/H/NB

$59k – $126k per year (Estimated) • Senior • In office (Paris, France) • Full-Time • Senior • French: C1
Python
Rust
AI/ML
LLM
TensorRT
TensorRT-LLM
Triton Inference Server
vLLM
DevOps
gRPC
Kubernetes
Apply
Report

Bjak

bjak.my
Bjak is a Malaysia-based insurtech company founded in 2019. It operates Southeast Asia's largest online insurance platform, letting users compare, buy, and renew insurance policies digitally. The company is headquartered in Petaling Jaya and focuses on making insurance more accessible and convenient.
bjak.my • HQ: Petaling Jaya, Malaysia • Financial Services • FinTech • Insurance • Est. 2019
HQ: Petaling Jaya, Malaysia • Financial Services • FinTech • Insurance • Est. 2019
Verified live · 8 hours ago 21 day ago

Machine Learning Platform Engineer

$116k – $194k per year (Estimated)Zurich • Remote (Switzerland) • Full-Time
Python
AI/ML
JAX
LLM
PyTorch
SGLang
TensorRT
TensorRT-LLM
vLLM
DevOps
Vector
Apply
Report

Eka Care

eka.care
Eka Care provides a personal health record and clinic management platform. Patients store prescriptions and reports while doctors manage consultations digitally. The company integrates with national digital health infrastructure.
eka.care • HQ: Bengaluru, India • Digital Health • Hospitals & Clinics • Health Care • Est. 2020
HQ: Bengaluru, India • Digital Health • Hospitals & Clinics • Health Care • Est. 2020
$32k – $62k per year (Estimated) • Senior • 2+ years expIn office (Bengaluru, India) • Full-Time • Senior
AWQ
CUDA
CUDA Toolkit
GPTQ
LLM
Perplexity
SGLang
TensorRT
TensorRT-LLM
vLLM
Edge AI
Mixture of Experts
Apply
Report

Adobe

adobe.com
Adobe is an American software company founded in 1982 by John Warnock and Charles Geschke and headquartered in San Jose, California. It created the PostScript and PDF formats and built the industry-standard creative toolset around Photoshop, Illustrator, Premiere Pro, After Effects and InDesign, now delivered as the Creative Cloud subscription. The company also runs Document Cloud for electronic signatures and document workflows, Experience Cloud for enterprise marketing analytics and personalisation, and the commercially safe Firefly family of generative AI models.
adobe.com • HQ: San Jose, United States • Sales & Marketing • Document Management • Artificial Intelligence • LLM & Generative AI • Software • Design & Creative • Graphic Design • Productivity Software • Digital Design • 5000+ employees • Est. 1982
HQ: San Jose, United States • Sales & Marketing • Document Management • Artificial Intelligence • LLM & Generative AI • Software • Design & Creative • Graphic Design • Productivity Software • Digital Design • 5000+ employees • Est. 1982
Top 25% payVerified live · 1 day ago 22 days ago

Staff Applied Scientist - VLLM Inference

$164k – $313k per year • Equity • Staff+ • In office (San Jose, United States) • Master's Degree • Full-Time • Staff
Python
Python
Dask
AI/ML
LLM
Multimodal AI
PyTorch
Quantization
Ray
SGLang
Spark
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
vLLM
VLM
OCR
TGI
RAG
DevOps
Kubernetes
Apply
Top 25% payVerified live · 1 day ago 22 days ago

Senior Machine Learning Engineer, AI Platform

$152k – $265k per year • Equity • Senior • In office (San Jose, United States) • Full-Time • Senior
C++
Go
Java
Python
Rust
C++
PyTorch C++
AI/ML
DeepSpeed
LLM
PyTorch
Ray
Ray Serve
TensorRT
TensorRT-LLM
Triton
vLLM
FSDP
DevOps
Kubernetes
Apply
Report

SambaNova Systems

sambanova.ai
SambaNova Systems is an American artificial intelligence semiconductor and cloud platform company headquartered in Palo Alto, California. Founded in 2017 by Stanford professors Kunle Olukotun and Christopher Ré alongside former Oracle executive Rodrigo Liang, the company builds full-stack hardware and software infrastructure purpose-built for enterprise AI, large-scale model training, and high-speed agentic AI inference.
sambanova.ai • HQ: Palo Alto, United States • Artificial Intelligence • Semiconductors • AI Infrastructure
HQ: Palo Alto, United States • Artificial Intelligence • Semiconductors • AI Infrastructure
Verified live · 4 hours ago 3 months ago

Senior Software Engineer - ML Infrastructure

$144k – $258k per year (Estimated) • Senior • 5+ years exp • Remote (United States) • Bachelor's Degree • Full-Time • Senior
Python
AI/ML
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Function Calling
Apply
Report

Deliveroo

deliveroo.co.uk
Deliveroo is an award-winning delivery service founded in 2013 by William Shu and Greg Orlowski. Deliveroo works with approximately 176,000 best-loved restaurants and grocery partners, as well as around 150,000 riders to provide the best food delivery experience in the world. Deliveroo is headquartered in London, with offices around the globe.
deliveroo.co.uk • Lucca • London • Paris • Manchester • Birmingham • Marketplaces • Transportation & Logistics • Delivery
Lucca • London • Paris • Manchester • Birmingham • Marketplaces • Transportation & Logistics • Delivery
Verified live · 1 hour ago 27 days ago

Senior Software Engineer, GenAI Platform

$94k – $195k per year (Estimated) • Senior • 5+ years expIn office (London, United Kingdom) • Bachelor's Degree • Full-Time • Senior
Python
AI/ML
Claude
Claude Code
Cursor
DeepSeek
Embeddings
Fine-tuning
LLM
LoRA
Qwen
Reinforcement Learning
RLHF
PEFT
DPO
LLM Guardrails
OpenAI Codex
Post-training
SFT
AI Agents
AWQ
GPTQ
RAG
SGLang
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
DevOps
AWS
GCP
Kubernetes
Vector
Apply
Verified live · 1 hour ago 2 months ago

Software Engineer, GenAI Platform

$65k – $151k per year (Estimated) • Middle • 3+ years expIn office (London, United Kingdom) • Bachelor's Degree • Full-Time • Middle
Python
AI/ML
Claude
Claude Code
Cursor
DeepSeek
Embeddings
Fine-tuning
LLM
LoRA
Qwen
Reinforcement Learning
RLHF
PEFT
DPO
LLM Guardrails
OpenAI Codex
Post-training
SFT
AI Agents
AWQ
GPTQ
Quantization
RAG
SGLang
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
DevOps
AWS
GCP
Kubernetes
Vector
Apply
Verified live · 1 hour ago 2 months ago

Software Engineer, Machine Learning Infrastructure

$68k – $191k per year (Estimated) • Middle • 3+ years expIn office (London, United Kingdom) • Bachelor's Degree • Full-Time • Middle
Python
AI/ML
Claude
Claude Code
Cursor
DeepSeek
Embeddings
Fine-tuning
LLM
LoRA
Qwen
Reinforcement Learning
RLHF
PEFT
DPO
LLM Guardrails
OpenAI Codex
Post-training
SFT
AI Agents
AWQ
GPTQ
Quantization
RAG
SGLang
TensorRT
TensorRT-LLM
vLLM
Model Context Protocol
DevOps
AWS
GCP
Kubernetes
Vector
Apply
Report

Twelve Labs

twelvelabs.io
Twelve Labs is an artificial intelligence technology company headquartered in San Francisco, California, and founded in 2021. The firm provides a video intelligence platform and API that utilizes multimodal foundation models to enable natural language search, automated summarization, and content categorization within large video libraries. Operating on a global scale, the company serves developers and enterprises in the media, advertising, and public sectors to streamline video workflows and extract insights from unstructured visual data.
twelvelabs.io • HQ: San Francisco, United States • Video & Image Processing Software • Software • LLM & Generative AI • AI Agents • Multimodal AI • Computer Vision • Artificial Intelligence • Est. 2021
HQ: San Francisco, United States • Video & Image Processing Software • Software • LLM & Generative AI • AI Agents • Multimodal AI • Computer Vision • Artificial Intelligence • Est. 2021
Verified live · 1 day ago 27 days ago

Tech Lead Manager, Jockey Core

Lead • Remote/Hybrid (Seoul, South Korea) • Master's Degree • Full-Time • Staff
Databricks
Snowflake
AI/ML
AI Agents
Claude
Embeddings
Gemini
Knowledge Distillation
LLM
Multimodal AI
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
Pre-training
Structured Outputs
Apply
Verified live · 1 day ago 27 days ago

Senior Machine Learning Engineer, Jockey Core

Senior • Remote/Hybrid (Seoul, South Korea) • Master's Degree • Full-Time • Senior
Databricks
Snowflake
AI/ML
AI Agents
Claude
Embeddings
Gemini
LLM
Multimodal AI
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
Pre-training
Structured Outputs
Apply
Report

IQVIA

iqvia.com
IQVIA (NYSE:IQV) is a leading global provider of advanced analytics, technology solutions, and clinical research services to the life sciences industry. IQVIA creates intelligent connections across all aspects of healthcare through its analytics, transformative technology, big data resources and extensive domain expertise. IQVIA Connected Intelligence™ delivers powerful insights with speed and agility - enabling customers to accelerate the clinical development and commercialization of innovative medical treatments that improve healthcare outcomes for patients.
iqvia.com • HQ: Durham, United States • Data Quality • Cybersecurity • Data & Analytics • Health Care • Clinical Research • 5000+ employees
HQ: Durham, United States • Data Quality • Cybersecurity • Data & Analytics • Health Care • Clinical Research • 5000+ employees
Verified live · 1 hour ago 28 days ago

Senior AI Platform Engineer

$94k – $194k per year (Estimated) • Senior • In office (London, United Kingdom) • Full-Time • Senior
AWQ
CUDA
CUDA Toolkit
GPTQ
LLM
LoRA
NVIDIA NIM
PyTorch
Ray
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
cuDNN
Knowledge Graph
NCCL
DevOps
AWS
Kubernetes
Platform Engineering
SLURM
Apply
Report

Lavendo

lavendo.com
lavendo.com • San Francisco • Newark • Austin • Sacramento • Mexico City • Sales Enablement • Sales & Marketing
San Francisco • Newark • Austin • Sacramento • Mexico City • Sales Enablement • Sales & Marketing
Top 25% payVerified live · 9 hours ago 28 days ago

AI Field Engineer, AI infrastructure (Remote - US)

$220k – $280k per year • In office (San Mateo, New York, United States) • Visa sponsorship • Full-Time
Python
SQL
AI/ML
Fine-tuning
LLM
Multimodal AI
SGLang
TensorRT
TensorRT-LLM
vLLM
DPO
Edge AI
SFT
DevOps
AWS
Azure
GCP
Kubernetes
Apply
Report

Rackspace Technology

rackspace.com
Rackspace Technology is a major global multi-cloud solutions, managed infrastructure, and enterprise AI engineering provider headquartered in San Antonio, Texas, USA. Founded in 1998 by Richard Yoo, Pat Condon, and Dirk Elmendorf, the company operates on a B2B managed cloud services, private/public cloud migration, enterprise software integration, and managed infrastructure subscription model led by CEO Amar Maletira.
rackspace.com • HQ: San Antonio, United States • Cybersecurity • Artificial Intelligence • Information Technology • 501-1000 employees • Est. 1998
HQ: San Antonio, United States • Cybersecurity • Artificial Intelligence • Information Technology • 501-1000 employees • Est. 1998
Verified live · 1 day ago 29 days ago

Software Developer IV (Golang & Kubernetes)

$18k – $44k per year (Estimated) • Senior • 9+ years expIndia • Remote (India) • Full-Time • Senior
Go
AI/ML
Claude
Claude Code
Copilot
DeepSeek
Embeddings
Fine-tuning
Gemma
Llama
LLM
LoRA
Mistral
PEFT
Prompt Engineering
Qwen
RAG
Semantic Search
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
vLLM
Transformers
AI Agents
Human-in-the-Loop
LLM Guardrails
OpenAI Codex
Semantic Search
SFT
DevOps
CI/CD
Docker
Git
Kubernetes
Platform Engineering
Vector
GitHub
AWS
Azure
GCP
Grafana
OpenTelemetry
Prometheus
Service Mesh
Apply
Report

Meesho

meesho.com
Meesho is an Indian e-commerce company headquartered in Bengaluru and founded in 2015. The company operates an online marketplace that enables small businesses and individual entrepreneurs to sell products such as apparel, home goods, and electronics directly to consumers. Originally established as a social commerce platform facilitating sales through networks like WhatsApp and Facebook, it has evolved into a large-scale retail ecosystem serving millions of users across India.
meesho.com • HQ: Bengaluru, India • B2B Commerce • Apparel & Fashion • Consumer Goods • Commerce • Marketplaces • Social Commerce • Est. 2015
HQ: Bengaluru, India • B2B Commerce • Apparel & Fashion • Consumer Goods • Commerce • Marketplaces • Social Commerce • Est. 2015
Verified live · 10 hours ago 2 months ago

Engineering Manager – AI Engineering

$42k – $93k per year (Estimated) • Lead • 9+ years expIn office (Bengaluru, India) • Bachelor's Degree • Full-Time • Staff
C++
Python
Rust
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
DeepSpeed
Flink
Knowledge Distillation
LLM
PyTorch
Quantization
Ray
SGLang
Spark
TensorRT
TensorRT-LLM
vLLM
AI Agents
FSDP
LLMOps
Megatron-LM
DevOps
FinOps
Google GKE
Kubernetes
GCP
Apply
Report

CoreWeave

coreweave.com
coreweave.com • San Francisco • Data Centers • Information Technology • Cloud Computing
San Francisco • Data Centers • Information Technology • Cloud Computing
Top 25% pay 2 months ago

Applied AI Engineer, Inference

$188k – $275k per year • Equity • Middle • 4+ years expIn office (San Francisco, United States) • Relocation • Bachelor's Degree • Full-Time • Senior
Python
AI/ML
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
CoreWeave
Apply
Report

Parasail

parasail.com
Access to www.domaineasy.com was denied. You don't have authorization to view this page.
parasail.com • San Mateo • San Francisco • Information Technology
San Mateo • San Francisco • Information Technology
Verified live · 1 day ago 2 months ago

Senior Inference Reliability Engineer

$146k – $282k per year (Estimated) • Senior • 5+ years expIn office (San Mateo, United States) • Full-Time • Senior
C++
Java
Python
Rust
AI/ML
Anomaly Detection
LLM
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
vLLM
DevOps
CI/CD
Kubernetes
Platform Engineering
SRE
Apply
Report

Apple

apple.com
Apple is an American multinational technology company founded in 1976 by Steve Jobs, Steve Wozniak and Ronald Wayne, and headquartered in Cupertino, California. It designs and sells consumer hardware including the iPhone, Mac, iPad, Apple Watch, AirPods and Vision Pro, together with the operating systems and silicon that run them. A growing services division built around the App Store, iCloud, Apple Music, Apple TV+ and Apple Pay now contributes a large share of profit, making Apple one of the most valuable companies in the world.
apple.com • HQ: Cupertino, United States • Payments • Smartwatches & Fitness Trackers • Hardware • Artificial Intelligence • Film Distribution • Commerce • Software • Consumer Electronics • Smartphones • 5000+ employees • Est. 1976
HQ: Cupertino, United States • Payments • Smartwatches & Fitness Trackers • Hardware • Artificial Intelligence • Film Distribution • Commerce • Software • Consumer Electronics • Smartphones • 5000+ employees • Est. 1976
$158k – $320k per year (Estimated) • Equity • Staff+ • 10+ years expIn office (Cupertino, United States) • Visa sponsorship • Bachelor's Degree • Staff
C++
Java
Python
C++
PyTorch C++
Databases
Apache Kafka
ElasticSearch
FAISS
OpenSearch
Cassandra
AI/ML
Embeddings
Flink
PyTorch
Quantization
Spark
XGBoost
DVC
LLM
MLFlow
ONNX
SGLang
TensorRT
TensorRT-LLM
vLLM
Weights & Biases
TorchServe
DevOps
AWS
Docker
Kubernetes
Apply
Report

MMLab@NTU

mmlab-ntu.com
MMLab@NTU is an academic research laboratory housed within Nanyang Technological University (NTU Singapore). Founded in 2018 and led by prominent faculty including Prof. Chen Change Loy and Prof. Ziwei Liu, the lab comprises roughly 40 researchers pioneering artificial intelligence, computer vision, multimodal foundation models, spatial intelligence (3D/4D generation), and embodied AI (vision-language-action robotics).
mmlab-ntu.com • Singapore • Artificial Intelligence • 1001-5000 employees • Est. 2018
Singapore • Artificial Intelligence • 1001-5000 employees • Est. 2018
Verified live · 8 hours ago 2 months ago

AI Engineer, Platforms

$66k – $195k per year (Estimated) • Junior • 1+ year expIn office (Singapore) • Bachelor's Degree • Full-Time • Junior
Python
C++
Rust
AI/ML
AI Agents
Claude
Copilot
Cursor
LLM
Model Context Protocol
RAG
SGLang
TensorRT
TensorRT-LLM
Transformers
vLLM
Multimodal AI
DevOps
CI/CD
Docker
Kubernetes
Platform Engineering
Rest API
Terraform
Apply
Report

Cerebras Systems

cerebras.ai
Cerebras Systems is an American computer hardware company founded in 2016 and headquartered in Sunnyvale, California that builds accelerators for artificial intelligence at wafer scale. Instead of assembling clusters from many small chips, it manufactures a single processor the size of an entire silicon wafer, the Wafer Scale Engine, which removes most of the communication overhead in large model training and inference. The company sells CS-series systems to research laboratories and enterprises, operates its own inference cloud known for very high token throughput, and has built large supercomputers with partners including the Gulf technology group G42.
cerebras.ai • HQ: Sunnyvale, United States • Machine Learning • Artificial Intelligence • Hardware • Semiconductors • AI Infrastructure • 1001-5000 employees • Est. 2016
HQ: Sunnyvale, United States • Machine Learning • Artificial Intelligence • Hardware • Semiconductors • AI Infrastructure • 1001-5000 employees • Est. 2016
Verified live · 12 hours ago 2 months ago

Staff Software Engineer, GPU Inference

$230k – $398k per year (Estimated) • Staff+ • 8+ years expRemote/Hybrid (Toronto, Canada, Sunnyvale, United States) • Bachelor's Degree • Full-Time • Staff
C++
Python
C++
PyTorch C++
AI/ML
Cerebras
LLM
Multimodal AI
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton Inference Server
vLLM
Edge AI
OpenAI
ROCm
AI Agents
CUDA Toolkit
Mixture of Experts
DevOps
CI/CD
Kubernetes
Apply
Verified live · 12 hours ago 10 months ago

Software Engineer, GPU Inference

$237k – $459k per year (Estimated) • Senior • 5+ years expIn office • Bachelor's Degree • Full-Time • Senior
C++
Python
C++
PyTorch C++
AI/ML
Cerebras
LLM
Multimodal AI
PyTorch
SGLang
TensorRT
TensorRT-LLM
vLLM
Quantization
Triton
Triton Inference Server
Edge AI
OpenAI
ROCm
AI Agents
CUDA
CUDA Toolkit
Mixture of Experts
DevOps
CI/CD
Kubernetes
Apply
Report
TensorRT-LLM Jobs - Remote & On-site
Frequently asked questions
TensorRT-LLM: How many jobs are available now?
There are 142 TensorRT-LLM jobs listed on Alion right now; listings are refreshed daily.
TensorRT-LLM: What is the typical salary?
The average listed base salary for TensorRT-LLM roles on Alion is about $289,947 USD, based on current listings.
TensorRT-LLM: Are remote, relocation, or visa sponsorship options available?
Many TensorRT-LLM roles include remote or hybrid flexibility, and some companies may offer relocation or visa support-review each posting for details.
TensorRT-LLM: How do I apply on Alion?
Click a TensorRT-LLM job to view the employer's application link, apply through that link or via Alion's apply flow, and sign in to track your applications.