CUDA Toolkit Jobs - Remote & On-site

CUDA Toolkit - curated CUDA Toolkit jobs and roles at vetted startups and product companies, updated daily. Find remote and hybrid positions on Alion and narrow results by domain, seniority, and GPU expertise. Browse and apply today.

AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

vCluster

vcluster.com
vCluster is a Kubernetes infrastructure company headquartered in San Francisco, California, and founded in 2019. The company builds virtual clusters that run inside a single physical Kubernetes cluster, giving each team or tenant what looks like its own isolated control plane without the cost of separate infrastructure. It maintains the open source vCluster project alongside a commercial platform and previously operated as Loft Labs.
vcluster.com • HQ: San Francisco, United States • Cloud Management • Cloud Computing • Software • Virtualization • DevOps • Information Technology • Est. 2019
HQ: San Francisco, United States • Cloud Management • Cloud Computing • Software • Virtualization • DevOps • Information Technology • Est. 2019
Open 153 daysVerified live · 11 min ago 6 months ago

AI Infrastructure Engineer

$116k – $224k per year (Estimated) • Senior • 5+ years exp • Remote (United States, Germany, Australia) • Full-Time • Senior
Python
AI/ML
CUDA
CUDA Toolkit
InfiniBand
LLM
OpenAI
DevOps
Kubernetes
GitHub
GitLab
Marketing
Salesforce
Apply
Report

Meshy

meshy.ai
Meshy generates three-dimensional models and textures from text prompts or reference images. Founded in 2021 by former Nvidia and Microsoft graphics researchers, it targets game studios and product designers who need assets quickly. The output is exported in standard formats ready for game engines.
meshy.ai • HQ: San Jose, United States • Design & Creative • Generative 3D • Artificial Intelligence • Est. 2021
HQ: San Jose, United States • Design & Creative • Generative 3D • Artificial Intelligence • Est. 2021
Open 154 daysVerified live · 11 hours ago 6 months ago

Machine Learning Engineer

Remote/Hybrid (Shenzhen, China) • Full-Time
C++
Python
AI/ML
CUDA
CUDA Toolkit
Quantization
Apply
Open 157 daysVerified live · 11 hours ago 6 months ago

AI 3D Dataset Engineer

In office (Shanghai, China) • PhD • Full-Time
C++
Python
Databases
Databricks
AI/ML
CUDA
CUDA Toolkit
DevOps
CI/CD
Vector
GitHub
Apply
Open 362 daysVerified live · 11 hours ago 12 months ago

Machine Learning Engineer – HPC

Remote/Hybrid (Shanghai, China) • Full-Time
C++
Python
AI/ML
CUDA
CUDA Toolkit
Quantization
DevOps
HPC
Apply
Report

E2E Networks

e2enetworks.com
E2E Networks is a listed Indian cloud provider that pivoted to accelerated computing for artificial intelligence workloads. Founded in 2009 in Delhi, it operates large GPU clusters in Indian data centres for startups, research institutions and enterprises. Its TIR platform provides managed notebooks, fine-tuning and inference endpoints.
e2enetworks.com • HQ: New Delhi, India • Cloud Computing • Artificial Intelligence • AI Infrastructure • Est. 2009
HQ: New Delhi, India • Cloud Computing • Artificial Intelligence • AI Infrastructure • Est. 2009
Open 159 daysVerified live · 1 day ago 6 months ago

Senior Solutions Architect

$31k – $75k per year (Estimated) • Staff+ • In office (Delhi, India) • Architect
CUDA
CUDA Toolkit
Kubeflow
ONNX
PyTorch
Ray
TensorFlow
TensorRT
cuDNN
CVAT
InfiniBand
DevOps
CentOS Stream
Debian
Docker
Kubernetes
KVM
SLURM
Ubuntu
VMWare
Apply
Report

Zyphra

zyphra.com
COMPANY RESEARCH CLOUD Zyphra Cloud Login Two sides. Two sides.
zyphra.com • HQ: San Francisco, United States • AI Infrastructure • Artificial Intelligence • LLM & Generative AI
HQ: San Francisco, United States • AI Infrastructure • Artificial Intelligence • LLM & Generative AI
Open 167 daysVerified live · 2 days ago 6 months ago

Research Engineer - AI Performance & Kernel Optimization

$156k – $342k per year (Estimated)In office (San Francisco, United States) • Relocation • Full-Time
CUDA Toolkit
CUDA
Triton
AWS Trainium
TPU
Mixture of Experts
DevOps
AWS
HPC
Apply
Report

Prime Robotics

primerobotics.com
Unlock the potential of your warehouse with our end-to-end warehouse automation. Experience efficiency with autonomous mobile robots and robotic automation systems.
primerobotics.com • Denver • Lakewood • Warehousing • Robotics • Logistics & Warehouse Robotics
Denver • Lakewood • Warehousing • Robotics • Logistics & Warehouse Robotics
Open 168 daysVerified live · 1 day ago 6 months ago

Senior Robotics Engineer

$150k per year • Senior • 8+ years expIn office (Denver, United States) • Full-Time • Senior
C++
Java
Python
C++
Protobuf
Java
Spring Boot
AI/ML
Computer Vision
CUDA Toolkit
CUDA
DevOps
Docker
gRPC
Robotics
Isaac Sim
Motion Planning
Path Planning
ROS
ROS2
Sensor Fusion
SLAM
Apply
Report

Fal

fal.ai
Fal is a generative media AI infrastructure platform headquartered in San Francisco, California. Founded in 2021 by former Coinbase and Amazon engineers Burkay Gur and Gorkem Yurtseven, the enterprise is backed by prominent venture investors including Andreessen Horowitz (a16z), Bessemer Venture Partners, and Salesforce Ventures.
fal.ai • HQ: San Francisco, United States • Generative Video • LLM & Generative AI • Artificial Intelligence • Est. 2021
HQ: San Francisco, United States • Generative Video • LLM & Generative AI • Artificial Intelligence • Est. 2021
Open 170 daysVerified live · 1 min ago 6 months ago

Software Engineer, Platform

Middle • 3+ years expIn office • Full-Time • Middle
Python
AI/ML
CUDA
CUDA Toolkit
InfiniBand
NVLink
DevOps
Ansible
Configuration Management
Terraform
KVM
QEMU
Cybersecurity
ISO 27001
SOC 2
Tcpdump
Apply
Open 188 daysTop 25% pay 7 months ago

Software Engineer, Platform

$180k – $250k per year • Middle • 3+ years expIn office (San Francisco, United States) • Relocation • Full-Time • Middle
Python
AI/ML
CUDA
CUDA Toolkit
InfiniBand
NVLink
DevOps
Ansible
Configuration Management
Terraform
KVM
QEMU
Cybersecurity
ISO 27001
SOC 2
Tcpdump
Apply
Report

Tether

tether.to
Tether is a digital asset company founded in 2014 that issues USDT, the largest stablecoin in circulation and the most heavily traded asset in crypto markets. Each token is intended to hold a one-to-one peg with the United States dollar and is backed by a reserve portfolio dominated by short-dated Treasury bills, with regular attestation reports published on the reserve composition. Headquartered in San Salvador, El Salvador, the group has expanded beyond stablecoin issuance into gold-backed tokens, Bitcoin mining, energy, peer-to-peer communications and local artificial intelligence infrastructure.
tether.to • HQ: San Salvador, El Salvador • Crypto Payments • Data Centers • Artificial Intelligence • Blockchain & Crypto • Cryptocurrencies • Stablecoins • 1001-5000 employees • Est. 2014
HQ: San Salvador, El Salvador • Crypto Payments • Data Centers • Artificial Intelligence • Blockchain & Crypto • Cryptocurrencies • Stablecoins • 1001-5000 employees • Est. 2014
Verified live · 1 day ago 6 months ago

AI Inference Engineer QVAC (100% remote Worldwide)

Remote (Colombia) • English: B2
C++
JavaScript
AI/ML
CUDA Toolkit
Diffusion Models
Llama
llama.cpp
LocalAI
OpenCL
Tokenization
CUDA
Edge AI
LLM
Web3
Bitcoin
Apply
Report

Parspec

parspec.io
Parspec is a technology company that leverages AI to help sales agents and distributors by simplifying the process of discovering and sourcing the best available construction products and materials.
parspec.io • Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Open 174 daysVerified live · 1 day ago 6 months ago

AI Ops Engineer

$30k – $75k per year (Estimated) • Senior • 5+ years expIn office (Bengaluru, India) • Bachelor's Degree • Full-Time • Senior
Python
Python
Asyncio
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
Weaviate
PostgreSQL
Qdrant
AI/ML
AWS Bedrock
Kubeflow
LiteLLM
LLM
MLFlow
Portkey
Ray
vLLM
AWQ
CUDA
CUDA Toolkit
Embeddings
GPTQ
Hallucination
Langfuse
LoRA
Prompt Engineering
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
PEFT
Amazon SageMaker
AWS Trainium
LLM Guardrails
LLMOps
NCCL
NVLink
TGI
Multimodal AI
AI Agents
DevOps
AIOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
GitHub Actions
Grafana
Kubernetes
OpenTelemetry
Prometheus
Terraform
Vector
Karpenter
Platform Engineering
Amazon EventBridge
Amazon S3
API Gateway
IAM
AWS Step Functions
GitHub
Cybersecurity
Least Privilege
Apply
Report

D-Wave

dwavequantum.com
D-Wave is an industry leader in the development and production of quantum computing systems, software, and services. As the first commercial supplier of quantum computers, D-Wave is at the forefront of the dual development of quantum annealing and gate-model systems. D-Wave helps companies use quantum tech to solve real-world problems right now.
dwavequantum.com • HQ: Palo Alto, United States • Hardware • Artificial Intelligence • Science & Engineering • Quantum Computing Hardware • Quantum Science & Computing
HQ: Palo Alto, United States • Hardware • Artificial Intelligence • Science & Engineering • Quantum Computing Hardware • Quantum Science & Computing
Verified live · 1 hour ago 6 months ago

Senior Quantum Software Engineer, Compiler [410]

$125k – $187k per year • Senior • 5+ years expRemote/Hybrid (New Haven, United States) • Master's Degree • Full-Time • Senior
C++
Python
Q#
SQL
C++
LLVM
Q#
Quantum Intermediate Representation
Databases
Oracle
PostgreSQL
AI/ML
CUDA Toolkit
CUDA
DevOps
CI/CD
Git
Quantum
Cirq
Qiskit
Apply
Report

ThingTrax

thingtrax.com
Built for the realities of the factory floor thingtrax was founded by engineers and technologists who've lived the challenges of manufacturing. From machine monitoring to full-line optimisation, our journey has always been about solving real problems with practical, scalable technology.
thingtrax.com • Lahore • Gurgaon • Liverpool • Industrial IoT (IIoT) • Manufacturing • Industrial Automation
Lahore • Gurgaon • Liverpool • Industrial IoT (IIoT) • Manufacturing • Industrial Automation
Open 192 daysVerified live · 1 day ago 7 months ago

Computer Vision & Hardware Engineer

$28k – $115k per year (Estimated) • Middle • 3+ years expIn office (Gurgaon, India) • Bachelor's Degree • Full-Time • Middle
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
Computer Vision
CUDA Toolkit
OpenCL
CUDA
AI Agents
Image Segmentation
PyTorch
TensorFlow
DevOps
Git
Apply
Report

Elastix

elastix.ai
Elastix AI delivers scalable, energy-efficient AI inference through machine learning, system software, and reconfigurable hardware.
elastix.ai • Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Open 193 daysVerified live · 1 hour ago 7 months ago

AI Software Engineer

$138k – $259k per year (Estimated) • Equity • Middle • 3+ years expIn office (Seattle, United States) • Bachelor's Degree • Full-Time • Middle
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
CUDA
DevOps
Docker
Kubernetes
Apply
Report

Generalist

generalist.com
The people, companies, and technologies shaping the future. Click to read The Generalist, by Mario Gabriele, a Substack publication with hundreds of thousands of subscribers.
generalist.com • San Francisco • Boston
San Francisco • Boston
Open 199 daysTop 25% pay 7 months ago

Software Engineer: ML Optimization

$200k – $350k per year • In office (San Francisco, United States) • Full-Time
Python
AI/ML
ChatGPT
CUDA
CUDA Toolkit
Gemini
Multimodal AI
NumPy
GPT-4
NVLink
OpenAI
Apply
Report

HeyGen

heygen.com
HeyGen is a company founded in 2020 that generates presenter videos from text using synthetic avatars and cloned voices. Its translation feature re-voices existing footage in other languages while matching lip movement, which made it popular with marketing and training teams. The company reports fast revenue growth serving businesses that produce large volumes of localised video.
heygen.com • HQ: Los Angeles, United States • Semiconductors • Embedded Systems • Cybersecurity • Est. 2020
HQ: Los Angeles, United States • Semiconductors • Embedded Systems • Cybersecurity • Est. 2020
Open 318 daysVerified live · 2 days ago 11 months ago

Tech Lead, AI Compute Infrastructure

$118k – $282k per year (Estimated) • Lead • 5+ years expIn office (Los Angeles, United States) • Bachelor's Degree • Full-Time • Staff
C++
Python
C++
PyTorch C++
TensorFlow C++
Databases
LanceDB
AI/ML
Accelerate
CUDA
CUDA Toolkit
JAX
Multimodal AI
PyTorch
Ray
Spark
TensorFlow
Scale AI
Diffusion Models
NCCL
DevOps
Kubernetes
HPC
Apply
Open 271 dayVerified live · 2 days ago 9 months ago

Software Engineer, AI Compute Infrastructure

$83k – $235k per year (Estimated) • Senior • 5+ years expIn office (Los Angeles, United States) • Bachelor's Degree • Full-Time • Senior
C++
Python
C++
PyTorch C++
TensorFlow C++
Databases
LanceDB
AI/ML
Accelerate
CUDA
CUDA Toolkit
JAX
Multimodal AI
PyTorch
Ray
Spark
TensorFlow
Scale AI
Diffusion Models
NCCL
DevOps
Kubernetes
HPC
Apply
Report

Architect AI

architect.ai
Architect AI is an artificial intelligence platform designed to transform urban design, architecture, and interior concepts into high-resolution visual renderings. The tool allows architects, designers, and real estate professionals to generate detailed 3D designs and explore spatial ideas rapidly using text prompts, sketch transformations, and mood boards. By streamlining the initial conceptual and visualization stages of building design, it helps creative teams accelerate project planning and present polished design options to clients.
architect.ai • Palo Alto • Bengaluru
Palo Alto • Bengaluru
Open 201 dayVerified live · 1 day ago 7 months ago

Member of Technical Staff - ML Research

$180k – $365k per year (Estimated) • Equity • Staff+ • In office (Palo Alto, United States) • Bachelor's Degree • Full-Time • Staff
CUDA
CUDA Toolkit
Fine-tuning
PyTorch
QLoRA
Reinforcement Learning
LoRA
PEFT
Anthropic
Edge AI
OpenAI
Post-training
Function Calling
Apply
Report

AHEAD

ahead.com
AHEAD helps enterprises build modern, secure, and scalable digital platforms by combining cloud, data, AI, and automation. Their consulting and managed services drive real business impact through smarter IT.
ahead.com • Gurgaon • New York • Palo Alto • Chicago • Minneapolis • Artificial Intelligence • Cloud Computing • Information Technology
Gurgaon • New York • Palo Alto • Chicago • Minneapolis • Artificial Intelligence • Cloud Computing • Information Technology
Open 212 daysVerified live · 3 hours ago 8 months ago

AI Platform Engineer

$23k – $63k per year (Estimated) • Middle • 4+ years exp • Remote (India) • Middle
Python
AI/ML
CUDA Toolkit
Kubeflow
MLFlow
PyTorch
TensorFlow
TensorRT
CUDA
Triton
NVIDIA NeMo
DevOps
Ansible
AWS
Azure
CI/CD
Configuration Management
GCP
Kubernetes
Terraform
HPC
Apply
Report

Inference

inference.ai
Models. Agents. GPUs. Whatever you need — we've got it. Ghost agent VMs, Maestro model routing, Engine wholesale GPUs, and Academy, on one platform.
inference.ai • San Francisco
San Francisco
Open 221 dayTop 25% pay 8 months ago

Senior Software Engineer - Model Performance

$220k – $320k per year • Senior • 2+ years expIn office (San Francisco, United States) • Full-Time • Senior
C++
Python
C#
C++
PyTorch C++
C#
.NET
AI/ML
CUDA Toolkit
Knowledge Distillation
LLM
LoRA
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
CUDA
GPT-5
Embeddings
DevOps
Docker
Kubernetes
GitHub
Apply
Report

Dexmate

dexmate.com
Dexmate is a robotics company headquartered in Santa Clara, California, and founded in 2024 by doctoral researchers from MIT, UC San Diego, and Carnegie Mellon. The company builds Vega, a foldable wheeled humanoid with two arms, dexterous hands, and an extending torso that reaches from floor level to overhead. It sells the platform to research groups and commercial pilots working on dexterous manipulation, and has taken strategic investment from NEC.
dexmate.com • HQ: Santa Clara, United States • Logistics & Warehouse Robotics • Robotic Software & Control Systems • Robotics AI • Humanoid Robots • Artificial Intelligence • Robotics • 201-500 employees • Est. 2024
HQ: Santa Clara, United States • Logistics & Warehouse Robotics • Robotic Software & Control Systems • Robotics AI • Humanoid Robots • Artificial Intelligence • Robotics • 201-500 employees • Est. 2024
Open 223 daysTop 25% pay 8 months ago

Senior robotics navigation engineer

$120k – $300k per year • Senior • 5+ years expIn office (Fremont, United States) • Master's Degree • Full-Time • Senior
C++
AI/ML
Multimodal AI
CUDA Toolkit
CUDA
Robotics
Ceres Solver
GTSAM
ROS
Sensor Fusion
SLAM
Visual-Inertial Odometry
Cartographer
LIO-SAM
ORB-SLAM3
RTAB-Map
Apply
Report

Prosync

prosync.com
ProSync, Professionals In-Sync, is an Employee Owned and Veteran Owned Small Business defense contractor that provides technology solutions and professional services to the federal government. Our primary focus is people and ensuring they have a ...
prosync.com • HQ: San Diego, United States • Cybersecurity • Military • Information Technology
HQ: San Diego, United States • Cybersecurity • Military • Information Technology
Open 228 daysVerified live · 1 day ago 8 months ago

Software Engineer III (DevOps)

$84k – $177k per year (Estimated) • Middle • 4+ years expIn office • Bachelor's Degree • Contractor • Middle
C++
Python
AI/ML
CUDA Toolkit
CUDA
DevOps
CI/CD
Docker
Git
GitLab CI
Helm
Jenkins
Kubernetes
GitLab
Apply
Report

Etched

etched.com
Etched. We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads.
etched.com • San Jose • Taipei • Austin • Artificial Intelligence • AI Infrastructure • Semiconductors
San Jose • Taipei • Austin • Artificial Intelligence • AI Infrastructure • Semiconductors
Top 25% payVerified live · 8 hours ago 8 months ago

Kernel Driver Software Engineer

$150k – $275k per year • In office (San Jose, United States) • Relocation • Full-Time
C++
Rust
C++
PyTorch C++
TensorFlow C++
AI/ML
CUDA Toolkit
OpenCL
PyTorch
TensorFlow
CUDA
DevOps
Git
CI/CD
Docker
Kubernetes
Apply
Report

Cohere

cohere.com
Cohere is a Canadian AI company founded in 2019 and headquartered in Toronto. It builds secure, enterprise-focused large language models and AI tools for businesses and regulated industries. The company is known for emphasizing privacy, private deployment, and "sovereign AI" rather than consumer chat products
cohere.com • HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
Verified live · 29 min ago 9 months ago

Senior ML Systems Engineer, Frameworks & Tooling

$98k – $200k per year (Estimated) • Senior • LondonNew YorkParisSan FranciscoToronto • Remote (United States, France, United Kingdom, Canada) • Full-Time • Senior
Cohere SDK
CUDA Toolkit
JAX
LLM
Ray
CUDA
FSDP
NCCL
DeepSpeed
PyTorch
TensorRT
TensorRT-LLM
vLLM
xFormers
Megatron-LM
DevOps
Docker
Kubernetes
SLURM
HPC
Apply
Verified live · 29 min ago 10 months ago

Member of Technical Staff, Model Efficiency

$173k – $323k per year (Estimated) • Staff+ • 5+ years expNew YorkSan FranciscoTorontoMontreal • Remote (United States, Canada) • Full-Time • Staff
C++
Python
Rust
AI/ML
Cohere SDK
CUDA Toolkit
LLM
SGLang
vLLM
CUDA
Mixture of Experts
Apply
Open 556 daysVerified live · 29 min ago 1 year ago

Member of Technical Staff, Training Performance Engineer

$124k – $271k per year (Estimated) • Staff+ • Remote/Hybrid (London, United Kingdom, New York, United States, Paris, France, Toronto, Montreal, Canada) • Full-Time • Staff
Python
AI/ML
CUDA Toolkit
JAX
PyTorch
Transformers
CUDA
Triton
Pre-training
Apply
Report
CUDA Toolkit Jobs - Remote & On-site
Frequently asked questions
CUDA Toolkit: How many jobs are available now?
There are 740 active CUDA Toolkit jobs listed on Alion right now.
CUDA Toolkit: What is the typical salary for CUDA Toolkit roles?
The average listed salary for CUDA Toolkit roles on Alion is $270,354 USD.
CUDA Toolkit: Are remote, relocation, or visa options available?
Many CUDA Toolkit listings include remote or hybrid options; some employers may offer relocation or visa sponsorship-review each posting for specifics.
CUDA Toolkit: How do I apply to CUDA Toolkit jobs on Alion?
Open the CUDA Toolkit listing on Alion and click the application link or follow the employer's instructions; sign in to save roles and get alerts.