DeepSpeed Jobs - Remote & On-site

DeepSpeed jobs curated from vetted startups and product companies, updated daily. Many DeepSpeed roles offer remote or hybrid schedules plus transparent pay ranges to streamline decision-making. Browse and apply today.

AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

Xponentiate

xponentiate.com
Founded in 2022 out of a venture studio, Xponentiate is a specialized recruitment and talent acquisition firm dedicated exclusively to the healthcare sector. The company provides tailored hiring solutions to build C-suite leadership, technical, clinical, and operational teams for digital health, medtech, and pharmaceutical organizations globally. By leveraging deep industry network insights, it connects health-tech startups and enterprise healthcare clients with vetted professionals to accelerate organizational growth.
xponentiate.com • Delhi • Mumbai • Bengaluru • New Delhi • San Francisco • Professional Services
Delhi • Mumbai • Bengaluru • New Delhi • San Francisco • Professional Services
Open 141 dayVerified live · 1 day ago 5 months ago

Senior AI Engineer

$31k – $79k per year (Estimated) • Senior • 4+ years expIn office (Bengaluru, India) • Full-Time • Senior
Python
AI/ML
Accelerate
DeepSpeed
Embeddings
Falcon
Fine-tuning
Hallucination
Knowledge Distillation
LLM
LoRA
Mistral
Multimodal AI
NLP
Prompt Engineering
PyTorch
QLoRA
Qwen
RAG
Tokenization
PEFT
Triton
Hugging Face
Human-in-the-Loop
LLM Guardrails
SFT
AI Agents
DevOps
Docker
Kubernetes
Apply
Report

Andromeda

andromeda.ai
Andromeda connects AI teams with high-performance compute fast, at scale, and on terms that work. Buy, sell, and operate gpu clusters without complexity.
andromeda.ai • San Francisco • AI Infrastructure
San Francisco • AI Infrastructure
Open 144 daysVerified live · 11 hours ago 5 months ago

Senior Site Reliability Engineer

$81k – $170k per year (Estimated) • Senior • San Francisco • Remote (United States, Canada) • Full-Time • Senior
Bash
Python
AI/ML
CUDA Toolkit
DeepSpeed
PyTorch
CUDA
FSDP
InfiniBand
Megatron-LM
NCCL
NVLink
DevOps
Ansible
Grafana
Helm
Incident Management
Kubernetes
Prometheus
Self-Healing
SLURM
Terraform
HPC
Apply
Report

Dyna Robotics

dynarobotics.com
Dyna Robotics is a robotics foundation model company headquartered in Redwood City, California, and founded in 2024. The company trains general purpose manipulation models that let a single robot arm learn commercial tasks such as folding, packing, and handling from demonstration rather than explicit programming. It targets repetitive physical work in laundries, restaurants, and light manufacturing, and publishes model progress through its DYNA series of releases.
dynarobotics.com • HQ: Redwood City, United States • Multimodal AI • Service & Hospitality Robotics • Robotics AI • Artificial Intelligence • Robotic Software & Control Systems • Robotics • 51-200 employees • Est. 2024
HQ: Redwood City, United States • Multimodal AI • Service & Hospitality Robotics • Robotics AI • Artificial Intelligence • Robotic Software & Control Systems • Robotics • 51-200 employees • Est. 2024
Open 153 daysTop 25% pay 6 months ago

ML Infrastructure Engineer, Training

$220k – $320k per year • Senior • 7+ years expIn office (Redwood City, United States) • Full-Time • Senior
Accelerate
DeepSpeed
Knowledge Distillation
Multimodal AI
PyTorch
Quantization
TensorRT
Triton
FSDP
NCCL
DevOps
AWS
GCP
Kubernetes
SLURM
HPC
Apply
Report

Arcade

arcade.agency
arcade.agency • San Francisco • Design & Creative • Sales & Marketing
San Francisco • Design & Creative • Sales & Marketing
Open 154 daysTop 25% pay 6 months ago

Senior Machine Learning Engineer

$180k – $300k per year • Senior • 5+ years expIn office (San Francisco, United States) • Relocation • Full-Time • Senior
Python
JavaScript
TypeScript
AI/ML
DeepSpeed
Fine-tuning
LoRA
PEFT
PyTorch
Ray
Transformers
Hugging Face
Frontend
React.js
Next.js
Apply
Report

Reflection AI

reflection.ai
Reflection AI is a company founded in 2024 by former Google DeepMind researchers who worked on AlphaGo and large language models. It builds autonomous coding agents and has committed to releasing frontier open-weight models as an American counterweight to Chinese open model releases. The company raised a very large round in 2025 to fund training at frontier scale.
reflection.ai • HQ: New York, United States • 501-1000 employees • Est. 2024
HQ: New York, United States • 501-1000 employees • Est. 2024
Open 160 daysVerified live · 4 hours ago 6 months ago

Member of Technical Staff - Pre-Training Infra

$208k – $360k per year (Estimated) • Equity • Staff+ • In office (San Francisco, New York, United States, London, United Kingdom) • Visa sponsorship • Full-Time • Staff
DeepSpeed
Megatron-LM
NCCL
Pre-training
Apply
Report

Hark

hark.com
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.
hark.com • HQ: San Jose, United States • Machine Learning • Artificial Intelligence
HQ: San Jose, United States • Machine Learning • Artificial Intelligence
Open 161 dayTop 25% pay 6 months ago

Member of Technical Staff, Pretraining

$180k – $450k per year • Staff+ • In office (San Jose, United States) • Full-Time • Staff
DeepSpeed
LLM
Multimodal AI
Synthetic Data
Megatron-LM
AI Agents
Apply
Open 161 dayTop 25% pay 6 months ago

Member of Technical Staff, Multimodal Vision

$180k – $450k per year • Staff+ • In office (San Jose, United States) • Full-Time • Staff
DeepSpeed
Fine-tuning
Multimodal AI
Reinforcement Learning
RLHF
DPO
FSDP
GRPO
Megatron-LM
Post-training
PPO
SFT
AI Agents
Apply
Report

Build

build.io
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.
build.io • HQ: London, United Kingdom • Web Development • Software • Data Centers • AI Agents • Artificial Intelligence • DevOps • Information Technology • Cloud Computing • Est. 2023
HQ: London, United Kingdom • Web Development • Software • Data Centers • AI Agents • Artificial Intelligence • DevOps • Information Technology • Cloud Computing • Est. 2023
Open 164 daysVerified live · 1 day ago 6 months ago

Member of Technical Staff - Post Training

$63k – $138k per year (Estimated) • Staff+ • Remote/Hybrid (Yokohama, Japan) • Full-Time • Staff
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
Report

JetBrains

jetbrains.com
JetBrains is a software company founded in 2000 that builds integrated development environments and team tools for professional programmers. Its IDE family covers most major languages, including IntelliJ IDEA for the Java virtual machine, PyCharm, WebStorm, GoLand, CLion and Rider, all sharing a code-intelligence engine that gave the company its reputation for refactoring and static analysis. Headquartered in Prague with development offices across Europe, it also created the Kotlin programming language, now Google's preferred language for Android, and ships team products such as TeamCity, YouTrack and its AI coding assistants.
jetbrains.com • HQ: Prague, Czech Republic • Project Management • Artificial Intelligence • Software • Code Intelligence • Productivity Software • 1001-5000 employees • Est. 2000
HQ: Prague, Czech Republic • Project Management • Artificial Intelligence • Software • Code Intelligence • Productivity Software • 1001-5000 employees • Est. 2000
Verified live · 1 hour ago 6 months ago

Staff Research Engineer (LLM Pre-Training)

Staff+ • Amsterdam • Remote (Germany, United Kingdom, Cyprus, Czech Republic, Netherlands) • Staff
Python
AI/ML
DeepSpeed
Fine-tuning
Kubeflow
LLM
NLP
PyTorch
RLHF
TensorRT
vLLM
Hugging Face
Pre-training
DevOps
CI/CD
Git
Kubernetes
TeamCity
Apply
Verified live · 1 hour ago 8 months ago

Senior MLOps Engineer (ML Workflows Engineering)

Senior • 3+ years expAmsterdam • Remote (Germany, Cyprus, Czech Republic, Netherlands, Poland) • Senior
Java
Kotlin
Python
AI/ML
Dagster
DeepSpeed
Fine-tuning
Langfuse
LLM
MLFlow
NLP
TensorRT
vLLM
DevOps
AWS
CI/CD
GCP
GitHub Actions
Kubernetes
TeamCity
GitHub
Apply
Report

Elastix

elastix.ai
Elastix AI delivers scalable, energy-efficient AI inference through machine learning, system software, and reconfigurable hardware.
elastix.ai • Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Open 194 daysVerified live · 1 day ago 7 months ago

AI Software Engineer

$138k – $259k per year (Estimated) • Equity • Middle • 3+ years expIn office (Seattle, United States) • Bachelor's Degree • Full-Time • Middle
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
CUDA
DevOps
Docker
Kubernetes
Apply
Report

Pear VC

pear.vc
Pear VC is a venture capital firm specializing in pre-seed and seed-stage investments in startups. Founded by Pejman Nozad and Mar Hershenson, the firm provides hands-on support in areas such as product development, recruiting, fundraising, and go-to-market strategy. Pear has backed companies including DoorDash, Gusto, Aurora Solar, Vanta, and Guardant Health.
pear.vc • HQ: San Francisco, United States • Financial Services • Venture Capital
HQ: San Francisco, United States • Financial Services • Venture Capital
Open 219 daysVerified live · 1 day ago 8 months ago

Member of Technical Staff, Frontend

Staff+ • In office (San Francisco, United States) • Full-Time • Staff
TypeScript
JavaScript
Databases
Snowflake
AI/ML
DeepSpeed
ONNX
Frontend
Next.js
React.js
Framer Motion
Tailwind CSS
Design
Figma
Framer
Apply
Open 307 daysVerified live · 1 day ago 11 months ago

Member of Technical Staff, Backend - NomadicML

$238k – $482k per year (Estimated) • Staff+ • In office (San Francisco, United States) • Full-Time • Staff
Go
Python
TypeScript
JavaScript
Databases
Snowflake
AI/ML
Dagster
DeepSpeed
ONNX
Ray
Multimodal AI
Ray Serve
Triton
Frontend
Next.js
React.js
DevOps
AWS
Azure
GCP
gRPC
Kubernetes
Amazon S3
IAM
Apply
Open 307 daysVerified live · 1 day ago 11 months ago

Member of Technical Staff, Machine Learning - NomadicML

$243k – $491k per year (Estimated) • Staff+ • In office (San Francisco, United States) • Full-Time • Staff
Python
Databases
Snowflake
AI/ML
DeepSpeed
Embeddings
Fine-tuning
Multimodal AI
ONNX
PyTorch
AI Agents
Kubeflow
MLFlow
Ray
Hugging Face
Apply
Report

Baseten

baseten.co
Baseten is an AI infrastructure platform designed to help developers and machine learning teams deploy, serve, and scale open-source and custom AI models. The platform provides performant, low-latency inference infrastructure alongside developer tools like Truss, an open-source model packaging framework. Headquartered in San Francisco, California, Baseten enables companies to run state-of-the-art models in production seamlessly without managing underlying cloud infrastructure.
baseten.co • HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
Top 25% payVerified live · 2 hours ago 8 months ago

Software Engineer - Training Product

$165k – $330k per year • Senior • 5+ years expRemote/Hybrid (San Francisco, New York, United States) • Full-Time • Senior
Cursor
DeepSeek
Reinforcement Learning
vLLM
PEFT
Axolotl
DeepSpeed
Fine-tuning
LoRA
PyTorch
Synthetic Data
FSDP
Megatron-LM
NCCL
Post-training
SFT
DevOps
Kubernetes
Apply
Open 367 daysTop 25% pay 1 year ago

Software Engineer - Training Infrastructure

$165k – $330k per year • Remote/Hybrid (San Francisco, New York, United States) • Bachelor's Degree • Full-Time
Go
Python
AI/ML
Cursor
Reinforcement Learning
Spark
DeepSpeed
PyTorch
FSDP
Megatron-LM
NCCL
DevOps
AWS
DigitalOcean
GCP
Kubernetes
Apply
Report
Firmus Technologies builds immersion-cooled artificial intelligence factories that run large GPU fleets on renewable power. Founded in 2021 in Singapore, it develops both the data centre design and the cloud service on top. Its Project Southgate campuses in Australia are among the region's largest planned artificial intelligence sites.
firmus.ai • HQ: Singapore • Artificial Intelligence • Est. 2021
HQ: Singapore • Artificial Intelligence • Est. 2021
Open 228 daysVerified live · 8 hours ago 8 months ago

AI Engineer, AI & Applications

$90k – $227k per year (Estimated) • Senior • 5+ years expIn office (Singapore) • Full-Time • Senior
DeepSpeed
Fine-tuning
JAX
Llama
LoRA
PyTorch
QLoRA
PEFT
FSDP
LLM Evaluation
Megatron-LM
NCCL
DevOps
Kubernetes
SLURM
Apply
Report

Inference

inference.ai
Models. Agents. GPUs. Whatever you need — we've got it. Ghost agent VMs, Maestro model routing, Engine wholesale GPUs, and Academy, on one platform.
inference.ai • San Francisco
San Francisco
Open 238 daysTop 25% pay 8 months ago

Machine Learning Researcher

$250k – $350k per year • Middle • 3+ years expIn office (San Francisco, United States) • Full-Time • Middle
C#
C#
.NET
AI/ML
DeepSpeed
Knowledge Distillation
LLM
PyTorch
Reinforcement Learning
RLHF
Transformers
TRL
DPO
GPT-5
Hugging Face
Megatron-LM
Post-training
SFT
Multimodal AI
DevOps
GitHub
Apply
Open 238 daysTop 25% pay 8 months ago

Applied Machine Learning Engineer

$220k – $320k per year • Junior • 2+ years expIn office (San Francisco, United States) • Full-Time • Junior
C#
C#
.NET
AI/ML
Axolotl
DeepSpeed
Knowledge Distillation
LLM
PyTorch
Transformers
GPT-5
Hugging Face
Post-training
SFT
Multimodal AI
DevOps
GitHub
Analytics
ETL/ELT
Apply
Report

Liquid AI

liquid.ai
Liquid AI is an artificial intelligence company headquartered in Boston, Massachusetts, and founded in 2023 as a spin-off from the MIT Computer Science and Artificial Intelligence Laboratory. The company builds Liquid Foundation Models, an architecture derived from liquid neural networks that aims to match transformer quality at a fraction of the memory and compute. It targets on-device and edge deployment where models must run on phones, vehicles, and embedded hardware rather than in a data center.
liquid.ai • HQ: Boston, United States • AI Infrastructure • Hardware • Machine Learning • Edge AI • LLM & Generative AI • Artificial Intelligence • Est. 2023
HQ: Boston, United States • AI Infrastructure • Hardware • Machine Learning • Edge AI • LLM & Generative AI • Artificial Intelligence • Est. 2023
Open 258 daysVerified live · 2 hours ago 9 months ago

Member of Technical Staff - Multi-Modal, Audio

$199k – $404k per year (Estimated) • Staff+ • Remote/Hybrid (San Francisco, Boston, United States) • Full-Time • Staff
DeepSpeed
Multimodal AI
PyTorch
FSDP
Text-to-Speech
Speech Recognition
Apply
Open 312 daysVerified live · 2 hours ago 11 months ago

Member of Technical Staff - Multi-Modal, Vision

$194k – $394k per year (Estimated) • Staff+ • Remote/Hybrid (San Francisco, United States) • Master's Degree • Full-Time • Staff
Python
AI/ML
Computer Vision
DeepSpeed
Multimodal AI
Reinforcement Learning
VLM
FSDP
Hugging Face
Megatron-LM
Post-training
SFT
DevOps
GitHub
Apply
Open 398 daysVerified live · 2 hours ago 1 year ago

Member of Technical Staff - Distributed Training Engineer

$199k – $404k per year (Estimated) • Staff+ • Remote/Hybrid (San Francisco, United States) • Full-Time • Staff
DeepSpeed
Multimodal AI
PyTorch
FSDP
Megatron-LM
NCCL
Mixture of Experts
Apply
Report

Tahoe Therapeutics

tahoebio.ai
Tahoe Therapeutics generates very large single-cell perturbation datasets to train foundation models of how cells respond to drugs. Founded in 2023 in San Francisco, it released the Tahoe-100M dataset openly to the research community. Its models aim to predict drug effects across cancer cell contexts before laboratory testing.
tahoebio.ai • HQ: San Francisco, United States • Artificial Intelligence • Biotechnology • Health Care • Est. 2023
HQ: San Francisco, United States • Artificial Intelligence • Biotechnology • Health Care • Est. 2023
Open 273 days 9 months ago

Senior Machine Learning Engineer

$150k – $290k per year (Estimated) • Senior • In office (South San Francisco, United States) • Senior
JAX
Keras
Multimodal AI
PyTorch
TensorFlow
Transformers
Accelerate
DeepSpeed
Megatron-LM
Apply
Report

Cohere

cohere.com
Cohere is a Canadian AI company founded in 2019 and headquartered in Toronto. It builds secure, enterprise-focused large language models and AI tools for businesses and regulated industries. The company is known for emphasizing privacy, private deployment, and "sovereign AI" rather than consumer chat products
cohere.com • HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
Verified live · 9 min ago 9 months ago

Senior ML Systems Engineer, Frameworks & Tooling

$98k – $200k per year (Estimated) • Senior • LondonNew YorkParisSan FranciscoToronto • Remote (United States, France, United Kingdom, Canada) • Full-Time • Senior
Cohere SDK
CUDA Toolkit
JAX
LLM
Ray
CUDA
FSDP
NCCL
DeepSpeed
PyTorch
TensorRT
TensorRT-LLM
vLLM
xFormers
Megatron-LM
DevOps
Docker
Kubernetes
SLURM
HPC
Apply
Report

SambaNova Systems

sambanova.ai
SambaNova Systems is an American artificial intelligence company founded in 2017 by Stanford professors Kunle Olukotun and Christopher Re together with former Oracle executive Rodrigo Liang. It designs the Reconfigurable Dataflow Unit, a processor whose datapath is configured to match a model's computation graph rather than executing fixed instructions, and pairs it with a three-tier memory system that keeps very large models resident. Headquartered in Palo Alto, California, the company sells complete systems and a hosted cloud service to enterprises and governments that want to run frontier open models on their own terms rather than through a public model provider.
sambanova.ai • HQ: Palo Alto, United States • Data Centers • LLM & Generative AI • Hardware • Artificial Intelligence • Semiconductors • AI Infrastructure • 201-500 employees • Est. 2017
HQ: Palo Alto, United States • Data Centers • LLM & Generative AI • Hardware • Artificial Intelligence • Semiconductors • AI Infrastructure • 201-500 employees • Est. 2017
Open 283 daysVerified live · 6 hours ago 10 months ago

Senior AI Performance Engineer

$151k – $293k per year (Estimated) • Senior • 3+ years expIn office (San Jose, United States) • Bachelor's Degree • Full-Time • Senior
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
DeepSeek
JAX
Llama
PyTorch
Quantization
Qwen
TensorFlow
CUDA
CUDA Toolkit
DeepSpeed
LLM
Multimodal AI
OpenCL
TensorRT
Triton
vLLM
cuDNN
Megatron-LM
Apply
Report

webAI

webai.com
webAI is an artificial intelligence platform company headquartered in Austin, Texas, and founded in 2019 by David Stout, Tyler Mauer, and Ethan Baird. The company builds software that trains and runs large models on hardware the customer already owns, including its Navigator model building tool, webFrame for spreading models across local clusters, and a device native assistant. It targets organizations in aviation, healthcare, manufacturing, financial services, and the public sector that need AI without sending data to a public cloud.
webai.com • HQ: Austin, United States • IT Infrastructure • Information Technology • AI Infrastructure • Hardware • LLM & Generative AI • Edge AI • Artificial Intelligence • Est. 2019
HQ: Austin, United States • IT Infrastructure • Information Technology • AI Infrastructure • Hardware • LLM & Generative AI • Edge AI • Artificial Intelligence • Est. 2019
Open 402 daysVerified live · 6 hours ago 1 year ago

Staff R&D AI Engineer

$175k – $354k per year (Estimated) • Equity • Staff+ • 7+ years expIn office (Austin, United States) • PhD • Full-Time • Staff
Python
AI/ML
Computer Vision
Fine-tuning
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
TensorFlow
Edge AI
PPO
DeepSpeed
Knowledge Distillation
Quantization
DevOps
AWS
Azure
GCP
Robotics
Reinforcement Learning
Imitation Learning
Sim-to-Real
Apply
Report

Entrata

entrata.com
Entrata provides an operating system for multifamily property management companies. Its modules cover leasing, accounting, resident payments, marketing websites and facilities. Large apartment owners use it as a single platform across their portfolios.
entrata.com • HQ: Lehi, United States • PropTech • Real Estate • Property Management • Est. 2003
HQ: Lehi, United States • PropTech • Real Estate • Property Management • Est. 2003
Open 403 daysTop 25% pay 1 year ago

Senior Machine Learning Engineer

$144k – $233k per year • Senior • 5+ years expRemote/Hybrid (Lehi, United States) • Bachelor's Degree • Full-Time • Senior
Python
AI/ML
Fine-tuning
LLM
PyTorch
Synthetic Data
Post-training
SFT
AI Agents
DeepSpeed
vLLM
FSDP
Apply
Report
DeepSpeed Jobs - Remote & On-site
Frequently asked questions
DeepSpeed: How many jobs are available now?
There are 129 DeepSpeed job openings listed on Alion right now; listings are updated daily.
DeepSpeed: What is the typical salary for DeepSpeed roles?
The typical DeepSpeed role on Alion averages $278,569 USD annually.
DeepSpeed: Are remote, relocation, or visa sponsorship options available?
DeepSpeed listings feature remote, hybrid, and on-site roles, and some companies may offer relocation or visa sponsorship-check each listing for specifics.
DeepSpeed: How do I apply for DeepSpeed jobs on Alion?
Create an account on Alion, upload your resume, and apply directly via the DeepSpeed job listings; your application will be sent to the employer.