SGLang Jobs - Remote & On-site

SGLang roles at vetted startups and product companies. Listings updated daily and include remote-friendly positions and jobs with transparent salary information. Browse and apply today.

AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

T-Systems Iberia

t-systems.es
T-Systems Iberia is the Spanish and Portuguese division of T-Systems, the global enterprise IT services and digital transformation arm of German telecommunications giant Deutsche Telekom. Headquartered in Barcelona, the firm serves major public sector institutions and enterprise clients across Spain and Southern Europe.
t-systems.es • HQ: Barcelona, Spain • IT Consulting • Information Technology • IT Outsourcing • 5000+ employees
HQ: Barcelona, Spain • IT Consulting • Information Technology • IT Outsourcing • 5000+ employees
Verified live · 1 day ago 2 months ago

Lead Senior Backend Engineer (m/f/d)

$50k – $133k per year (Estimated) • Lead • 5+ years expRemote/Hybrid (Granada, Spain) • Bachelor's Degree • Full-Time • Senior
C++
Go
Java
Python
Rust
SQL
Python
Django
FastAPI
Flask
AI/ML
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
gRPC
Kubernetes
Rest API
WebSockets
API Gateway
GitOps
Grafana
OpenTelemetry
Prometheus
Vector
Apply
Report

Radical Numerics

radicalnumerics.ai
Radical Numerics is an artificial intelligence research lab dedicated to developing general biological intelligence and generative genomics models. Headquartered in San Francisco, California, the company builds multimodal platforms capable of reading, writing, and engineering biological sequences across DNA, RNA, and proteins. Its technology aims to accelerate biopharmaceutical research, enhance early disease diagnostics, and establish robust biodefense capabilities.
radicalnumerics.ai • HQ: San Francisco, United States
HQ: San Francisco, United States
Verified live · 1 day ago 2 months ago

Member of Technical Staff, Inference

$193k – $391k per year (Estimated) • Staff+ • In office (San Francisco, United States) • Full-Time • Staff
Python
AI/ML
CUDA
CUDA Toolkit
Multimodal AI
PyTorch
Triton
Mixture of Experts
DeepSpeed
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Apply
Report

Mirantis

mirantis.com
Mirantis is a B2B open-source cloud computing, container management, and AI infrastructure company headquartered in Campbell, California. Originally known as a core contributor to OpenStack, Mirantis now provides open-cloud software and managed services centered on Kubernetes, multi-cloud platforms, and enterprise AI workloads.
mirantis.com • Calgary • Warsaw • Brussels • Prague • Poznań • DevOps • Information Technology • Cloud Computing • 1001-5000 employees
Calgary • Warsaw • Brussels • Prague • Poznań • DevOps • Information Technology • Cloud Computing • 1001-5000 employees
Open 95 daysVerified live · 1 min ago 3 months ago

Product Manager - AI Inference & Model Serving

$145k – $248k per year (Estimated) • Equity • Senior • 7+ years expAustin • Remote (United States) • Full-Time • Senior
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Triton
DevOps
AWS
Kubernetes
Platform Engineering
Chips/EDA
PoC Library
Apply
Report

Machinify

machinify.com
Machinify is an artificial intelligence software company headquartered in Palo Alto, California, and founded in 2016. The company provides an AI-powered operating system designed to automate healthcare payment integrity processes, including claims auditing, subrogation, and pharmacy payment accuracy. It serves major health plans across the United States, managing data for over 270 million members to identify cost savings and reduce administrative friction for payers and providers.
machinify.com • HQ: Palo Alto, United States • Business Process Automation (BPA) • Software • Insurance • Financial Services • Machine Learning • Health Care • Digital Health • Artificial Intelligence • Est. 2016
HQ: Palo Alto, United States • Business Process Automation (BPA) • Software • Insurance • Financial Services • Machine Learning • Health Care • Digital Health • Artificial Intelligence • Est. 2016
Verified live · 1 hour ago 4 months ago

Staff AI Engineer

$210k – $280k per year • Staff+ • 3+ years exp • Remote (United States) • Staff
Python
TypeScript
Databases
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
Claude
Function Calling
Gemini
LangGraph
Llama
LLM
LoRA
Mistral
Prompt Engineering
QLoRA
Qwen
RAG
LangChain
PEFT
DPO
Human-in-the-Loop
LLM Guardrails
SFT
Structured Outputs
Fine-tuning
Model Context Protocol
SGLang
vLLM
TGI
DevOps
Vector
Apply
Report

Saviynt

saviynt.com
We're building revolutionary identity and security solutions to help the world's largest companies migrate to the cloud and solve the toughest security challenges in record time. What's our secret? Our people. We're a global group of innovators wh...
saviynt.com • San Francisco • Bengaluru • Milpitas • Delhi • Atlanta • Artificial Intelligence • Cybersecurity • Identity Management • 501-1000 employees
San Francisco • Bengaluru • Milpitas • Delhi • Atlanta • Artificial Intelligence • Cybersecurity • Identity Management • 501-1000 employees
Top 25% payVerified live · 6 hours ago 3 months ago

AI Platform Engineer, Training and Inference

$274k – $304k per year • Remote/Hybrid (Milpitas, United States) • Bachelor's Degree • Full-Time
Python
Databases
pgvector
Qdrant
PostgreSQL
AI/ML
Fine-tuning
LLM
MLFlow
ONNX
PyTorch
Quantization
RAG
Ray
Ray Serve
RLHF
RLlib
SGLang
TensorRT
vLLM
Triton
Flyte
FSDP
GRPO
NCCL
PPO
AWQ
Bitsandbytes
GPTQ
Post-training
DevOps
Google GKE
Vector
GCP
Kubernetes
Amazon S3
QA
k6
Apply
Report

SemiWiki

semiwiki.com
SemiWiki is an open community-driven semiconductor web portal and knowledge-sharing platform dedicated to the global electronic design automation (EDA) and semiconductor manufacturing industries. Founded in 2010 by industry veterans Daniel Nenni, Paul McLellan, and David Manners, SemiWiki provides analysis, expert blogging, and commentary on advanced node semiconductor technology, chip design methodologies, foundry economics, and supply chain dynamics.
semiwiki.com • HQ: Danville, United States • Cybersecurity • Manufacturing • Hardware • 1001-5000 employees • Est. 2010
HQ: Danville, United States • Cybersecurity • Manufacturing • Hardware • 1001-5000 employees • Est. 2010
Staff+ • 5+ years expIn office (Hsinchu, Taiwan) • Master's Degree • Full-Time • Staff
C++
Python
C++
LLVM
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
LLM
OpenCL
SGLang
Triton
vLLM
ROCm
TGI
PyTorch
Apply
Report

Fathom

fathom.ai
Fathom captures, transcribes, and summarizes Zoom, Google Meet, and Microsoft Teams calls. Free for individuals, with AI-powered CRM updates for teams.
fathom.ai • Productivity Software • LLM & Generative AI • Artificial Intelligence • Est. 2020
Productivity Software • LLM & Generative AI • Artificial Intelligence • Est. 2020
Open 117 daysVerified live · 8 hours ago 3 months ago

AI Engineer - Model Performance

$138k – $302k per year (Estimated) • Remote (United States) • Master's Degree • Full-Time
Python
AI/ML
Axolotl
CUDA Toolkit
Fine-tuning
LLM
LoRA
Multimodal AI
Prompt Engineering
QLoRA
Quantization
Ray
Ray Serve
SGLang
TensorRT
TensorRT-LLM
torchtune
vLLM
PEFT
PyTorch
CUDA
DPO
SFT
Speech Recognition
DevOps
GitHub
Management
Notion
Slack
Marketing
HubSpot
Apply
Report

H Company

hcompany.ai
H Company is an artificial intelligence research company headquartered in Paris, France, and founded in 2023 by Charles Kantor with a founding team drawn largely from Google DeepMind. The company develops agentic AI systems, including Runner H, an agent that carries out multi-step tasks in a browser on the user's behalf. It raised one of the largest seed rounds in European history and positions itself as a French alternative in the agent tooling market.
hcompany.ai • HQ: Paris, France • Business Process Automation (BPA) • Software • Multimodal AI • LLM & Generative AI • AI Agents • Artificial Intelligence • 201-500 employees • Est. 2023
HQ: Paris, France • Business Process Automation (BPA) • Software • Multimodal AI • LLM & Generative AI • AI Agents • Artificial Intelligence • 201-500 employees • Est. 2023
Open 138 daysVerified live · 1 day ago 4 months ago

Research Engineer, Model Inference & Serving - Paris

$77k – $183k per year (Estimated)Remote/Hybrid (Paris, France) • Full-Time
C++
Python
Rust
C++
PyTorch C++
AI/ML
AI Agents
JAX
Multimodal AI
PyTorch
SGLang
vLLM
Edge AI
CUDA Toolkit
llama.cpp
LLM
MLX ML
ONNX
Quantization
TensorRT
TensorRT-LLM
CUDA
Triton
DevOps
Kubernetes
Apply
Open 142 daysVerified live · 1 day ago 4 months ago

Research Engineer, Model Inference & Serving - London

$104k – $219k per year (Estimated)Remote/Hybrid (London, United Kingdom) • Full-Time
C++
Python
Rust
C++
PyTorch C++
AI/ML
AI Agents
JAX
Multimodal AI
PyTorch
SGLang
vLLM
Edge AI
CUDA Toolkit
llama.cpp
LLM
MLX ML
ONNX
Quantization
TensorRT
TensorRT-LLM
CUDA
Triton
DevOps
Kubernetes
Apply
Report

Reflection AI

reflection.ai
Reflection AI is a company founded in 2024 by former Google DeepMind researchers who worked on AlphaGo and large language models. It builds autonomous coding agents and has committed to releasing frontier open-weight models as an American counterweight to Chinese open model releases. The company raised a very large round in 2025 to fund training at frontier scale.
reflection.ai • HQ: New York, United States • 501-1000 employees • Est. 2024
HQ: New York, United States • 501-1000 employees • Est. 2024
Open 159 daysVerified live · 2 hours ago 5 months ago

Member of Technical Staff - Mid-Training Infra

$202k – $350k per year (Estimated) • Equity • Staff+ • In office (San Francisco, New York, United States, London, United Kingdom) • Visa sponsorship • Full-Time • Staff
LLM
Reinforcement Learning
SGLang
Synthetic Data
Megatron-LM
Apply
Report

Zyphra

zyphra.com
COMPANY RESEARCH CLOUD Zyphra Cloud Login Two sides. Two sides.
zyphra.com • HQ: San Francisco, United States • AI Infrastructure • Artificial Intelligence • LLM & Generative AI
HQ: San Francisco, United States • AI Infrastructure • Artificial Intelligence • LLM & Generative AI
Open 167 daysVerified live · 1 day ago 5 months ago

Platform Engineer

$125k – $280k per year (Estimated)In office (San Francisco, United States) • Relocation • Full-Time
Ray
SGLang
vLLM
Triton
DevOps
Ansible
AWS
CI/CD
Docker
GCP
Kubernetes
SLURM
Terraform
Chaos Engineering
Apply
Report

Parspec

parspec.io
Parspec is a technology company that leverages AI to help sales agents and distributors by simplifying the process of discovering and sourcing the best available construction products and materials.
parspec.io • Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Bengaluru • San Mateo • Commerce • Marketplaces • Artificial Intelligence • Est. 2020
Open 174 daysVerified live · 11 hours ago 5 months ago

AI Ops Engineer

$30k – $75k per year (Estimated) • Senior • 5+ years expIn office (Bengaluru, India) • Bachelor's Degree • Full-Time • Senior
Python
Python
Asyncio
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
Weaviate
PostgreSQL
Qdrant
AI/ML
AWS Bedrock
Kubeflow
LiteLLM
LLM
MLFlow
Portkey
Ray
vLLM
AWQ
CUDA
CUDA Toolkit
Embeddings
GPTQ
Hallucination
Langfuse
LoRA
Prompt Engineering
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
Triton Inference Server
PEFT
Amazon SageMaker
AWS Trainium
LLM Guardrails
LLMOps
NCCL
NVLink
TGI
Multimodal AI
AI Agents
DevOps
AIOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
GitHub Actions
Grafana
Kubernetes
OpenTelemetry
Prometheus
Terraform
Vector
Karpenter
Platform Engineering
Amazon EventBridge
Amazon S3
API Gateway
IAM
AWS Step Functions
GitHub
Cybersecurity
Least Privilege
Apply
Report

Elastix

elastix.ai
Elastix AI delivers scalable, energy-efficient AI inference through machine learning, system software, and reconfigurable hardware.
elastix.ai • Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Seattle • Hardware • AI Infrastructure • Artificial Intelligence
Open 193 days 6 months ago

AI Software Engineer

$138k – $259k per year (Estimated) • Equity • Middle • 3+ years expIn office (Seattle, United States) • Bachelor's Degree • Full-Time • Middle
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
CUDA
DevOps
Docker
Kubernetes
Apply
Report

Percepta

percepta.com
Headquartered in Dearborn, Michigan, Percepta is a global customer experience (CX) and BPO services provider specializing in the automotive sector. Originally established as a joint venture between Ford Motor Company and TeleTech (now TTEC), the company delivers customer care, technical support, warranty administration, and dealer relations services. Additionally, it operates contact centers and support networks worldwide to help automotive brands optimize customer loyalty and streamline operations.
percepta.com • HQ: Dearborn, United States • Professional Services
HQ: Dearborn, United States • Professional Services
Open 220 daysVerified live · 2 hours ago 7 months ago

Research Engineer / Scientist – Reinforcement Learning (RL)

$195k – $427k per year (Estimated)In office (New York, Boston, United States) • Master's Degree • Full-Time
Python
AI/ML
LLM
Ray
Reinforcement Learning
SGLang
vLLM
Anthropic
Post-training
AI Agents
DevOps
Amazon EKS
AWS
Kubernetes
Apply
Report

Inference

inference.ai
Models. Agents. GPUs. Whatever you need — we've got it. Ghost agent VMs, Maestro model routing, Engine wholesale GPUs, and Academy, on one platform.
inference.ai • San Francisco
San Francisco
Open 221 dayTop 25% pay 7 months ago

Senior Software Engineer - Model Performance

$220k – $320k per year • Senior • 2+ years expIn office (San Francisco, United States) • Full-Time • Senior
C++
Python
C#
C++
PyTorch C++
C#
.NET
AI/ML
CUDA Toolkit
Knowledge Distillation
LLM
LoRA
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
PEFT
CUDA
GPT-5
Embeddings
DevOps
Docker
Kubernetes
GitHub
Apply
Report

Cohere

cohere.com
Cohere is a Canadian AI company founded in 2019 and headquartered in Toronto. It builds secure, enterprise-focused large language models and AI tools for businesses and regulated industries. The company is known for emphasizing privacy, private deployment, and "sovereign AI" rather than consumer chat products
cohere.com • HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
HQ: Toronto, Canada • Energy & Utilities • Biotechnology • Natural Language Processing • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2019
Verified live · 1 hour ago 9 months ago

Audio Inference Engineer, Model Efficiency

$132k – $289k per year (Estimated)New YorkSan FranciscoTorontoMontreal • Remote (United States, Canada) • Full-Time
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
Cohere SDK
LLM
PyTorch
SGLang
TensorFlow
vLLM
Apply
Verified live · 1 hour ago 9 months ago

Member of Technical Staff, Model Efficiency

$173k – $323k per year (Estimated) • Staff+ • 5+ years expNew YorkSan FranciscoTorontoMontreal • Remote (United States, Canada) • Full-Time • Staff
C++
Python
Rust
AI/ML
Cohere SDK
CUDA Toolkit
LLM
SGLang
vLLM
CUDA
Mixture of Experts
Apply
Report

Specter

specter.com
Specter tracks growth signals across millions of private companies to help investors find opportunities early. Founded in 2021, it combines web, hiring, app and social data into company momentum scores. Venture funds and corporate development teams use it for sourcing.
specter.com • HQ: Singapore • Est. 2021
HQ: Singapore • Est. 2021
Open 331 dayVerified live · 7 hours ago 10 months ago

Software Engineer - ML Infrastructure

$158k – $346k per year (Estimated)In office (San Francisco, United States) • Full-Time
C++
Python
Rust
C++
PyTorch C++
TensorFlow C++
Databases
LanceDB
Qdrant
AI/ML
AI Agents
Computer Vision
LLM
Multimodal AI
PyTorch
Ray
SGLang
Spark
TensorFlow
TensorRT
TensorRT-LLM
VLM
Robotics
Sensor Fusion
Analytics
A/B Testing
Apply
Report

Sesame AI

sesame.com
Sesame AI is a company founded in 2022 that develops conversational speech models with unusually natural timing and prosody. Its demonstration voices drew wide attention for sounding present in a conversation rather than reading text aloud, and it released a base speech model openly. The company is also building lightweight eyewear intended to keep an assistant available throughout the day.
sesame.com • HQ: San Francisco, United States • Est. 2022
HQ: San Francisco, United States • Est. 2022
Open 534 daysTop 25% pay 1 year ago

ML Model Serving Engineer

$175k – $280k per year • Equity • In office (San Francisco, New York, Bellevue, United States) • Full-Time
LLM
PyTorch
SGLang
Ray
DevOps
AWS
Azure
GCP
Kubernetes
Apply
Report

Cartesia

cartesia.ai
Cartesia is an artificial intelligence research and technology company that builds real-time, low-latency generative voice models for interactive applications. Founded by Stanford researchers, the company utilizes State Space Models (SSMs) to power its flagship text-to-speech platform, Sonic, which enables sub-second voice synthesis, instant voice cloning, and emotional expression across dozens of languages. By providing real-time speech generation and voice agent infrastructure, Cartesia allows developers and enterprises to power conversational AI systems, virtual assistants, and customer service platforms.
cartesia.ai • HQ: San Francisco, United States • LLM & Generative AI • Artificial Intelligence • Speech & Audio AI • 11-50 employees • Est. 2022
HQ: San Francisco, United States • LLM & Generative AI • Artificial Intelligence • Speech & Audio AI • 11-50 employees • Est. 2022
Open 626 daysTop 25% pay 2 years ago

Inference Engineer

$180k – $250k per year • Equity • In office (San Francisco, United States) • Visa sponsorship • Full-Time
Databricks
AI/ML
CUDA Toolkit
Multimodal AI
SGLang
Transformers
vLLM
CUDA
Triton
Edge AI
Apply
Report

Binance

binance.com
Binance is the world's largest cryptocurrency exchange by trading volume, founded in 2017 by Changpeng Zhao and Yi He. The platform offers spot, margin and derivatives trading across hundreds of digital assets, alongside staking, savings products, payments, an institutional custody arm and a self-custodial Web3 wallet. The group also created BNB Chain, one of the most used smart contract networks, and now operates under a licensed regional structure after a 2023 settlement with United States authorities that installed new leadership and compliance oversight.
binance.com • HQ: Abu Dhabi, United Arab Emirates • Mining & Staking • Crypto Payments • Web3 & dApps • Blockchain & Crypto • Cryptocurrencies • Crypto Exchanges • 1001-5000 employees • Est. 2017
HQ: Abu Dhabi, United Arab Emirates • Mining & Staking • Crypto Payments • Web3 & dApps • Blockchain & Crypto • Cryptocurrencies • Crypto Exchanges • 1001-5000 employees • Est. 2017
Open 1451 dayVerified live · 13 min ago 4 years ago

LLM Data Scientist/Algorithm Engineer (Fully Remote)

$64k – $190k per year (Estimated) • Junior • 2+ years exp • Remote (Japan, Taiwan, Hong Kong, Singapore, Australia) • Master's Degree • Full-Time • Junior • Chinese: B2
AutoGen
Chain-of-Thought
CrewAI
Hallucination
LangGraph
LLM
NLP
Prompt Engineering
Quantization
RAG
SGLang
vLLM
LangChain
DPO
SFT
AI Agents
Apply
Report
SGLang Jobs - Remote & On-site
Frequently asked questions
SGLang: How many jobs are available now?
There are 219 SGLang jobs on Alion right now.
SGLang: What is the average salary for these roles?
The average salary for SGLang roles on Alion is $295,494 per year.
SGLang: Are remote, relocation, or visa options available?
Many SGLang listings are remote-friendly or hybrid, and several employers offer relocation or visa sponsorship.
SGLang: How do I apply on Alion?
Create an Alion account, view the SGLang job posting, and click Apply to submit your resume and profile.