738,400open jobs
44,340companies
105,262added this week
Browse all
Salary
$200k – $400k per year
Location
In office (Palo Alto)
Seniority
Staff · 4+ years exp

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Jul 23, 2026. RadixArk scores C on the Alion truth index.

Overview
Company
Impact
Profile match
RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible.

About the Role

RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible on modern GPU hardware. Our systems sit at the center of how modern AI is served and trained: SGLang is a high-performance inference engine that serves trillions of tokens daily across leading AI companies and research labs, and Miles is our reinforcement-learning post-training framework for large-scale LLM and MoE models. Your work directly advances our mission to democratize AI: every improvement you ship lowers the cost and raises the ceiling of what developers everywhere can build.

As our technical face to a community of expert users and partners, you'll push the performance of SGLang and Miles through the lens of real production workloads. You'll profile and optimize GPU performance, enable new models and hardware, build kernels, deliver day-0 model support, and push the limits of inference and training. Working in close partnership with leading teams across the ecosystem, you'll turn their hardest, most ambiguous problems into concrete wins and clear guidance, and feed those improvements back into our systems and future roadmap.

Key Responsibilities

  • Accelerate AI workloads. Profile and optimize GPU performance for real production workloads on current and next-generation hardware, root-causing bottlenecks from kernels to distributed multi-node systems.

  • Go deep in one or two focus areas. The team collectively covers the full stack; each engineer specializes in one or two tracks:

    • Inference performance: engine tuning, benchmarking, long-context and multi-turn optimization, parallelism strategy, production debugging

    • Kernels and model/hardware enablement: custom CUDA/ROCm/Triton kernels, low-precision quantization, day-0 support for new models on new silicon

    • Speculative decoding: draft-model training, acceptance-rate tuning, cross-platform kernel adaptation

    • Training systems: RL post-training with Miles, FP8 training, elasticity, long-rollout and long-context efficiency

  • Partner directly with the ecosystem. Turn ambiguous, high-stakes problems from expert engineers at our key partners into concrete wins, clear technical guidance, and reproducible cookbooks.

  • Enhance SGLang and Miles. Feed user-driven improvements back into our open-source systems and roadmap, so every win compounds across the ecosystem.

Qualifications

Minimum Requirements

  • 4+ years of experience in GPU systems, LLM infrastructure, or performance engineering.

  • Strong profiling and debugging skills: able to root-cause performance and correctness issues across the stack.

  • Hands-on GPU programming experience in at least one of CUDA, ROCm, or Triton, and willingness to work across platforms.

  • Strong programming skills in Python plus C++ or CUDA.

  • Comfortable making progress on hard, ambiguous problems with little context to start from, and fast to ramp into unfamiliar systems, codebases, and domains.

  • Ability to translate ambiguous asks into clear technical plans, verified cookbooks, and actionable recommendations, and to communicate credibly with expert engineering audiences.

Preferred (Bonus) Qualifications

  • Deep familiarity with LLM inference internals: distributed serving, parallelism, routing, KV-cache management, scheduling.

  • Experience with low-precision quantization and inference/training (FP8, INT8/INT4; NVFP4 or MXFP4 a strong plus).

  • Experience writing and optimizing custom GPU kernels.

  • Practical familiarity with speculative decoding methods such as Eagle, DFlash, or DSpark.

  • Working knowledge of large-scale distributed training: pre-training, SFT, RL post-training, elasticity, long-context workloads.

  • Experience optimizing across both NVIDIA and AMD platforms.

  • Hands-on experience with SGLang, Miles, vLLM, TensorRT-LLM, Megatron, or comparable frameworks; contributions to open-source AI/ML projects.

About RadixArk

RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
738,400 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Palo Alto
Senior ML Engineer 4 hours ago
$188k – $200k per year • Remote/Hybrid • 5+ years exp • Master's Degree • Palo Alto
Python
C++
AI/ML
Fine-tuning
AI Agents
Amazon SageMaker
Machine Learning
DevOps
GCP
AWS
Apply
$130k – $180k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Toronto
Python
Databases
Databricks
AI/ML
Copilot
Prompt Engineering
AI Agents
RAG
OpenAI
Knowledge Graph
Copilot Studio
Machine Learning
DevOps
Azure
GitHub
Apply
$118k – $231k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Toronto
AI/ML
AI Agents
LLMOps
Multi-Agent Systems
Machine Learning
DevOps
Platform Engineering
Incident Management
Apply
$78k – $165k per year (Estimated) • Remote/Hybrid • 2+ years exp • High School Diploma • Montreal
Python
Go
Java
Databases
PostgreSQL
Apache Kafka
AI/ML
LLM
Edge AI
Machine Learning
DevOps
gRPC
CI/CD
Apply
$51k – $70k per year • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Toronto
Python
SQL
Python
Flask
FastAPI
AI/ML
CatBoost
Model Context Protocol
XGBoost
Scikit-learn
AI Agents
Gradio
Pandas
NumPy
RAG
Streamlit
OpenAI
Machine Learning
DevOps
Rest API
Azure
CI/CD
Apply
$70k – $235k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Columbus • Tampa • Dallas • Atlanta
Python
Java
AI/ML
Claude
Vertex AI
AI Agents
RAG
OpenAI
Anthropic
Context Engineering
LLM Evaluation
Agentic Workflows
DevOps
Terraform
Helm
CI/CD
Docker
Kubernetes
Apply
$94k – $194k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Toronto
Python
JavaScript
Java
TypeScript
SQL
Node JS
Databases
Databricks
Firestore
Google BigQuery
BigQuery
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Vertex AI
Prompt Engineering
CrewAI
Gemini
LLM
RAG
Google ADK
Multi-Agent Systems
Machine Learning
Frontend
Angular
React.js
Mobile
Firebase
DevOps
Rest API
GCP
Azure DevOps
Dynatrace
Azure
CI/CD
Kubernetes
Google Cloud Run
Cybersecurity
SIEM
Management
Agile
Apply
$11k – $12k per year • In office • 1+ year exp • Zelenograd
Python
JavaScript
SQL
Python
Flask
SQLAlchemy
Databases
PostgreSQL
Frontend
Bootstrap
JQuery
DevOps
Docker
Linux
Apply
$70k – $100k per year • Equity • In office • Full-Time • Master's Degree • Boston
Python
DevOps
Linux
Apply
$195k – $323k per year • In office • Full-Time • New York • Arlington
Python
AI/ML
Fine-tuning
Embeddings
TensorFlow
PyTorch
Apply
$200k – $400k per year • In office • 4+ years exp • Palo Alto
Python
C++
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
LLM
CUDA
Triton
Megatron-LM
TPU
NCCL
ROCm
MLIR
Apache TVM
XLA
DevOps
GitHub
HPC
Apply
$200k – $400k per year • In office • 4+ years exp • Palo Alto
Python
C++
AI/ML
CUDA Toolkit
SGLang
LLM
CUDA
DevOps
GitHub
Management
Discord
Apply
$200k – $400k per year • In office • Palo Alto
Python
C++
AI/ML
CUDA Toolkit
SGLang
LLM
Mixture of Experts
CUDA
Triton
TPU
ROCm
XLA
Speculative Decoding
KV Cache
DevOps
GitHub
Apply
$200k – $400k per year • In office • 5+ years exp • Bachelor's Degree • Palo Alto
Python
Go
JavaScript
Rust
TypeScript
C++
C++
PyTorch C++
AI/ML
DeepSpeed
vLLM
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Ray
Megatron-LM
Edge AI
DevOps
GitHub
Apply
$200k – $400k per year • In office • Palo Alto
Python
Bash
AI/ML
CUDA Toolkit
Quantization
SGLang
LLM
CUDA
NCCL
InfiniBand
ROCm
KV Cache
DevOps
GitHub Actions
GitLab CI
CI/CD
Jenkins
Docker
Buildkite
GitHub
Linux
Management
Slack
Apply
Senior ML Engineer 4 hours ago
$188k – $200k per year • Remote/Hybrid • 5+ years exp • Master's Degree • Palo Alto
Python
C++
AI/ML
Fine-tuning
AI Agents
Amazon SageMaker
Machine Learning
DevOps
GCP
AWS
Apply
$440k per year • In office • Palo Alto
Python
Rust
C++
AI/ML
Reinforcement Learning
LLM
Apply
$440k per year • In office • Palo Alto
Python
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
PyTorch
LLM
CUDA
Apply
$600k per year • In office • Palo Alto
AI/ML
RLHF
Reinforcement Learning
DPO
Post-training
Reward Modeling
Apply
$440k per year • In office • Palo Alto
Python
Rust
C++
C++
PyTorch C++
AI/ML
Spark
Fine-tuning
JAX
Multimodal AI
Function Calling
AI Agents
PyTorch
Ray
Tokenization
SFT
Post-training
Pre-training
TPU
XLA
Tool Use
Reward Modeling
DevOps
Kubernetes
Apply
See all jobs
This is one of many
738,400 more open roles from verified company boards, updated every day.