722,097open jobs
43,116companies
103,109added this week
Browse all
Salary
$200k – $400k per year
Location
In office (Palo Alto)
Seniority
Staff · 3+ years exp

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Feb 17, 2026. RadixArk scores C on the Alion truth index.

Overview
Company
Impact
Profile match
RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible.

About the Role

As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training - spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.

This role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.

In This Role, You Will

  • Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
  • Profile and optimize performance across the full-stack - model implementation, parallelism strategies, communication libraries, and GPU kernels - to improve throughput, latency, memory efficiency, hardware utilization, and cost.
  • Investigate numerical correctness and low-precision issues in distributed training and inference, including train-inference consistency for reinforcement learning.
  • Improve the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.
  • Build supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.
  • Contribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.

Minimum Qualifications

  • 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.
  • Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).
  • Experience in at least two of the following areas:
    • Performance, efficiency, and scalability of multi-GPU, multi-node workloads
    • Numerical correctness or low precision
    • Stability, reliability, or fault tolerance
    • Post-training algorithm recipes and orchestration infrastructure for large training runs
    • Multimodal training or inference, including vision-language models and multimodal generation
    • Agent infrastructure, including sandboxes, harnesses, and eval systems
    • Building and maintaining open-source projects widely adopted in industry and academia

Preferred Qualifications

  • Familiarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.
  • Experience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).
  • Key open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).
  • GPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).
  • Experience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).
  • Top-tier publications in ML systems or other systems fields.

Even if you don't meet every qualification above, we encourage you to apply - we care most about demonstrated ability to build and reason about large-scale systems.

About RadixArk

RadixArk builds open-source and production infrastructure for large language models and multimodal post-training. Our systems - including Miles, an enterprise-grade reinforcement learning training framework, and SGLang, a widely deployed high-performance LLM inference engine - power distributed post-training across clusters of 10k-100k+ GPUs.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 to $400,000, plus equity.

Benefits include a 401(k) plan and unlimited PTO.

RadixArk sponsors employment visas (e.g., H-1B, O-1) for eligible candidates.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
722,097 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$149k – $289k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Santa Clara
Python
C++
C++
TensorFlow C++
PyTorch C++
LLVM
AI/ML
CUDA Toolkit
Quantization
JAX
AI Agents
OpenCL
TensorFlow
PyTorch
CUDA
Triton
MLIR
Apache TVM
Physical AI
DevOps
CI/CD
Apply
$82k – $221k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • Dubai
AI/ML
AI Agents
Machine Learning
Apply
$95k – $191k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree • United Kingdom
Python
JavaScript
TypeScript
SQL
PowerShell
Node JS
AI/ML
Cursor
Claude Code
Vertex AI
Scikit-learn
AI Agents
AWS Bedrock
NumPy
PyTorch
Human-in-the-Loop
Agentic Workflows
DevOps
GCP
Azure
AWS
IAM
Cybersecurity
LDAP
Management
ServiceNow
Apply
$41k – $85k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Pune
Python
Java
Java
Maven
Hibernate
Gradle
Databases
MySQL
PostgreSQL
Redis
DynamoDB
Apache Kafka
Amazon Aurora
AI/ML
Cursor
Claude Code
Vertex AI
Scikit-learn
AI Agents
AWS Bedrock
NumPy
PyTorch
Human-in-the-Loop
LLM Guardrails
Agentic Workflows
DevOps
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Apply
$123k – $227k per year (Estimated) • In office • 5+ years exp • Raleigh
Java
C#
Databases
ActiveMQ
Apache Kafka
Amazon Aurora
AI/ML
Copilot
Cursor
Claude
AI Agents
AWS Strands Agents
DevOps
Rest API
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
Amazon ECS
API Gateway
Apply
$208k – $426k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Palo Alto
AI/ML
SGLang
LLM
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Kubernetes
GitHub
IAM
Linux
Unix
Cybersecurity
OWASP Top 10
SOC 2
HIPAA
Threat Modeling
SIEM
Apply
$200k – $400k per year • In office • 4+ years exp • Bachelor's Degree • Palo Alto
Python
Go
Rust
AI/ML
SGLang
LLM
DevOps
Rest API
gRPC
GCP
Datadog
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Platform Engineering
GitHub
Apply
$160k – $300k per year • In office • 10+ years exp • Palo Alto
AI/ML
SGLang
LLM
DevOps
GitHub
Apply
Product Manager 7 days ago
$160k – $300k per year • Equity • In office • 3+ years exp • Bachelor's Degree • Palo Alto
Python
C++
AI/ML
SGLang
LLM
OpenAI
DevOps
GitHub
Apply
$160k – $300k per year • In office • 4+ years exp • Bachelor's Degree • Palo Alto
Python
C++
AI/ML
CUDA Toolkit
SGLang
LLM
CUDA
TPU
AWS Trainium
DevOps
CI/CD
AWS
GitHub
Apply
$143k – $301k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • Palo Alto
Apex
Apex
MuleSoft
DevOps
SLI/SLO/SLA
Apply
$52k – $143k per year (Estimated) • In office • High School Diploma • Palo Alto
Apply
$59k – $114k per year (Estimated) • In office • 1+ year exp • Palo Alto
Apply
Chief of Staff 1 day ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Management
Stripe
Apply
$48k – $96k per year • In office • Internship • Palo Alto
Apply
See all jobs
This is one of many
722,097 more open roles from verified company boards, updated every day.