{"id":827760,"url":"https://alion.io/job/radixark-member-of-technical-staff-training","title":"Member of Technical Staff — Training","company":{"id":686458,"name":"RadixArk","domain":"radixark.com","url":"https://alion.io/company/radixark","size_band":null,"is_staffing_agency":false,"is_intermediary":false,"ats_vendor":"Greenhouse","truth_index":{"grade":"C","score":55,"open_postings":20,"ghost_share":0.75,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-24T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"staff","employment_type":null,"work_mode":"on_site","remote_scope":null,"hiring_geo_confidence":"structured","locations":["Palo Alto, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":200000,"max":400000,"currency":"USD","period":"year","gross":null,"usd_annual":400000},"salary_estimate":null,"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"FSDP","optional":false},{"name":"LLM","optional":false},{"name":"LLM Evaluation","optional":false},{"name":"Machine Learning","optional":false},{"name":"Megatron-LM","optional":false},{"name":"Multimodal AI","optional":false},{"name":"Post-training","optional":false},{"name":"Reinforcement Learning","optional":false},{"name":"SGLang","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"vLLM","optional":false},{"name":"CUDA","optional":true},{"name":"CUDA Toolkit","optional":true},{"name":"CUTLASS","optional":true},{"name":"GRPO","optional":true},{"name":"Mixture of Experts","optional":true},{"name":"NCCL","optional":true},{"name":"NVLink","optional":true},{"name":"PPO","optional":true},{"name":"TensorRT","optional":true},{"name":"Triton","optional":true}],"status":"live","first_seen_at":"2026-02-17T11:28:04Z","employer_posted_date":"2026-08-27","last_verified_at":"2026-09-24T08:38:16Z","board_verified":true,"closed_at":null,"days_open":218,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":218},"description":"About the Role\nAs a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training - spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.\nThis role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.\nIn This Role, You Will\nDesign, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.\nProfile and optimize performance across the full-stack - model implementation, parallelism strategies, communication libraries, and GPU kernels - to improve throughput, latency, memory efficiency, hardware utilization, and cost.\nInvestigate numerical correctness and low-precision issues in distributed training and inference, including train-inference consistency for reinforcement learning.\nImprove the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.\nBuild supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.\nContribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.\nMinimum Qualifications\n3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.\nHands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).\nExperience in at least two of the following areas:Performance, efficiency, and scalability of multi-GPU, multi-node workloads\nNumerical correctness or low precision\nStability, reliability, or fault tolerance\nPost-training algorithm recipes and orchestration infrastructure for large training runs\nMultimodal training or inference, including vision-language models and multimodal generation\nAgent infrastructure, including sandboxes, harnesses, and eval systems\nBuilding and maintaining open-source projects widely adopted in industry and academia\n\nPreferred Qualifications\nFamiliarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.\nExperience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).\nKey open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).\nGPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).\nExperience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).\nTop-tier publications in ML systems or other systems fields.\nEven if you don't meet every qualification above, we encourage you to apply - we care most about demonstrated ability to build and reason about large-scale systems.\nAbout RadixArk\nRadixArk builds open-source and production infrastructure for large language models and multimodal post-training. Our systems - including Miles, an enterprise-grade reinforcement learning training framework, and SGLang, a widely deployed high-performance LLM inference engine - power distributed post-training across clusters of 10k-100k+ GPUs.\nCompensation\nDepending on background, skills, and experience, the expected annual salary range for this position is $200,000 to $400,000, plus equity.\nBenefits include a 401(k) plan and unlimited PTO.\nRadixArk sponsors employment visas (e.g., H-1B, O-1) for eligible candidates.\nEqual Opportunity\nRadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.","description_format":"text","description_chars":4317,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Unlimited PTO"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["AI Infrastructure"],"lifecycle":[{"event":"open","at":"2026-09-12T15:26:28Z"}],"liveness":{"score":8,"band":"cold","label":"Long shot","p_open":1,"p_active":0.27,"p_room":0.28,"age_days":218,"expected_fill_days":31,"reasons":["conf:1","stale_co","ghost","win:tail","crowd:"],"computed_at":"2026-09-24T05:45:00Z"},"pay":{"stated_usd_annual":400000,"is_top_pay":true},"html_url":"https://alion.io/job/radixark-member-of-technical-staff-training","json_url":"https://alion.io/job/radixark-member-of-technical-staff-training.json","meta":{"generated_at":"2026-09-24T09:38:40Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}