{"id":827765,"url":"https://alion.io/job/radixark-member-of-technical-staff-inference-kernel-compiler-communication","title":"Member of Technical Staff — Inference-Kernel, Compiler & Communication","company":{"id":686458,"name":"RadixArk","domain":"radixark.com","url":"https://alion.io/company/radixark","size_band":null,"is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"C","score":55,"open_postings":20,"ghost_share":0.75,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-24T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"staff","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Palo Alto, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":200000,"max":400000,"currency":"USD","period":"year","gross":null,"usd_annual":400000},"salary_estimate":null,"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Apache TVM","optional":false},{"name":"C++","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"GitHub","optional":false},{"name":"HPC","optional":false},{"name":"InfiniBand","optional":false},{"name":"LLM","optional":false},{"name":"MLIR","optional":false},{"name":"NCCL","optional":false},{"name":"NVLink","optional":false},{"name":"Python","optional":false},{"name":"SGLang","optional":false},{"name":"Triton","optional":false},{"name":"XLA","optional":false}],"status":"live","first_seen_at":"2026-02-17T11:34:17Z","employer_posted_date":"2026-08-12","last_verified_at":"2026-09-24T12:13:24Z","board_verified":true,"closed_at":null,"days_open":219,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":218},"description":"About the Role\nRadixArk is seeking a Member of Technical Staff - Kernel / Compiler / Communication to push the limits of performance for frontier AI systems.\nYou will work at the lowest layers of the stack - kernels, runtimes, compilers, and communication libraries - to unlock maximum efficiency from modern accelerators and interconnects.\nThis role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth matter. Your work will directly shape the performance envelope of next-generation AI systems.\nThis is a deeply technical role for engineers who enjoy working close to hardware and solving performance problems that most engineers never encounter.\nRequirements\n5+ years of experience in systems, compiler, or performance engineering\n\nStrong expertise in CUDA or accelerator programming\n\nDeep understanding of GPU architecture and memory hierarchy\n\nExperience writing or optimizing high-performance kernels\n\nStrong background in compilers, runtimes, or code generation\n\nExperience with distributed communication libraries (NCCL, MPI, RCCL, etc.)\n\nSolid knowledge of networking and interconnect technologies\n\nProficiency in C++ and Python\n\nStrong debugging and profiling skills at system level\n\nStrong Plus\nExperience with Triton, TVM, XLA, or MLIR\n\nExperience building compiler passes or IR transformations\n\nFamiliarity with NVLink, InfiniBand, or RDMA\n\nExperience optimizing collective communication at scale\n\nBackground in HPC or performance-critical systems\n\nContributions to kernel/compiler/ML systems open source\n\nExperience scaling workloads to 1000+ GPUs\n\nExperience with mixed-precision or quantized kernels\n\nResponsibilities\nDesign and implement high-performance kernels for AI workloads\n\nOptimize compiler and runtime stacks for ML systems\n\nImprove communication efficiency across large GPU clusters\n\nReduce latency and increase throughput for distributed workloads\n\nProfile and eliminate system bottlenecks across the stack\n\nCollaborate with training and inference teams on performance optimization\n\nDevelop tooling for profiling and performance analysis\n\nContribute to long-term architecture for performance-critical systems\n\nPush the limits of hardware-software co-design\n\nAbout RadixArk\nRadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.\nCompensation\nDepending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\nEqual Opportunity\nRadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.","description_format":"text","description_chars":3344,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["AI Infrastructure"],"lifecycle":[{"event":"open","at":"2026-09-12T15:26:28Z"}],"liveness":{"score":8,"band":"cold","label":"Long shot","p_open":1,"p_active":0.27,"p_room":0.28,"age_days":218,"expected_fill_days":31,"reasons":["conf:1","stale_co","ghost","win:tail","crowd:"],"computed_at":"2026-09-24T05:45:00Z"},"pay":{"stated_usd_annual":400000,"is_top_pay":true},"html_url":"https://alion.io/job/radixark-member-of-technical-staff-inference-kernel-compiler-communication","json_url":"https://alion.io/job/radixark-member-of-technical-staff-inference-kernel-compiler-communication.json","meta":{"generated_at":"2026-09-24T17:17:41Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}