{"id":1212434,"url":"https://alion.io/job/rebellions-npu-software-engineer-runtime","title":"NPU Software Engineer - Runtime","company":{"id":173,"name":"Rebellions","domain":"rebellions.ai","url":"https://alion.io/company/rebellions","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greeting","truth_index":{"grade":"C","score":66,"open_postings":28,"ghost_share":0.571,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-26T05:45:00Z"}},"role":"Backend","role_family":"Backend","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Seongnam, South Korea"],"countries":["KR"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":60000,"max_usd":157000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":4445},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"C++","optional":false},{"name":"KV Cache","optional":false},{"name":"LLM","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"SGLang","optional":false},{"name":"TensorRT","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"TPU","optional":false},{"name":"Transformers","optional":false},{"name":"vLLM","optional":false}],"status":"live","first_seen_at":"2026-03-18T05:07:21Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-25T08:02:56Z","board_verified":true,"closed_at":null,"days_open":193,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":192},"description":"We are seeking a highly skilled NPU Runtime Software Engineer to join our team. You will be responsible for designing and implementing the software layer that bridges high-level ML frameworks with our proprietary NPU hardware - enabling the next generation of real-time AI applications. Your work will ensure that state-of-the-art models - with a heavy focus on LLMs - run with industry-leading efficiency, low latency, and high throughput. You will sit at the intersection of compilers, system drivers, and distributed inference frameworks, spanning the full runtime stack from graph execution and compiler integration to inference serving.\nResponsibilities and Opportunities\nDesign and implement the RBLN runtime module that interfaces with compiler and driver components, including the graph executor and runtime APIs, to enable ML model deployment through the RBLN SDK\nArchitect and maintain native PyTorch execution support within the runtime, including torch.compile integration and RBLN compiler toolchains, to enable seamless NPU acceleration with minimal user-side code changes\nDesign and implement a user-facing profiler that provides actionable performance insights, delivered as part of the RBLN SDK\nDevelop and extend vLLM to enhance inference performance on NPUs, including support for key vLLM features such as advanced memory management, parallelism, and dynamic batching\nDesign and optimize distributed inference across multi-NPU setups, including collective communication operations (CCL) to support various parallelism strategies\nConduct benchmarking and profiling to evaluate runtime system performance and implement optimizations to improve overall system efficiency\nCollaborate with ML engineers and infrastructure teams to deploy and scale inference services\nKey Qualifications\nOver 5 years of experience in software engineering, with significant work on ML frameworks, inference runtimes, or AI accelerator toolchains in production environments\nBachelor's degree or higher in Computer Science, Electrical Engineering, or a related field\nStrong proficiency in C++ and Python\nStrong understanding of deep learning fundamentals and LLM architectures, including Transformer-based models, generative AI, and inference optimization techniques\nHands-on experience with LLM serving frameworks (e.g., vLLM, TensorRT-LLM)\nSolid understanding of model optimization techniques (tensor parallelism, KV cache optimizations, memory-efficient execution)\nFamiliarity with system software components, including compilers, runtimes, drivers, and firmware\nFamiliarity with hardware acceleration (GPUs, NPUs, TPUs) and efficient memory management techniques\nStrong debugging and performance profiling skills for high-throughput inference environments\nAbility to work effectively across compiler, driver, and ML engineering teams\nExcellent written and verbal communication skills\nIdeal Qualifications\nPractical experience with AI accelerator runtimes and driver APIs (e.g., GPUs)\nDirect contribution or production experience with ML frameworks and serving systems such as PyTorch, vLLM, SGLang, TensorRT, and TensorRT-LLM\nUnderstanding of torch.compile and graph optimizations\nStrong understanding of operating systems, resource management, and high-performance computing concepts\nAdvanced proficiency in modern C++ for developing efficient, high-performance systems\nExperience with multithreading and parallel programming\nExperience deploying LLMs in distributed environments\n전형절차\n서류전형 > On-line 인터뷰 > On-site 인터뷰(과제 포함) > Culture-fit인터뷰 > 처우 협의 > 최종 합격\n전형절차는 직무별로 다르게 운영될 수 있으며, 일정 및 상황에 따라 변동될 수 있습니다.\n전형 일정 및 결과는 지원 시 작성하신 이메일로 개별 안내드립니다.\n참고사항\n본 공고는 모집 완료 시 조기 마감될 수 있습니다.\n지원서 내용 중 허위사실이 있는 경우에는 합격이 취소될 수 있습니다.\n채용 및 업무 수행과 관련하여 요구되는 법령 상 자격이 갖추어지지 않은 경우 채용이 제한될 수 있습니다.\n보훈 대상자 및 장애인 여부는 채용 과정에서 어떠한 불이익도 미치지 않습니다.\n담당 업무 범위는 후보자의 전반적인 경력과 경험 등 제반사정을 고려하여 변경될 수 있습니다. 이러한 변경이 필요할 경우, 최종 합격 통지 전 적절한 시기에 후보자와 커뮤니케이션 될 예정입니다.\n채용 관련 문의사항은 아래 메일 주소로 문의바랍니다. \n","description_format":"text","description_chars":3981,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Processors, MCUs & AI Chips","AI Compute & Inference","AI Chips & Accelerators"],"lifecycle":[{"event":"open","at":"2026-09-25T05:02:56Z"}],"liveness":{"score":5,"band":"cold","label":"Long shot","p_open":1,"p_active":0.189,"p_room":0.28,"age_days":192,"expected_fill_days":39,"reasons":["conf:21","stale_co","velocity","win:tail","crowd:"],"computed_at":"2026-09-26T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/rebellions-npu-software-engineer-runtime","json_url":"https://alion.io/job/rebellions-npu-software-engineer-runtime.json","meta":{"generated_at":"2026-09-27T05:38:12Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4997,"day_limit":5000,"remaining_today":3,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}