824,647open jobs
53,159companies
134,960added this week
Browse all
Salary
≈ $60k – $157k per year (Estimated)
Location
In office (Seongnam)
Seniority
Senior · 5+ years exp

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Mar 18, 2026. Rebellions scores C on the Alion truth index.

Overview
Company
Impact
Profile match
Rebellions is a South Korean semiconductor company founded in 2020 by former Wall Street quantitative engineers and IBM chip designers to build accelerators specialised for AI inference rather than training. Its ATOM processor targets efficiency per watt in data centre inference, and the follow-on REBEL design uses chiplet packaging to scale to large language model workloads. The company merged with Sapeon Korea, the AI chip unit of the SK group, to consolidate the domestic effort against foreign GPUs, becoming South Korea's first AI semiconductor unicorn and working closely with SK Telecom, KT and Samsung Foundry.

We are seeking a highly skilled NPU Runtime Software Engineer to join our team. You will be responsible for designing and implementing the software layer that bridges high-level ML frameworks with our proprietary NPU hardware - enabling the next generation of real-time AI applications. Your work will ensure that state-of-the-art models - with a heavy focus on LLMs - run with industry-leading efficiency, low latency, and high throughput. You will sit at the intersection of compilers, system drivers, and distributed inference frameworks, spanning the full runtime stack from graph execution and compiler integration to inference serving.

Responsibilities and Opportunities

  • Design and implement the RBLN runtime module that interfaces with compiler and driver components, including the graph executor and runtime APIs, to enable ML model deployment through the RBLN SDK
  • Architect and maintain native PyTorch execution support within the runtime, including torch.compile integration and RBLN compiler toolchains, to enable seamless NPU acceleration with minimal user-side code changes
  • Design and implement a user-facing profiler that provides actionable performance insights, delivered as part of the RBLN SDK
  • Develop and extend vLLM to enhance inference performance on NPUs, including support for key vLLM features such as advanced memory management, parallelism, and dynamic batching
  • Design and optimize distributed inference across multi-NPU setups, including collective communication operations (CCL) to support various parallelism strategies
  • Conduct benchmarking and profiling to evaluate runtime system performance and implement optimizations to improve overall system efficiency
  • Collaborate with ML engineers and infrastructure teams to deploy and scale inference services

Key Qualifications

  • Over 5 years of experience in software engineering, with significant work on ML frameworks, inference runtimes, or AI accelerator toolchains in production environments
  • Bachelor's degree or higher in Computer Science, Electrical Engineering, or a related field
  • Strong proficiency in C++ and Python
  • Strong understanding of deep learning fundamentals and LLM architectures, including Transformer-based models, generative AI, and inference optimization techniques
  • Hands-on experience with LLM serving frameworks (e.g., vLLM, TensorRT-LLM)
  • Solid understanding of model optimization techniques (tensor parallelism, KV cache optimizations, memory-efficient execution)
  • Familiarity with system software components, including compilers, runtimes, drivers, and firmware
  • Familiarity with hardware acceleration (GPUs, NPUs, TPUs) and efficient memory management techniques
  • Strong debugging and performance profiling skills for high-throughput inference environments
  • Ability to work effectively across compiler, driver, and ML engineering teams
  • Excellent written and verbal communication skills

Ideal Qualifications

  • Practical experience with AI accelerator runtimes and driver APIs (e.g., GPUs)
  • Direct contribution or production experience with ML frameworks and serving systems such as PyTorch, vLLM, SGLang, TensorRT, and TensorRT-LLM
  • Understanding of torch.compile and graph optimizations
  • Strong understanding of operating systems, resource management, and high-performance computing concepts
  • Advanced proficiency in modern C++ for developing efficient, high-performance systems
  • Experience with multithreading and parallel programming
  • Experience deploying LLMs in distributed environments

전형절차

  • 서류전형 > On-line 인터뷰 > On-site 인터뷰(과제 포함) > Culture-fit인터뷰 > 처우 협의 > 최종 합격
  • 전형절차는 직무별로 다르게 운영될 수 있으며, 일정 및 상황에 따라 변동될 수 있습니다.
  • 전형 일정 및 결과는 지원 시 작성하신 이메일로 개별 안내드립니다.

참고사항

  • 본 공고는 모집 완료 시 조기 마감될 수 있습니다.
  • 지원서 내용 중 허위사실이 있는 경우에는 합격이 취소될 수 있습니다.
  • 채용 및 업무 수행과 관련하여 요구되는 법령 상 자격이 갖추어지지 않은 경우 채용이 제한될 수 있습니다.
  • 보훈 대상자 및 장애인 여부는 채용 과정에서 어떠한 불이익도 미치지 않습니다.
  • 담당 업무 범위는 후보자의 전반적인 경력과 경험 등 제반사정을 고려하여 변경될 수 있습니다. 이러한 변경이 필요할 경우, 최종 합격 통지 전 적절한 시기에 후보자와 커뮤니케이션 될 예정입니다.
  • 채용 관련 문의사항은 아래 메일 주소로 문의바랍니다.
  • [email protected]
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
824,647 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Seongnam
In office • Full-Time • 3+ years exp • Seoul
Python
C
C++
C
ZeroMQ
AI/ML
Physical AI
DevOps
RTOS
WebRTC
WebSockets
CI/CD
Git
Docker
Linux
Robotics
ROS2
Zenoh
Teleoperation
Apply
$37k – $58k per year • Remote (Mexico) • Full-Time • Monterrey
Python
Python
FastAPI
Asyncio
Pydantic
Databases
PostgreSQL
AI/ML
Cursor
Claude Code
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
DeepEval
AWS Bedrock
LLM
OpenAI Agents SDK
AWS Strands Agents
A2A
OCR
Multi-Agent Systems
Tool Use
DevOps
Terraform
Git
AWS
Docker
Kubernetes
Trunk-Based Development
kubectl
Amazon EKS
IAM
QA
Pytest
Apply
≈ $67k – $197k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • London
Java
Kotlin
AI/ML
Machine Learning
DevOps
CI/CD
AWS
Kubernetes
Apply
In office • Full-Time • Riga
Apply
In office • 12+ years exp • Bachelor's Degree
Apply
Staff Engineer I 2 days ago
≈ $58k – $125k per year (Estimated) • Remote (India) • Full-Time • 10+ years exp • Bengaluru
Java
C++
Java
Spring Framework
Spring MVC
Databases
DynamoDB
Amazon Redshift
AI/ML
Prompt Engineering
AI Agents
LLM
RAG
Machine Learning
Frontend
GraphQL
DevOps
Rest API
CI/CD
Jenkins
Git
AWS
Self-Healing
AWS Lambda
Amazon EC2
Amazon S3
Cybersecurity
LDAP
Management
Jira
Agile
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
C++
DevOps
CI/CD
Git
Gerrit
Linux
Wi-Fi
Management
Agile
Scrum
Apply
≈ $80k – $191k per year (Estimated) • In office • Full-Time • Singapore
Python
JavaScript
SQL
DevOps
Rest API
Azure
Analytics
Power BI
Management
SharePoint
Apply
≈ $40k – $74k per year (Estimated) • In office • Contractor • 1+ year exp • Bachelor's Degree • Singapore
Python
JavaScript
TypeScript
AI/ML
Computer Vision
DevOps
Azure
CI/CD
Git
AWS
Docker
Linux
Management
Scrum
Apply
$21k per year (net) • In office • Full-Time • Kazan
Python
1C
Databases
MS SQL
Analytics
Power BI
ETL/ELT
Management
Google Sheets
Apply
≈ $59k – $155k per year (Estimated) • In office • 5+ years exp • Master's Degree • Seongnam
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
Multimodal AI
AI Agents
Diffusers
TensorRT
TensorRT-LLM
Transformers
TensorFlow
PyTorch
LLM
CUDA
Triton
Hugging Face
MLIR
Apache TVM
DevOps
GitHub Actions
CI/CD
Docker
Kubernetes
Buildkite
Apply
In office • Master's Degree • Seongnam
Python
C++
AI/ML
LLM
Apply
In office • Master's Degree • Seongnam
Python
C++
AI/ML
CUDA Toolkit
Computer Vision
Speech Recognition
OpenCL
CUDA
Apply
In office • Bachelor's Degree • Seongnam
Python
Perl
Apply
≈ $60k – $159k per year (Estimated) • In office • 5+ years exp • Master's Degree • Seongnam
C
C++
C
MPI
AI/ML
TPU
NCCL
NVLink
DevOps
HPC
Apply
In office • Full-Time • 3+ years exp • Seongnam
Databases
PostgreSQL
Milvus
pgvector
FAISS
AI/ML
CUDA Toolkit
RAG
CUDA
Apply
In office • Full-Time • Seongnam
Apply
SW검증 팀장 1 day ago
≈ $44k – $94k per year (Estimated) • In office • Full-Time • 3+ years exp • Seongnam
AI/ML
Copilot
Claude
Management
Slack
Jira
QA
Playwright
Postman
Apply
≈ $67k – $181k per year (Estimated) • In office • Full-Time • Seongnam
AI/ML
TensorFlow
PyTorch
Apply
임상팀 팀장 1 day ago
≈ $57k – $119k per year (Estimated) • In office • Full-Time • 7+ years exp • Seongnam
Apply
See all jobs
This is one of many
824,647 more open roles from verified company boards, updated every day.