368,746open jobs
9,444companies
47,506added this week
Browse all
Location
In office (Seoul)
Seniority
Senior · 3+ years exp
Overview
Company
Impact
Profile match
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. 

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence  for every enterprise.

About the Role

Designs and implements the low-level runtime stack that drives FuriosaAI's NPU hardware to its theoretical limits - from device driver interfaces and DMA-based I/O to kernel execution scheduling, multi-node inference, and embedded firmware.

Key Responsibilities

  • Develops the low-level runtime responsible for DMA-based I/O operations and kernel execution scheduling, maximizing inference throughput while minimizing end-to-end latency.

  • Builds and optimizes asynchronous execution pipelines that orchestrate data movement and compute across the NPU hardware.

  • Enables multi-node inference by implementing foundational communication primitives, including RDMA-based data transfer for low-latency, high-bandwidth inter-node operations.

  • Develops embedded firmware (PERT) that runs on the NPU's integrated ARM core, managing on-device scheduling, synchronization, and hardware resource control.

  • Profiles and tunes system-level performance across the full runtime stack - from firmware to user-space - to eliminate bottlenecks in real-world inference workloads.

Minimum Qualifications

  • BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience

  • 3+ years of relevant industry experience or equivalent practical experience in systems programming using Rust, C, or C++

  • Solid understanding of computer architecture fundamentals, including memory hierarchy, cache coherency, operating systems, DMA, interrupts, and MMIO

  • Strong communication skills, with the ability to gather requirements and drive technical alignment across teams

Preferred Qualifications

  • Deep expertise in low-latency runtime systems, embedded firmware development, or high-performance I/O - especially in the context of accelerator hardware.

  • Experience designing and implementing low-latency asynchronous execution models and scheduling systems.

  • Experience with DMA engines, scatter-gather I/O, or other zero-copy data transfer mechanisms.

  • Experience developing embedded firmware for ARM-based processors (bare-metal or lightweight RTOS environments).

  • Familiarity with RDMA technologies and high-performance networking for distributed or multi-node systems.

  • Experience with CUDA low-level runtime internals such as CUDA Graphs, stream-based execution, and asynchronous kernel launch optimization.

  • Experience with kernel-level performance optimizations (e.g., Linux kernel modules, eBPF, perf, ftrace).

  • Understanding of deep learning inference workloads and their hardware execution characteristics.

  • Experience with profiling and performance tuning of system software on accelerator or SoC platforms.

Contact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$200k – $288k per year • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Menlo Park
C++
Go
Java
Rust
Databases
PostgreSQL
Redis
Snowflake
AI/ML
AI Agents
DevOps
Ansible
AWS
Azure
CloudFormation
GCP
Kubernetes
Platform Engineering
Terraform
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$71k – $112k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austria
C#
C++
Python
Visual Basic
Apply
$150k per year • In office • Full-Time • New York
C++
Python
Apply
Controls Assurance AVP 10 hours ago
In office • Full-Time • Bachelor's Degree • Gurgaon
C++
COBOL
Java
SQL
Apply
In office • Contractor • Bachelor's Degree • Seoul
AI/ML
Mamba
Multimodal AI
TPU
Apply
In office • Bachelor's Degree • Seoul
AI/ML
Multimodal AI
RAG
AI Agents
DevOps
GitHub
Analytics
A/B Testing
Apply
In office • 3+ years exp • Bachelor's Degree • Seoul
Apply
In office • Seoul
C++
Rust
Apply
Remote/Hybrid • Hwaseong
C++
Rust
Apply
Product Manager 1 day ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Seoul
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Seoul
Python
DevOps
AWS
CI/CD
GCP
Marketing
Salesforce
Apply
In office • Full-Time • Seoul
C++
SystemVerilog
Verilog
Apply
In office • Contractor • Bachelor's Degree • Seoul
AI/ML
Mamba
Multimodal AI
TPU
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.