489,618open jobs
16,287companies
72,295added this week
Browse all
Location
In office (Seoul)
Seniority
Middle
Overview
Company
Impact
Profile match
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. 

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence  for every enterprise.

About the Role

The compiler plays a central role in FuriosaAI’s mission to build high-performance, energy-efficient AI systems. As modern deep learning models continue to evolve in structure and execution behavior, the compiler’s abstractions and optimization strategies must advance in parallel. Our goal is to build a production compiler that delivers high performance across a broad range of workloads.

The middle-end tackles key optimization problems that determine how tensor programs execute on FuriosaAI hardware. These involve a large, tightly coupled space of decisions, including tiling, fusion, data layout, hardware mapping, and scheduling. We formulate these problems at the right level of abstraction for the TCP architecture and optimize them jointly.

In this role, you will identify and take ownership of compiler problems in real models and production workloads. You will turn those insights into general compiler solutions. You will design IR abstractions, optimization passes, program analyses, or planning algorithms, then integrate and validate them in the compiler.

Key Responsibilities

  • IR, Semantics, and Lowering. Design and evolve intermediate representations, semantics, and lowering transformations that connect tensor programs to efficient hardware execution.

  • Global Optimization and Execution Planning. Develop performance models, search and solver techniques, and optimization algorithms that produce efficient execution plans across the program.

  • Scheduling and Resource Management. Optimize execution schedules and hardware resource allocation.

  • Program Analysis and Verification. Build dependence and alias analyses, race detectors, and compiler checks that validate transformations and catch correctness issues early.

  • Hardware/Software Co-design. Contribute compiler and workload insights to future hardware generations, and evaluate proposed architectures before implementation in silicon.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Mathematics, or a related technical field.

  • Foundational knowledge of compiler design and optimization passes.

  • Ability to abstract complex hardware/software constraints into logical algorithms and solve problems through rigorous reasoning.

Preferred Qualifications

  • Master's or PhD in Programming Languages, Compilers, Program Analysis, or related fields.

  • Research or industry experience with compiler infrastructures like LLVM or MLIR.

  • Experience developing code generators, instruction schedulers, or high-performance kernels for specialized accelerators (NPU, GPU, etc).

  • Experience applying search, constraint solving, or mathematical optimization (ILP, SMT, dynamic programming, heuristic search) to compilation, scheduling, or resource allocation problems.

  • Experience building analytical or learned performance models for hardware.

  • Experience applying program analysis techniques to optimize performance or ensure program correctness.

  • Experience with functional programming languages, particularly in the design of large-scale software systems.

Why Join FuriosaAI

The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.

With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.

At Furiosa, you will:

Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.

Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.

Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.

Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.

Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition. 

Contact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
489,618 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$243k – $297k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
AI/ML
vLLM
CUDA Toolkit
AI Agents
SGLang
PyTorch
CUDA
ROCm
MLIR
DevOps
Prometheus
Kubernetes
Grafana
Apply
$152k – $242k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • United States
AI/ML
CUDA Toolkit
CUDA
MLIR
DevOps
HPC
Apply
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
C
C++
C
MPI
C++
TensorFlow C++
PyTorch C++
LLVM
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
JAX
TensorRT
TensorFlow
PyTorch
CUDA
Triton
NCCL
ROCm
MLIR
CUTLASS
DevOps
GCP
Azure
AWS
HPC
Apply
$80k – $107k per year • Remote • 7+ years exp • Bachelor's Degree
C++
C++
LLVM
AI/ML
vLLM
CUDA Toolkit
TensorRT
CUDA
Triton
NCCL
MLIR
CUTLASS
Apply
$100k – $150k per year • Remote • 6+ years exp • Bachelor's Degree
C++
C++
PyTorch C++
LLVM
AI/ML
vLLM
CUDA Toolkit
JAX
TensorRT
PyTorch
CUDA
Triton
NCCL
MLIR
CUTLASS
DevOps
HPC
Apply
$34k – $83k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Seoul
Apply
$36k – $93k per year (Estimated) • In office • Seoul
Rust
C++
Apply
$35k – $102k per year (Estimated) • Remote/Hybrid • Hwaseong
Rust
C++
Apply
$35k – $102k per year (Estimated) • In office • Seoul
C
C++
C
Embedded C
DevOps
RTOS
HPC
Apply
$35k – $84k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
AI/ML
CUDA Toolkit
Quantization
PyTorch
LLM
CUDA
Triton
Hugging Face
KV Cache
DevOps
Git
Apply
Remote • Part-Time • Seoul
Apply
In office • Full-Time • 3+ years exp • Seoul
Apply
$36k – $76k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Seoul
DevOps
GCP
Apply
$51k – $117k per year (Estimated) • In office • Seoul
Python
AI/ML
LLM
DevOps
AWS
Kubernetes
Apply
$32k – $73k per year (Estimated) • In office • Seoul
Analytics
Tableau
Power BI
Microsoft Excel
Management
Agile
Apply
See all jobs
This is one of many
489,618 more open roles from verified company boards, updated every day.