686,755open jobs
40,412companies
95,106added this week
Browse all
Salary
$61k – $176k per year (Estimated)
Location
Remote/Hybrid (Seoul, South Korea)
Overview
Company
Impact
Profile match
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. 

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence  for every enterprise.

Software Engineer, AI Model Enablement & Inference

Location: Seoul, South Korea (Hybrid)

About the Job

FuriosaAI is seeking a Software Engineer to join our MLSys team. We apply deep expertise in model architectures and algorithms to build efficient, production-ready implementations of frontier LLMs tailored to our Tensor Contraction Processor (TCP) architecture.

Our work makes these models available through FuriosaAI's software stack, enabling developers to deploy them efficiently on our NPUs with Furiosa-LLM (FLM). You will develop correct, efficient model kernels in Tensor Contraction Language (TCL), FuriosaAI's Python-based DSL for tensor computations, and work closely with the Inference Engine and Compiler teams on kernel implementation strategies and FLM integration.

Key Responsibilities

  • Analyze model architectures, algorithms, and reference implementations to identify new model features and document their implementation requirements and trade-offs. Work with the Inference Engine and Compiler teams to agree on model implementation and integration strategies.
  • Design, implement, and optimize model-specific TCL kernels, including attention and mixture-of-experts computations, for high performance and efficient use of NPU resources.
  • Integrate new models and TCL kernels into Furiosa-LLM to enable correct, efficient inference.
  • Improve the process for adding new models by building reusable analysis and integration tools, automating validation and benchmarking, and documenting repeatable workflows.
  • Validate kernel and model correctness on NPUs against reference implementations, investigate numerical differences, and build regression tests for supported configurations.
  • Evaluate techniques from frameworks such as vLLM and SGLang, adapt and apply relevant approaches to TCL kernels and Furiosa-LLM integration, and document validated findings.
  • Continuously study state-of-the-art models to deepen expertise in model architectures and algorithms. Share insights across the company and provide technical guidance to teams on model capabilities, architectural trade-offs, and inference requirements.

Minimum Qualifications

  • Deep understanding of transformer-based LLMs, including attention variants, mixture-of-experts architectures, and KV-cache behavior.
  • Strong Python programming skills and hands-on experience reading, modifying, and debugging model implementations in PyTorch or a comparable framework.
  • Hands-on experience implementing, debugging, and optimizing tensor operations or accelerator kernels, with an understanding of compute, memory, and numerical correctness.
  • Understanding of LLM inference performance, including prefill/decode, batching, and latency-throughput trade-offs, with experience in quantitative performance evaluation.
  • Ability to read and reason about Rust or C++ code when working on inference systems.
  • Clear technical communication and cross-team collaboration skills, including the ability to turn analysis into actionable engineering decisions.

Preferred Qualifications

  • Experience bringing up or optimizing models on GPUs, NPUs, TPUs, or other AI accelerators.
  • Familiarity with the internals of vLLM, SGLang, TensorRT-LLM, or similar inference frameworks, including scheduling, caching, or model parallelism.
  • Familiarity with ML compilers and optimizations such as fusion, tiling, and scheduling.
  • Experience developing and optimizing accelerator kernels using CUDA, Triton, or a tensor DSL.
  • Experience developing software in Rust, building model validation or benchmarking tools, or contributing to open-source model and inference projects.

Why Join FuriosaAI

The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.

With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.

At Furiosa, you will:

Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.

Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.

Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.

Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.

Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition. 

Contact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
686,755 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$68k – $196k per year (Estimated) • In office • Internship • 2+ years exp • Des Moines
Python
JavaScript
SQL
C#
C++
C#
.NET
DevOps
SOAP
Management
Agile
Apply
$94k – $232k per year (Estimated) • In office • Full-Time
Python
Databases
Databricks
AI/ML
Claude Code
AI Agents
LLM
RAG
Time Series Forecasting
Feature Store
Agentic Workflows
Machine Learning
DevOps
CI/CD
AWS
Vector
Apply
Remote/Hybrid
Python
AI/ML
Diffusion Models
PyTorch
Synthetic Data
World Models
Robotics
NVIDIA Drive
CARLA
Sensor Fusion
Sim-to-Real
Apply
Remote/Hybrid • Bachelor's Degree
Python
Java
AI/ML
LLM
DevOps
GitLab CI
CI/CD
SLI/SLO/SLA
Cybersecurity
GDPR
Apply
$121k – $267k per year (Estimated) • In office • Contractor • Bachelor's Degree • Singapore
Python
JavaScript
Java
TypeScript
C#
Node JS
C#
.NET
Databases
PostgreSQL
pgvector
Pinecone
Azure Cosmos DB
AI/ML
Copilot
AutoGen
LangChain
Claude
ChatGPT
LlamaIndex
Embeddings
Prompt Engineering
Function Calling
Semantic Kernel
CrewAI
RAG
Hallucination
OpenAI
Human-in-the-Loop
Red Teaming
LLM Guardrails
Tool Use
Machine Learning
Frontend
React.js
Mobile
Clean Architecture
DevOps
Rest API
Azure DevOps
GitHub Actions
Azure
CI/CD
Platform Engineering
Vector
GitHub
Cybersecurity
Least Privilege
Microsoft Entra ID
Management
SharePoint
Agile
Apply
$68k – $176k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
C++
AI/ML
vLLM
SGLang
TensorRT
TensorRT-LLM
LLM
Speculative Decoding
KV Cache
DevOps
Istio
OpenTelemetry
Prometheus
GitOps
Kubernetes
Grafana
Service Mesh
SLI/SLO/SLA
Linux
Apply
$76k – $205k per year (Estimated) • Remote/Hybrid • Seoul
Python
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
Quantization
SGLang
TensorRT
TensorRT-LLM
Transformers
PyTorch
LLM
Hugging Face
TPU
Speculative Decoding
KV Cache
DevOps
Linux
Apply
$50k – $144k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
C++
Bash
Cython
Rust
PyO3
C++
CMake
Cython
Manylinux
PyBind11
DevOps
GitHub Actions
CI/CD
Git
Docker
Ubuntu
Bazel
CentOS Stream
Apply
$63k – $156k per year (Estimated) • In office • Seoul
Rust
C++
Apply
$60k – $172k per year (Estimated) • Remote/Hybrid • Hwaseong
Rust
C++
DevOps
Linux
Apply
$35k – $65k per year (Estimated) • Remote/Hybrid • Internship • Bachelor's Degree • Seoul
Apply
$55k – $118k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Seoul
Management
Agile
Marketing
Salesforce
Apply
$72k – $194k per year (Estimated) • In office • Seoul
Python
Go
JavaScript
Kotlin
TypeScript
Node JS
AI/ML
Cursor
Claude
Claude Code
Model Context Protocol
vLLM
LLM
Triton
OpenAI Codex
Frontend
React.js
DevOps
GCP
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Apply
$100k – $250k per year (Estimated) • In office • 5+ years exp • Seoul
DevOps
SLI/SLO/SLA
Marketing
Salesforce
Apply
$31k – $76k per year (Estimated) • In office • Contractor • 2+ years exp • Bachelor's Degree • Seoul
Apply
See all jobs
This is one of many
686,755 more open roles from verified company boards, updated every day.