About FuriosaAI
FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.
Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.
Software Engineer, AI Model Enablement & Inference
Location: Seoul, South Korea (Hybrid)
About the Job
FuriosaAI is seeking a Software Engineer to join our MLSys team. We apply deep expertise in model architectures and algorithms to build efficient, production-ready implementations of frontier LLMs tailored to our Tensor Contraction Processor (TCP) architecture.
Our work makes these models available through FuriosaAI's software stack, enabling developers to deploy them efficiently on our NPUs with Furiosa-LLM (FLM). You will develop correct, efficient model kernels in Tensor Contraction Language (TCL), FuriosaAI's Python-based DSL for tensor computations, and work closely with the Inference Engine and Compiler teams on kernel implementation strategies and FLM integration.
Key Responsibilities
- Analyze model architectures, algorithms, and reference implementations to identify new model features and document their implementation requirements and trade-offs. Work with the Inference Engine and Compiler teams to agree on model implementation and integration strategies.
- Design, implement, and optimize model-specific TCL kernels, including attention and mixture-of-experts computations, for high performance and efficient use of NPU resources.
- Integrate new models and TCL kernels into Furiosa-LLM to enable correct, efficient inference.
- Improve the process for adding new models by building reusable analysis and integration tools, automating validation and benchmarking, and documenting repeatable workflows.
- Validate kernel and model correctness on NPUs against reference implementations, investigate numerical differences, and build regression tests for supported configurations.
- Evaluate techniques from frameworks such as vLLM and SGLang, adapt and apply relevant approaches to TCL kernels and Furiosa-LLM integration, and document validated findings.
- Continuously study state-of-the-art models to deepen expertise in model architectures and algorithms. Share insights across the company and provide technical guidance to teams on model capabilities, architectural trade-offs, and inference requirements.
Minimum Qualifications
- Deep understanding of transformer-based LLMs, including attention variants, mixture-of-experts architectures, and KV-cache behavior.
- Strong Python programming skills and hands-on experience reading, modifying, and debugging model implementations in PyTorch or a comparable framework.
- Hands-on experience implementing, debugging, and optimizing tensor operations or accelerator kernels, with an understanding of compute, memory, and numerical correctness.
- Understanding of LLM inference performance, including prefill/decode, batching, and latency-throughput trade-offs, with experience in quantitative performance evaluation.
- Ability to read and reason about Rust or C++ code when working on inference systems.
- Clear technical communication and cross-team collaboration skills, including the ability to turn analysis into actionable engineering decisions.
Preferred Qualifications
- Experience bringing up or optimizing models on GPUs, NPUs, TPUs, or other AI accelerators.
- Familiarity with the internals of vLLM, SGLang, TensorRT-LLM, or similar inference frameworks, including scheduling, caching, or model parallelism.
- Familiarity with ML compilers and optimizations such as fusion, tiling, and scheduling.
- Experience developing and optimizing accelerator kernels using CUDA, Triton, or a tensor DSL.
- Experience developing software in Rust, building model validation or benchmarking tools, or contributing to open-source model and inference projects.
Why Join FuriosaAI
The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.
With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.
At Furiosa, you will:
Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.
Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.
Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.
Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.
Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition.

