About FuriosaAI
FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.
Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.
Software Engineer, LLM Performance & Evaluation
Location: Seoul, South Korea (Hybrid)
About the Job
FuriosaAI is seeking a Software Engineer to own the measurement and continuous tracking of LLM inference performance and model accuracy on our NPUs. You will build the benchmarks, evaluation pipelines, and analysis tools that show how individual inference features perform, how accurately each supported model runs, and how both change as Furiosa-LLM (FLM) and our software stack evolve.
Your work will give engineering teams reliable evidence for optimization priorities, model support, and release decisions. Working with the MLSys, Inference Engine, Compiler, and LLM Serving teams, you will establish reproducible baselines, investigate regressions, and make performance-accuracy trade-offs clear across models, workloads, and configurations.
Key Responsibilities
- Design and maintain benchmarks that isolate the impact of inference features such as continuous batching, prefix caching, speculative decoding, and distributed inference. Measure feature interactions and end-to-end behavior across representative input/output lengths, concurrency levels, and deployment configurations.
- Measure time to first token (TTFT), inter-token and end-to-end latency distributions, throughput, memory use, and accelerator utilization. Compare configurations under explicit workload and latency constraints, and quantify improvements against controlled baselines.
- Own model-level accuracy evaluation for supported LLMs, selecting representative datasets and task-appropriate metrics. Compare NPU results with trusted reference implementations and prior releases, and characterize the effects of quantization, numerical precision, and inference optimizations on model quality.
- Build reliable automation for benchmarks and accuracy evaluations, including scheduled runs and change-triggered checks. Version models, datasets, prompts, scoring logic, and execution configurations so results remain reproducible and comparable over time.
- Maintain dashboards and reports that track performance and accuracy by model, inference feature, hardware configuration, and software version. Define regression thresholds and release acceptance criteria with engineering owners, and integrate evaluation checks into CI and release workflows with the Build & Release team.
- Investigate regressions using profiling, controlled experiments, and targeted reproductions. Distinguish implementation changes from workload variation, evaluation errors, and infrastructure noise; work with component owners to identify causes and verify fixes.
- Improve measurement quality through repeated runs, statistical analysis, and validation of datasets and scoring methods. Maintain a focused regression suite and broader periodic evaluations that balance coverage, execution cost, and feedback speed.
- Communicate findings, supported operating ranges, and performance-accuracy trade-offs through clear technical reports. Use the evidence to recommend optimization priorities and keep benchmark methodology and model evaluation results documented.
Minimum Qualifications
- Strong Python programming skills and experience building maintainable automation, data pipelines, or engineering tools.
- Hands-on experience measuring and analyzing ML inference or complex systems performance, including profiling, latency/throughput analysis, and regression investigation.
- Understanding of transformer-based LLM inference, including prefill/decode, batching, KV-cache behavior, and the effects of workload and numerical precision on performance and model outputs.
- Practical experience evaluating model accuracy or numerical correctness using PyTorch, Hugging Face Transformers, or comparable tools, with the ability to investigate discrepancies against a reference implementation.
- Sound experimental design and data analysis skills, including controlled comparisons, variability analysis, and reproducible reporting.
- Ability to communicate quantitative findings clearly and collaborate across engineering teams to drive issues through resolution.
Preferred Qualifications
- Experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, including their benchmarking and profiling tools.
- Experience building LLM evaluation suites with tools such as lm-evaluation-harness or comparable frameworks, including dataset preparation, prompt configuration, and scoring validation.
- Familiarity with quantization, mixed precision, speculative decoding, or distributed inference, and their implications for performance and accuracy.
- Experience benchmarking GPU, NPU, or TPU systems and analyzing memory bandwidth, kernel execution, or communication bottlenecks.
- Experience operating automated evaluation workloads in Linux, container, or cluster environments and integrating results with CI, experiment tracking, or dashboards.
- Ability to read and instrument Rust or C++ code, or contributions to open-source inference, benchmarking, or evaluation projects.
Why Join FuriosaAI
The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.
With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.
At Furiosa, you will:
Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.
Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.
Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.
Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.
Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition.

