687,847open jobs
40,453companies
95,608added this week
Browse all
Salary
$76k – $205k per year (Estimated)
Location
Remote/Hybrid (Seoul, South Korea)
Overview
Company
Impact
Profile match
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. 

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence  for every enterprise.

Software Engineer, LLM Performance & Evaluation

Location: Seoul, South Korea (Hybrid)

About the Job

FuriosaAI is seeking a Software Engineer to own the measurement and continuous tracking of LLM inference performance and model accuracy on our NPUs. You will build the benchmarks, evaluation pipelines, and analysis tools that show how individual inference features perform, how accurately each supported model runs, and how both change as Furiosa-LLM (FLM) and our software stack evolve.

Your work will give engineering teams reliable evidence for optimization priorities, model support, and release decisions. Working with the MLSys, Inference Engine, Compiler, and LLM Serving teams, you will establish reproducible baselines, investigate regressions, and make performance-accuracy trade-offs clear across models, workloads, and configurations.

Key Responsibilities

  • Design and maintain benchmarks that isolate the impact of inference features such as continuous batching, prefix caching, speculative decoding, and distributed inference. Measure feature interactions and end-to-end behavior across representative input/output lengths, concurrency levels, and deployment configurations.
  • Measure time to first token (TTFT), inter-token and end-to-end latency distributions, throughput, memory use, and accelerator utilization. Compare configurations under explicit workload and latency constraints, and quantify improvements against controlled baselines.
  • Own model-level accuracy evaluation for supported LLMs, selecting representative datasets and task-appropriate metrics. Compare NPU results with trusted reference implementations and prior releases, and characterize the effects of quantization, numerical precision, and inference optimizations on model quality.
  • Build reliable automation for benchmarks and accuracy evaluations, including scheduled runs and change-triggered checks. Version models, datasets, prompts, scoring logic, and execution configurations so results remain reproducible and comparable over time.
  • Maintain dashboards and reports that track performance and accuracy by model, inference feature, hardware configuration, and software version. Define regression thresholds and release acceptance criteria with engineering owners, and integrate evaluation checks into CI and release workflows with the Build & Release team.
  • Investigate regressions using profiling, controlled experiments, and targeted reproductions. Distinguish implementation changes from workload variation, evaluation errors, and infrastructure noise; work with component owners to identify causes and verify fixes.
  • Improve measurement quality through repeated runs, statistical analysis, and validation of datasets and scoring methods. Maintain a focused regression suite and broader periodic evaluations that balance coverage, execution cost, and feedback speed.
  • Communicate findings, supported operating ranges, and performance-accuracy trade-offs through clear technical reports. Use the evidence to recommend optimization priorities and keep benchmark methodology and model evaluation results documented.

Minimum Qualifications

  • Strong Python programming skills and experience building maintainable automation, data pipelines, or engineering tools.
  • Hands-on experience measuring and analyzing ML inference or complex systems performance, including profiling, latency/throughput analysis, and regression investigation.
  • Understanding of transformer-based LLM inference, including prefill/decode, batching, KV-cache behavior, and the effects of workload and numerical precision on performance and model outputs.
  • Practical experience evaluating model accuracy or numerical correctness using PyTorch, Hugging Face Transformers, or comparable tools, with the ability to investigate discrepancies against a reference implementation.
  • Sound experimental design and data analysis skills, including controlled comparisons, variability analysis, and reproducible reporting.
  • Ability to communicate quantitative findings clearly and collaborate across engineering teams to drive issues through resolution.

Preferred Qualifications

  • Experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, including their benchmarking and profiling tools.
  • Experience building LLM evaluation suites with tools such as lm-evaluation-harness or comparable frameworks, including dataset preparation, prompt configuration, and scoring validation.
  • Familiarity with quantization, mixed precision, speculative decoding, or distributed inference, and their implications for performance and accuracy.
  • Experience benchmarking GPU, NPU, or TPU systems and analyzing memory bandwidth, kernel execution, or communication bottlenecks.
  • Experience operating automated evaluation workloads in Linux, container, or cluster environments and integrating results with CI, experiment tracking, or dashboards.
  • Ability to read and instrument Rust or C++ code, or contributions to open-source inference, benchmarking, or evaluation projects.

Why Join FuriosaAI

The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.

With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.

At Furiosa, you will:

Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.

Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.

Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.

Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.

Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition. 

Contact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
687,847 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$94k – $232k per year (Estimated) • In office • Full-Time
Python
Databases
Databricks
AI/ML
Claude Code
AI Agents
LLM
RAG
Time Series Forecasting
Feature Store
Agentic Workflows
Machine Learning
DevOps
CI/CD
AWS
Vector
Apply
Remote/Hybrid
Python
AI/ML
Diffusion Models
PyTorch
Synthetic Data
World Models
Robotics
NVIDIA Drive
CARLA
Sensor Fusion
Sim-to-Real
Apply
Remote/Hybrid • Bachelor's Degree
Python
Java
AI/ML
LLM
DevOps
GitLab CI
CI/CD
SLI/SLO/SLA
Cybersecurity
GDPR
Apply
Data Scientist 1 day ago
$101k – $193k per year (Estimated) • In office • 4+ years exp • Master's Degree • New York
Python
SQL
Databases
Snowflake
Databricks
AI/ML
LLM
Machine Learning
Apply
Remote/Hybrid
AI/ML
Weights & Biases
MLFlow
Fine-tuning
Computer Vision
ONNX
TensorRT
PyTorch
FSDP
Analytics
A/B Testing
Apply
$61k – $176k per year (Estimated) • Remote/Hybrid • Seoul
Python
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Mixture of Experts
CUDA
Triton
TPU
KV Cache
Apply
$68k – $176k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
C++
AI/ML
vLLM
SGLang
TensorRT
TensorRT-LLM
LLM
Speculative Decoding
KV Cache
DevOps
Istio
OpenTelemetry
Prometheus
GitOps
Kubernetes
Grafana
Service Mesh
SLI/SLO/SLA
Linux
Apply
$50k – $144k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
C++
Bash
Cython
Rust
PyO3
C++
CMake
Cython
Manylinux
PyBind11
DevOps
GitHub Actions
CI/CD
Git
Docker
Ubuntu
Bazel
CentOS Stream
Apply
$63k – $156k per year (Estimated) • In office • Seoul
Rust
C++
Apply
$60k – $172k per year (Estimated) • Remote/Hybrid • Hwaseong
Rust
C++
DevOps
Linux
Apply
$35k – $65k per year (Estimated) • Remote/Hybrid • Internship • Bachelor's Degree • Seoul
Apply
$55k – $118k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Seoul
Management
Agile
Marketing
Salesforce
Apply
$72k – $194k per year (Estimated) • In office • Seoul
Python
Go
JavaScript
Kotlin
TypeScript
Node JS
AI/ML
Cursor
Claude
Claude Code
Model Context Protocol
vLLM
LLM
Triton
OpenAI Codex
Frontend
React.js
DevOps
GCP
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Apply
$100k – $250k per year (Estimated) • In office • 5+ years exp • Seoul
DevOps
SLI/SLO/SLA
Marketing
Salesforce
Apply
$31k – $76k per year (Estimated) • In office • Contractor • 2+ years exp • Bachelor's Degree • Seoul
Apply
See all jobs
This is one of many
687,847 more open roles from verified company boards, updated every day.