686,162open jobs
40,351companies
95,115added this week
Browse all
Salary
$68k – $176k per year (Estimated)
Location
Remote/Hybrid (Seoul, South Korea)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. 

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence  for every enterprise.

Production Engineer, LLM Serving

Location: Seoul, South Korea (Hybrid)

About the Job

Owns the reliability, observability, and service-level performance of production LLM serving services on FuriosaAI's RNGD NPUs. You will improve latency, throughput, and resource efficiency through workload-driven measurement, bottleneck analysis, and software engineering.

Our stack uses llm-d and Furiosa-LLM, with Istio for traffic management in Kubernetes-based deployments and Mooncake for L3 KV-cache storage. You will own the serving service layer, collaborating with infrastructure, inference engine, runtime, and compiler teams on improvements that span their components.

Key Responsibilities

  • Operates and improves RNGD-based LLM serving services, defining SLIs/SLOs and production readiness criteria to guide reliability and capacity decisions.
  • Builds metrics, logs, traces, dashboards, and actionable alerts across the serving request path, tracking availability, errors, time to first token (TTFT), inter-token and end-to-end latency, throughput, and NPU utilization.
  • Designs benchmarks and load tests reflecting production input/output lengths, concurrency, and traffic patterns. Uses latency distributions, throughput, and utilization to establish baselines, detect regressions, and plan serving capacity.
  • Diagnoses bottlenecks across routing, queueing, batching, caching, networking, and inference execution. Improves service components and configurations to increase throughput and resource efficiency while meeting latency SLOs.
  • Improves deployment topologies, autoscaling, health checks, graceful degradation, and Istio traffic policies for serving workloads, partnering with the infrastructure team on shared Kubernetes and service mesh capabilities.
  • Automates serving deployments and validates changes through performance tests, canary rollouts, and rollback procedures, collaborating with the build and release team on release artifacts.
  • Participates in on-call and incident response, from mitigation to root-cause analysis and preventive improvements. Builds runbooks and recovery automation to reduce operational toil.

Minimum Qualifications

  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience, with 3+ years developing and operating production services.
  • Strong programming skills in Rust, C++, Python, or Go, with the ability to read, instrument, and improve existing production code.
  • Hands-on experience operating services in Kubernetes and solid understanding of Linux, containers, networking, concurrency, and distributed systems.
  • Understanding of LLM serving concepts, including prefill/decode, batching, KV-cache management, and request scheduling, and their impact on performance.
  • Experience profiling production services and analyzing telemetry to identify bottlenecks across the request path and validate improvements in tail latency, throughput, and resource efficiency.
  • Experience building observability systems and using metrics, logs, and traces to diagnose production incidents, restore services, and implement preventive improvements.
  • Ability to communicate quantitative findings and collaborate across teams to deliver measurable production improvements.

Preferred Qualifications

  • Experience operating LLM serving systems using llm-d, vLLM, SGLang, TensorRT-LLM, or similar technologies.
  • Experience optimizing request routing, load balancing, queueing, or autoscaling, and translating workload measurements into capacity plans or serving-cost estimates.
  • Experience with inference optimizations such as continuous batching, prefix KV-cache reuse, speculative decoding, distributed inference, or kernel optimization.
  • Experience with Istio traffic management, infrastructure as code, or GitOps, including automated rollout and rollback workflows.
  • Experience with Prometheus, Grafana, OpenTelemetry, and reliability practices such as SLO design, error budgets, and post-incident reviews.
  • Experience operating GPU/NPU workloads in on-premises or cloud environments and diagnosing issues across drivers, runtime, orchestration, and applications.
  • Experience developing software in Rust.

Why Join FuriosaAI

The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.

With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.

At Furiosa, you will:

Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.

Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.

Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.

Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.

Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition. 

Contact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
686,162 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$68k – $196k per year (Estimated) • In office • Internship • 2+ years exp • Des Moines
Python
JavaScript
SQL
C#
C++
C#
.NET
DevOps
SOAP
Management
Agile
Apply
$19k – $47k per year (Estimated) • In office • Full-Time • Kuala Lumpur
Python
Go
DevOps
Terraform
GitHub Actions
Datadog
Prometheus
CI/CD
GitOps
AWS
Kubernetes
Grafana
Amazon EKS
Incident Management
DNS
Apply
$94k – $232k per year (Estimated) • In office • Full-Time
Python
Databases
Databricks
AI/ML
Claude Code
AI Agents
LLM
RAG
Time Series Forecasting
Feature Store
Agentic Workflows
Machine Learning
DevOps
CI/CD
AWS
Vector
Apply
Remote/Hybrid
Python
AI/ML
Diffusion Models
PyTorch
Synthetic Data
World Models
Robotics
NVIDIA Drive
CARLA
Sensor Fusion
Sim-to-Real
Apply
Remote/Hybrid • Bachelor's Degree
Python
Java
AI/ML
LLM
DevOps
GitLab CI
CI/CD
SLI/SLO/SLA
Cybersecurity
GDPR
Apply
$61k – $176k per year (Estimated) • Remote/Hybrid • Seoul
Python
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Mixture of Experts
CUDA
Triton
TPU
KV Cache
Apply
$76k – $205k per year (Estimated) • Remote/Hybrid • Seoul
Python
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
Quantization
SGLang
TensorRT
TensorRT-LLM
Transformers
PyTorch
LLM
Hugging Face
TPU
Speculative Decoding
KV Cache
DevOps
Linux
Apply
$50k – $144k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
C++
Bash
Cython
Rust
PyO3
C++
CMake
Cython
Manylinux
PyBind11
DevOps
GitHub Actions
CI/CD
Git
Docker
Ubuntu
Bazel
CentOS Stream
Apply
$63k – $156k per year (Estimated) • In office • Seoul
Rust
C++
Apply
$60k – $172k per year (Estimated) • Remote/Hybrid • Hwaseong
Rust
C++
DevOps
Linux
Apply
$35k – $65k per year (Estimated) • Remote/Hybrid • Internship • Bachelor's Degree • Seoul
Apply
$55k – $118k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Seoul
Management
Agile
Marketing
Salesforce
Apply
$72k – $194k per year (Estimated) • In office • Seoul
Python
Go
JavaScript
Kotlin
TypeScript
Node JS
AI/ML
Cursor
Claude
Claude Code
Model Context Protocol
vLLM
LLM
Triton
OpenAI Codex
Frontend
React.js
DevOps
GCP
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Apply
$100k – $250k per year (Estimated) • In office • 5+ years exp • Seoul
DevOps
SLI/SLO/SLA
Marketing
Salesforce
Apply
$31k – $76k per year (Estimated) • In office • Contractor • 2+ years exp • Bachelor's Degree • Seoul
Apply
See all jobs
This is one of many
686,162 more open roles from verified company boards, updated every day.