About FuriosaAI
FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.
Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.
Production Engineer, LLM Serving
Location: Seoul, South Korea (Hybrid)
About the Job
Owns the reliability, observability, and service-level performance of production LLM serving services on FuriosaAI's RNGD NPUs. You will improve latency, throughput, and resource efficiency through workload-driven measurement, bottleneck analysis, and software engineering.
Our stack uses llm-d and Furiosa-LLM, with Istio for traffic management in Kubernetes-based deployments and Mooncake for L3 KV-cache storage. You will own the serving service layer, collaborating with infrastructure, inference engine, runtime, and compiler teams on improvements that span their components.
Key Responsibilities
- Operates and improves RNGD-based LLM serving services, defining SLIs/SLOs and production readiness criteria to guide reliability and capacity decisions.
- Builds metrics, logs, traces, dashboards, and actionable alerts across the serving request path, tracking availability, errors, time to first token (TTFT), inter-token and end-to-end latency, throughput, and NPU utilization.
- Designs benchmarks and load tests reflecting production input/output lengths, concurrency, and traffic patterns. Uses latency distributions, throughput, and utilization to establish baselines, detect regressions, and plan serving capacity.
- Diagnoses bottlenecks across routing, queueing, batching, caching, networking, and inference execution. Improves service components and configurations to increase throughput and resource efficiency while meeting latency SLOs.
- Improves deployment topologies, autoscaling, health checks, graceful degradation, and Istio traffic policies for serving workloads, partnering with the infrastructure team on shared Kubernetes and service mesh capabilities.
- Automates serving deployments and validates changes through performance tests, canary rollouts, and rollback procedures, collaborating with the build and release team on release artifacts.
- Participates in on-call and incident response, from mitigation to root-cause analysis and preventive improvements. Builds runbooks and recovery automation to reduce operational toil.
Minimum Qualifications
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience, with 3+ years developing and operating production services.
- Strong programming skills in Rust, C++, Python, or Go, with the ability to read, instrument, and improve existing production code.
- Hands-on experience operating services in Kubernetes and solid understanding of Linux, containers, networking, concurrency, and distributed systems.
- Understanding of LLM serving concepts, including prefill/decode, batching, KV-cache management, and request scheduling, and their impact on performance.
- Experience profiling production services and analyzing telemetry to identify bottlenecks across the request path and validate improvements in tail latency, throughput, and resource efficiency.
- Experience building observability systems and using metrics, logs, and traces to diagnose production incidents, restore services, and implement preventive improvements.
- Ability to communicate quantitative findings and collaborate across teams to deliver measurable production improvements.
Preferred Qualifications
- Experience operating LLM serving systems using
llm-d, vLLM, SGLang, TensorRT-LLM, or similar technologies. - Experience optimizing request routing, load balancing, queueing, or autoscaling, and translating workload measurements into capacity plans or serving-cost estimates.
- Experience with inference optimizations such as continuous batching, prefix KV-cache reuse, speculative decoding, distributed inference, or kernel optimization.
- Experience with Istio traffic management, infrastructure as code, or GitOps, including automated rollout and rollback workflows.
- Experience with Prometheus, Grafana, OpenTelemetry, and reliability practices such as SLO design, error budgets, and post-incident reviews.
- Experience operating GPU/NPU workloads in on-premises or cloud environments and diagnosing issues across drivers, runtime, orchestration, and applications.
- Experience developing software in Rust.
Why Join FuriosaAI
The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.
With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.
At Furiosa, you will:
Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.
Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.
Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.
Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.
Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition.

