599,985open jobs
30,608companies
86,410added this week
Browse all
Salary
$38k – $100k per year (Estimated)
Location
In office (Seoul)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About the job

FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.

Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.

Key Responsibilities

Cluster Architecture

  • Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades.

  • Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks.

  • Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants.

  • Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.

Networking

  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.

  • Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.

  • Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes.

Reliability & Collaboration

  • Define SLOs for platform-critical systems and lead post-incident hardening.

  • Deliver infrastructure as code with Terraform, Helm, and GitOps.

  • Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.

Qualifications

  • 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production.

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.

  • Proven experience operating large-scale, high-traffic network services in production.

  • Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd.

  • Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.

  • Proficiency with AWS, Terraform, Helm, and Ansible.

  • Programming skills in Go or Python, with the ability to build infrastructure tooling and automation.

  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.

  • Clear written and verbal communication, including the ability to document architectural decisions for other engineers.

Preferred Experience

  • Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.

  • Cilium and eBPF, including kube-proxy replacement or upstream contributions.

  • Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.

  • GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA).

  • High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning.

  • Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations.

  • Contributions to Kubernetes, Cilium, Istio, or other CNCF projects.

Benefits

  • Flexible working hours

  • Daily lunch and dinner provided; unlimited snacks and beverages

  • Supportive and highly collaborative work environment

  • Health check-up support and top-tier equipment/hardware support

  • A front-row seat to the generative AI infrastructure revolution

  • Competitive compensation, startup equity, health insurance, and other benefits.

About FriendliAI

FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads - long-context inference, real-time streaming, and accurate tool calling.

We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
599,985 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$60k – $107k per year (Estimated) • Remote • London
Python
Bash
Databases
MySQL
RabbitMQ
DevOps
Rest API
Apply
$65k – $146k per year (Estimated) • Remote • 5+ years exp • Dubai
Python
AI/ML
Claude
ChatGPT
Model Context Protocol
Structured Outputs
Management
Slack
Confluence
Jira
Power Automate
Google Drive
QuickBooks
Xero
Apply
Backend Developer 1 day ago
In office • Jaipur
Python
Java
Databases
MySQL
PostgreSQL
Redis
DevOps
CI/CD
Apply
$175k – $240k per year • In office • Full-Time • 5+ years exp • Boston • New York
Python
Go
JavaScript
TypeScript
Databases
PostgreSQL
Redis
AI/ML
LangGraph
LangChain
AI Agents
LangSmith
LLM
Frontend
React.js
DevOps
Cloudflare
Management
Monday.com
Apply
$158k – $268k per year (Estimated) • In office • Full-Time • 4+ years exp • Boston • New York
Python
Go
Databases
PostgreSQL
Redis
ClickHouse
AI/ML
LangGraph
LangChain
AI Agents
LangSmith
DevOps
GCP
Azure
AWS
Cloudflare
Management
Monday.com
Apply
$208k – $386k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • San Francisco
AI/ML
vLLM
Multimodal AI
Function Calling
AI Agents
SGLang
TensorRT-LLM
LLM
KV Cache
Tool Use
Apply
In office • Full-Time • 4+ years exp • Bachelor's Degree • Seoul
Python
JavaScript
TypeScript
SQL
Python
FastAPI
Databases
PostgreSQL
ClickHouse
AI/ML
Multimodal AI
Function Calling
AI Agents
LLM
Tool Use
Frontend
GraphQL
Next.js
React.js
DevOps
gRPC
OpenTelemetry
CI/CD
Kubernetes
Cybersecurity
Auth0
Management
Stripe
Apply
$166k – $299k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
Function Calling
NCCL
InfiniBand
Tool Use
DevOps
Terraform
Ansible
Helm
Cilium
Istio
etcd
GitOps
AWS
Kubernetes
Service Mesh
eBPF
Apply
Account Executive 3 months ago
$130k – $274k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
AI/ML
Multimodal AI
LLM
Hugging Face
Edge AI
DevOps
AWS
Apply
In office • Full-Time • 3+ years exp • Seoul
Python
Go
TypeScript
Python
Asyncio
AI/ML
Multimodal AI
LLM
Hugging Face
DevOps
gRPC
Docker
Kubernetes
Apply
$27k – $72k per year (Estimated) • In office • Seoul
Management
ServiceNow
Apply
In office • Part-Time • Seoul
Apply
Remote/Hybrid • 3+ years exp • Seoul
Python
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Qwen
DeepSeek
vLLM
SGLang
Langfuse
LiteLLM
Ollama
Portkey
TGI
Llama
Mistral
LLM
RAG
Kong AI Gateway
MiniMax
OpenAI
Text-to-Speech
LLM Evaluation
LLM Guardrails
KV Cache
Frontend
GraphQL
DevOps
OpenTelemetry
WebRTC
Prometheus
CI/CD
Kubernetes
Grafana
Graylog
IAM
Analytics
ETL/ELT
Apply
$38k – $98k per year (Estimated) • In office • 5+ years exp • Seoul
Mobile
Twilio
Marketing
Salesforce
Apply
$46k – $122k per year (Estimated) • In office • Internship • Seoul
Python
AI/ML
NLP
Apply
See all jobs
This is one of many
599,985 more open roles from verified company boards, updated every day.