624,530open jobs
30,918companies
85,628added this week
Browse all
Salary
$166k – $299k per year (Estimated)
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About the job

FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.

Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.

Key Responsibilities

Cluster Architecture

  • Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades.

  • Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks.

  • Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants.

  • Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.

Networking

  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.

  • Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.

  • Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes.

Reliability & Collaboration

  • Define SLOs for platform-critical systems and lead post-incident hardening.

  • Deliver infrastructure as code with Terraform, Helm, and GitOps.

  • Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.

Qualifications

  • 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production.

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.

  • Proven experience operating large-scale, high-traffic network services in production.

  • Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd.

  • Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.

  • Proficiency with AWS, Terraform, Helm, and Ansible.

  • Programming skills in Go or Python, with the ability to build infrastructure tooling and automation.

  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.

  • Clear written and verbal communication, including the ability to document architectural decisions for other engineers.

Preferred Experience

  • Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.

  • Cilium and eBPF, including kube-proxy replacement or upstream contributions.

  • Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.

  • GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA).

  • High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning.

  • Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations.

  • Contributions to Kubernetes, Cilium, Istio, or other CNCF projects.

Benefits

  • Flexible working hours

  • Daily lunch and dinner provided; unlimited snacks and beverages

  • Supportive and highly collaborative work environment

  • Health check-up support and top-tier equipment/hardware support

  • A front-row seat to the generative AI infrastructure revolution

  • Competitive compensation, startup equity, health insurance, and other benefits.

About FriendliAI

FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads - long-context inference, real-time streaming, and accurate tool calling.

We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
624,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$19k – $45k per year (Estimated) • In office • Full-Time • 5+ years exp • Indore
Python
DevOps
Rest API
Terraform
Ansible
CI/CD
Git
Cybersecurity
Zscaler
Apply
Remote/Hybrid
Python
Java
SQL
Groovy
Java
Maven
Spring Boot
Gradle
Apache Tomcat
Databases
Oracle
MS SQL
Apache Kafka
CouchDB
DevOps
Splunk
Ansible
Azure DevOps
Datadog
Azure
CI/CD
Jenkins
AppDynamics
GitHub
QA
JMeter
Swagger
Postman
Apply
Emulation Engineer 1 day ago
$25k – $59k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Python
C++
SystemC
DevOps
QEMU
Chips/EDA
Synopsys ZeBu
Cadence Palladium
Siemens Veloce
Apply
$26k – $61k per year (Estimated) • Remote/Hybrid • 7+ years exp • Bachelor's Degree • Pune
Python
SQL
Scala
Databases
Snowflake
Apache Kafka
AI/ML
Cursor
Spark
Claude Code
OpenAI Codex
DevOps
GitHub Actions
CI/CD
Jenkins
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon CloudWatch
Amazon Kinesis
Apply
$23k – $69k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Sofia
Python
Java
Java
Spring Boot
Databases
PostgreSQL
Apache Kafka
AI/ML
Copilot
Cursor
DevOps
Rest API
GCP
OpenShift
Azure
CI/CD
AWS
Kubernetes
Management
Agile
Apply
$208k – $386k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • San Francisco
AI/ML
vLLM
Multimodal AI
Function Calling
AI Agents
SGLang
TensorRT-LLM
LLM
KV Cache
Tool Use
Apply
In office • Full-Time • 4+ years exp • Bachelor's Degree • Seoul
Python
JavaScript
TypeScript
SQL
Python
FastAPI
Databases
PostgreSQL
ClickHouse
AI/ML
Multimodal AI
Function Calling
AI Agents
LLM
Tool Use
Frontend
GraphQL
Next.js
React.js
DevOps
gRPC
OpenTelemetry
CI/CD
Kubernetes
Cybersecurity
Auth0
Management
Stripe
Apply
$38k – $100k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seoul
Python
AI/ML
Function Calling
NCCL
InfiniBand
Tool Use
DevOps
Terraform
Ansible
Helm
Cilium
Istio
etcd
GitOps
AWS
Kubernetes
Service Mesh
eBPF
Apply
Account Executive 3 months ago
$130k – $274k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
AI/ML
Multimodal AI
LLM
Hugging Face
Edge AI
DevOps
AWS
Apply
In office • Full-Time • 3+ years exp • Seoul
Python
Go
TypeScript
Python
Asyncio
AI/ML
Multimodal AI
LLM
Hugging Face
DevOps
gRPC
Docker
Kubernetes
Apply
$115k – $180k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • Seattle • Raleigh • New York
Apply
$97k – $124k per year • In office • Full-Time • 1+ year exp • San Francisco
Design
Canva
Management
Slack
Google Workspace
Apply
$200k – $240k per year • Remote/Hybrid • San Francisco
AI/ML
AI Agents
LLM
RAG
Apply
$152k – $301k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Dallas • Austin • San Francisco • Fort Worth • Los Angeles
Marketing
Salesforce
Apply
$285k – $335k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco
DevOps
VMWare
containerd
Kubernetes
KVM
QEMU
Hyper-V
Apply
See all jobs
This is one of many
624,530 more open roles from verified company boards, updated every day.