545,534open jobs
20,145companies
74,873added this week
Browse all
Salary
$143k – $251k per year (Estimated)
Location
Remote (United States, Canada)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
vCluster, formerly Loft Labs, builds virtual Kubernetes clusters that give each team or tenant an isolated control plane running inside a shared physical cluster. The approach cuts the cost and operational overhead of running many real clusters while keeping namespaces, custom resources and even Kubernetes versions properly separated between workloads that would otherwise collide. Platform teams use it for development environments, multi-tenant products and AI infrastructure where isolation would otherwise mean duplicating everything.

As a Senior Inference Engineer at vCluster Labs, you are the first engineer we're hiring to own inference. You'll partner directly with our CTO to build the platform's inference layer from the ground up, taking models and turning them into a production-grade, query-to-response pipeline running at scale. From there, you will help lead the engineering direction of inference at vCluster, partnering with Product to shape what we build next as the space evolves.

As a Senior Inference Engineer, your role will include:

  • Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure, owning the full pipeline from a customer's query to the served response.

  • Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.

  • Optimizing for scale: Apply quantization, batching, caching, and routing to keep latency and cost in check as traffic grows.

  • Programming, not just configuring: Build real infrastructure in Python or Golang - this is an engineering role, not a research or data-science one.

  • Owning the roadmap: Build the first iteration alongside our CTO, then take the lead on the inference platform and partner with Product to decide what we build next.

This role could be a fit for you if you bring:

  • Production LLM serving experience: You've deployed and served LLMs using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale.

  • Inference optimization know-how: Hands-on experience with quantization, batching, caching, and routing, not just familiarity with the terms.

  • Hands-on programming experience: Strong engineering skills in Python or Golang, with real production code experience.

  • Communication: Strong communication skills, explaining technical concepts clearly to both engineers and non-technical stakeholders.

Bonus points for:

  • Familiarity with containerized environments (Docker, Kubernetes)

  • Hands-on generative AI experience with common ML frameworks (PyTorch, Transformers)

  • Good understanding of the GPU stack: CUDA, NCCL, drivers, and related libraries

  • Knowledge of model architectures and fine-tuning approaches

  • Experience with NVIDIA Dynamo

About vCluster Labs

We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.

We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.

We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Benefits

We offer the following benefits:

  • Competitive Salary: We offer a competitive compensation package, including equity.

  • Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

  • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At vCluster Labs, we value and stand for:

  • Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

  • Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

  • Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

  • Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy - the strongest ideas win, no matter who or where they come from.

  • Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
545,534 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$28k – $80k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Ho Chi Minh City
Python
Go
JavaScript
Java
PowerShell
C#
Node JS
Go
Chi
C#
.NET
Databases
MS SQL
Frontend
React.js
Sass
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Management
SharePoint
Agile
Apply
$41k – $88k per year (Estimated) • Remote/Hybrid • 10+ years exp • Bachelor's Degree • Bengaluru
Python
Go
DevOps
Terraform
Ansible
GCP
CircleCI
CloudFormation
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
GitHub
Apply
Remote/Hybrid • 8+ years exp
Python
Java
Databases
PostgreSQL
Redis
Snowflake
Apache Kafka
AI/ML
AI Agents
LLM
DevOps
AWS
Kubernetes
Apply
$52k – $87k per year • In office • Full-Time • Berlin
Python
Java
TypeScript
AI/ML
Model Context Protocol
Apply
Solutions Engineer 19 min ago
$160k – $220k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
JavaScript
PHP
Ruby
C#
Node JS
AI/ML
LangChain
LlamaIndex
LLM
Genkit
DevOps
Rest API
GCP
Vercel
Azure
AWS
IAM
Cybersecurity
Auth0
Apply
$225k – $280k per year • Remote • Full-Time
AI/ML
Ray
OpenAI
DevOps
GCP
SLURM
Azure
AWS
Kubernetes
GitHub
GitLab
Management
Stripe
Marketing
Salesforce
Apply
$225k – $280k per year • Remote • Full-Time • 8+ years exp
AI/ML
Ray
OpenAI
DevOps
SLURM
Kubernetes
Platform Engineering
GitHub
GitLab
Management
Stripe
Apply
$105k – $125k per year • Remote • Full-Time
AI/ML
Ray
OpenAI
DevOps
SLURM
Kubernetes
Platform Engineering
GitHub
GitLab
Management
Stripe
Marketing
Salesforce
LinkedIn
Apply
$135k – $231k per year (Estimated) • Remote • Full-Time • 5+ years exp
AI/ML
LLM
Ray
OpenAI
DevOps
SLURM
Kubernetes
GitHub
GitLab
Management
Stripe
Apply
$146k – $274k per year (Estimated) • Remote • Full-Time • 5+ years exp
AI/ML
Ray
OpenAI
DevOps
SLURM
Kubernetes
GitHub
GitLab
HPC
Management
Stripe
Apply
See all jobs
This is one of many
545,534 more open roles from verified company boards, updated every day.