1,443,753open jobs
85,355companies
221,329added this week
Browse all
Salary
$121k – $137k per year
Location
Remote (United States, Singapore)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 6, 2026. vCluster scores A on the Alion truth index.

Overview
Company
Impact
Profile match
vCluster, formerly Loft Labs, builds virtual Kubernetes clusters that give each team or tenant an isolated control plane running inside a shared physical cluster. The approach cuts the cost and operational overhead of running many real clusters while keeping namespaces, custom resources and even Kubernetes versions properly separated between workloads that would otherwise collide. Platform teams use it for development environments, multi-tenant products and AI infrastructure where isolation would otherwise mean duplicating everything.
Backed by Fusion Fund

As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer.

vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.

As an AI Infrastructure Engineer, your role will include:

  • Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.

  • Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.

  • Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.

  • Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.

  • Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.

  • Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.

  • Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

This role could be a fit for you if you bring:

  • Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.

  • GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.

  • Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.

  • Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.

  • Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.

  • Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.

Bonus points for:

  • Automation Skills: Experience writing automation scripts with Bash, Python, or Go.

  • Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.

  • AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.

  • Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.

About vCluster Labs

We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.

We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.

We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Benefits

We offer the following benefits:

  • Competitive Salary: We offer a competitive compensation package, including equity.

  • Premium Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

  • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At vCluster Labs, we value and stand for:

  • Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

  • Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

  • Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

  • Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy - the strongest ideas win, no matter who or where they come from.

  • Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,443,753 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Singapore
≈ $41k – $105k per year (Estimated) • Remote (Greece) • Full-Time
Python
PowerShell
Databases
Databricks
DevOps
Terraform
Azure DevOps
GitHub Actions
Prometheus
Azure
CI/CD
Docker
Kubernetes
Grafana
Bicep
Azure AKS
Cybersecurity
GDPR
Apply
≈ $40k – $101k per year (Estimated) • Remote (Serbia) • Serbia
DevOps
Terraform
GitHub Actions
Azure
CI/CD
AWS
Bicep
Cybersecurity
Least Privilege
Threat Modeling
Apply
$73k – $90k per year • Remote (France) • Full-Time • France
Python
Bash
Databases
PostgreSQL
ElasticSearch
Apache Kafka
DevOps
Terraform
Ansible
GCP
Prometheus
Grafana
Configuration Management
Linux
Apply
$88k – $123k per year • In office • 3+ years exp • Bachelor's Degree • Vista
SQL
Apply
$112k – $154k per year • In office • 4+ years exp • Bachelor's Degree • Vista
Python
PowerShell
DevOps
Windows Server
Cybersecurity
Active Directory
IoT
MQTT
Apply
≈ $67k – $152k per year (Estimated) • In office • Full-Time • Kiel
Python
Go
JavaScript
SQL
Ruby
C#
C++
Databases
Databricks
AI/ML
Prompt Engineering
DevOps
Docker Compose
CI/CD
Git
Docker
Kubernetes
Analytics
ETL/ELT
DataStage
Fivetran
Management
n8n
Zapier
Power Automate
Power Apps
Apply
≈ $137k – $316k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Tel Aviv
AI/ML
InfiniBand
DevOps
HPC
Linux
BGP
OSPF
Apply
In office • Full-Time • Bangkok
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
Databases
MySQL
PostgreSQL
Redis
Db2
Oracle
MS SQL
MariaDB
Apache Kafka
Frontend
Angular
JQuery
DevOps
GCP
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Apply
≈ $64k – $129k per year (Estimated) • Hybrid • Belfast
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Vertex AI
Gemini
DevOps
Terraform
GCP
Kubernetes
Google GKE
Google Cloud Run
IAM
Apply
≈ $93k – $182k per year (Estimated) • Hybrid • Manchester
Python
DevOps
Terraform
Puppet
GCP
Istio
Chef
Linkerd
Prometheus
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
SRE
Service Mesh
Configuration Management
Google GKE
IAM
Management
Outlook
Apply
$101k – $118k per year • Remote (United States) • Full-Time • 5+ years exp
Python
AI/ML
CUDA Toolkit
LLM
Ray
CUDA
OpenAI
InfiniBand
DevOps
SLURM
Kubernetes
GitHub
GitLab
Management
Stripe
Apply
≈ $127k – $235k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp
Python
AI/ML
CUDA Toolkit
LLM
Ray
CUDA
OpenAI
InfiniBand
DevOps
SLURM
Kubernetes
GitHub
GitLab
Management
Stripe
Apply
$125k – $145k per year • Remote (United States) • Full-Time • PhD
AI/ML
Ray
OpenAI
DevOps
SLURM
Kubernetes
Platform Engineering
GitHub
GitLab
Management
Stripe
Marketing
Salesforce
LinkedIn Ads
Google Ads
HubSpot
Marketo
LinkedIn
Apply
$112k – $129k per year • Remote (Germany) • Full-Time
AI/ML
Ray
OpenAI
DevOps
SLURM
Kubernetes
Platform Engineering
GitHub
GitLab
Management
Stripe
Marketing
Salesforce
Apply
$280k – $330k per year • Remote (United States) • Full-Time • 7+ years exp • United States
AI/ML
Ray
OpenAI
DevOps
SLURM
Kubernetes
GitHub
GitLab
Management
Stripe
Marketing
Salesforce
Apply
≈ $106k – $261k per year (Estimated) • Hybrid • 4+ years exp • Bachelor's Degree • Singapore
AI/ML
NCCL
InfiniBand
DevOps
Terraform
Helm
SLURM
Docker
Kubernetes
Platform Engineering
HPC
Apply
≈ $70k – $131k per year (Estimated) • In office • Contractor • 6+ years exp • Singapore
Apply
≈ $23k – $41k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Singapore
Analytics
Microsoft Excel
Apply
≈ $36k – $66k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Singapore
Python
Go
JavaScript
SQL
Cybersecurity
Threat Modeling
Apply
≈ $35k – $70k per year (Estimated) • In office • 1+ year exp • Bachelor's Degree • Singapore
Python
JavaScript
Rust
C++
Dart
AI/ML
LLM
Machine Learning
Mobile
Flutter
Apply
See all jobs
This is one of many
1,443,753 more open roles from verified company boards, updated every day.