368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$90k – $150k per year
Location
In office (Los Angeles)
Employment
Full-Time
Overview
Company
Impact
Profile match
Vast.ai is a cloud computing marketplace company headquartered in San Francisco, California, and founded in 2018. The company operates a platform where owners of idle GPU hardware, from individual hosts to full data centers, rent that capacity to customers who need it for machine learning training and inference. It uses a real time bidding and search model to undercut traditional cloud providers, and is used mainly by AI researchers, startups, and independent developers.

About Us

Vast.ai 's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing - reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This role focuses on troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, GPU workloads, Ubuntu, Docker, KVM based virtual machines, networking, hardware, BIOS, and firmware. You’ll investigate failures, reproduce issues, identify root causes, and propose practical solutions across the full infrastructure stack.

You’ll also serve as the engineering resource our L1 support team relies on when tickets go beyond frontline triage. You’ll own complex escalations end-to-end, gather technical evidence, coordinate with the appropriate teams, and communicate findings clearly to clients, infrastructure suppliers, and internal teams.

The best engineers in this role don’t just resolve individual issues-they recognize recurring patterns, improve diagnostic tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic Linux, GPU, and infrastructure problems.

Strong GPU troubleshooting experience, Linux systems knowledge, and technical support skills are the primary requirements. You should be comfortable working autonomously in Ubuntu environments and troubleshooting NVIDIA drivers, CUDA, containers, virtual machines, networking, hardware, and GPU workloads.

Vast.ai users or hosts strongly preferred.

Location and Schedule

This is a full-time position based in our Westwood, Los Angeles office.

Available schedules:

  • Monday-Friday: Fully on-site

  • Sunday-Thursday: Four days on-site and one day working from home

Key Responsibilities

  • Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments

  • Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks

  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads

  • Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

  • Provide managed support for supplier onboarding and ongoing machine management, including installation, configuration, and post setup troubleshooting

  • Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance

  • Provide coverage for L1 support overflow during peak periods or incidents

  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

  • Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues

You Are

  • Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line

  • Someone who enjoys debugging difficult problems and fixing broken systems

  • Methodical and focused on finding root causes, not just temporary fixes

  • Able to manage complex tickets independently

  • A clear written communicator with an interest in AI infrastructure and GPU computing

Must-Haves

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions

  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting

  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting

  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting

  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting

  • Python and Bash scripting skills for automation and diagnostic tooling

  • Strong written English communication that is clear, professional, and technically precise

  • Experience providing technical support in a customer-facing or internal help desk environment

  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice-to-Haves

  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers

  • Monitoring and observability experience (Prometheus, Grafana)

  • Relevant certifications: RHCSA, CompTIA Linux+, or similar

  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Interview Process (~1 week)

After you submit your application, our technical team will review your experience and qualifications. Selected candidates will proceed through the following stages:

  • 15 minutes - Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role

  • 45 minutes - Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience

  • 2 hours - Meet and Greet and Technical Assessment (On-site): Meet the team and complete an LLM-assisted Linux systems operations assessment

Annual Salary Range

$90,000 - $160,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits

  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Los Angeles
Remote • Full-Time • 3+ years exp • Prague
Bash
PowerShell
Python
DevOps
Splunk
Cybersecurity
Crowdstrike
IBM QRadar
Apply
Architect - Java 1 day ago
$45k – $98k per year (Estimated) • In office • Full-Time • 8+ years exp • Kochi
Java
SQL
Java
Hazelcast
Hibernate
Spring Boot
Spring Cloud
Databases
Apache Kafka
Cassandra
DynamoDB
MySQL
PostgreSQL
RabbitMQ
Redis
AI/ML
Copilot
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
Git
GitHub Actions
Grafana
Istio
Jenkins
Kong
Kubernetes
Linkerd
OpenTelemetry
Platform Engineering
Prometheus
Rest API
Service Mesh
API Gateway
GitHub
Marketing
Salesforce
Apply
$35k – $82k per year (Estimated) • In office • Full-Time • 13+ years exp • Kochi
Java
SQL
Java
Hazelcast
Hibernate
Spring Boot
Spring Cloud
Databases
Apache Kafka
Cassandra
DynamoDB
MySQL
PostgreSQL
RabbitMQ
Redis
AI/ML
Copilot
DevOps
AWS
Azure
Azure DevOps
CI/CD
CloudFormation
Docker
GCP
Git
GitHub Actions
Grafana
Istio
Jenkins
Kubernetes
Linkerd
OpenTelemetry
Platform Engineering
Prometheus
Service Mesh
Terraform
GitHub
Marketing
Salesforce
Apply
Cloud Engineer (AWS) 7 hours ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Dalian
Python
DevOps
AWS
CI/CD
CloudFormation
Docker
FinOps
GitHub Actions
Kubernetes
Terraform
GitHub
Apply
$28k – $65k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Bash
PowerShell
Python
Node JS
JavaScript
Node JS
Commander.js
AI/ML
AI Agents
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
Kubernetes
Amazon ECS
IAM
Cybersecurity
Crowdstrike
Zero Trust
Apply
$90k – $150k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
LLM
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
Apply
$170k – $240k per year • In office • Full-Time • 3+ years exp • Los Angeles
C++
Python
SQL
Databases
PostgreSQL
Redis
AI/ML
AI Agents
DevOps
AWS
Docker
Rest API
Terraform
Apply
$90k – $130k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
Apply
$200k – $320k per year • In office • Full-Time • 10+ years exp • San Francisco • Los Angeles
AI/ML
CUDA Toolkit
Apply
$180k – $300k per year • In office • Full-Time • Los Angeles • San Francisco
C++
Python
Databases
PostgreSQL
Redis
AI/ML
LLM
DevOps
AWS
Docker
gRPC
KVM
Terraform
Cybersecurity
Zero Trust
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
Network Engineer 3 hours ago
$189k – $227k per year • In office • Full-Time • 8+ years exp • Los Angeles
DevOps
AIOps
Ansible
AWS
Azure
CI/CD
Git
Kubernetes
Pulumi
Service Mesh
Terraform
Cybersecurity
Calico
Apply
Network Administrator 3 hours ago
$99k – $234k per year (Estimated) • In office • Full-Time • 5+ years exp • Los Angeles
DevOps
Ansible
AWS
Azure
CI/CD
Git
Kubernetes
Pulumi
Terraform
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.