368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$90k – $130k per year
Location
In office (Los Angeles)
Seniority
Middle
Employment
Full-Time
Overview
Company
Impact
Profile match
Vast.ai is a cloud computing marketplace company headquartered in San Francisco, California, and founded in 2018. The company operates a platform where owners of idle GPU hardware, from individual hosts to full data centers, rent that capacity to customers who need it for machine learning training and inference. It uses a real time bidding and search model to undercut traditional cloud providers, and is used mainly by AI researchers, startups, and independent developers.

About Us

Vast.ai 's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing - reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on when tickets get complex: diagnosing and resolving issues across the full stack - hardware/BIOS/firmware, networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM).

You'll handle higher-complexity issues, own escalation resolution end-to-end, and contribute to internal documentation and runbooks. The best engineers in this role don't just resolve tickets - they build the tooling and runbooks that eliminate recurring ones. You'll collaborate directly with the engineering team and host support team on systemic issues.

Strong technical depth and support experience are the primary requirements. You should be comfortable working autonomously across Ubuntu environments, diagnosing container and GPU issues, and communicating findings clearly to both technical and non-technical audiences.

Vast.ai users or hosts strongly preferred.

This role is full-time and onsite in our office in Westwood (LA)

Schedule: Sunday - Thursday.

Key Responsibilities

  • Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

  • Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM)

  • Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

  • Investigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks

  • Advise suppliers (hosts) on installation best practices - hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance

  • Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting

  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

  • Collaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issues

  • Assist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloads

  • Provide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation

You Are

  • Fluent in Linux - you navigate systems, read logs, and solve problems from the command line without hesitation

  • Methodical and thorough: you gather data, dig into root causes, and don't settle for surface-level fixes

  • A self-starter who can manage a queue of complex tickets with minimal supervision

  • Adaptable to a defined on-call rotation which may include weekend coverage

  • A clear written communicator: able to explain technical findings and write useful internal documentation

  • Genuinely curious about AI infrastructure, GPU computing, and distributed systems

Must-Haves

  • Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissions

  • Proficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem management

  • Experience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMs

  • Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshooting

  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential)

  • Scripting in Python and Bash for automation and diagnostic tooling

  • Strong English written communication: clear, professional, and technically precise

  • Experience providing technical support in a customer-facing or internal helpdesk context

  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice-to-Haves

  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers

  • Monitoring and observability experience (Prometheus, Grafana)

  • Relevant certifications: RHCSA, CompTIA Linux+, or similar

  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Annual Salary Range

$90,000 - $150,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits

  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Los Angeles
$142k – $213k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Jersey City
Java
Python
SQL
TypeScript
JavaScript
Java
Spring Boot
Python
Asyncio
FastAPI
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Devin
Fine-tuning
Gemini
Google ADK
Hybrid Search
Knowledge Graph
LangChain
LangGraph
RAG
Frontend
Angular
DevOps
CI/CD
Docker
Kubernetes
Rest API
Apply
$94k – $142k per year • Remote/Hybrid • Full-Time • 5+ years exp • Mississauga
JavaScript
Python
SQL
Java
Java
Gradle
Maven
Spring Boot
Databases
Oracle
Frontend
GraphQL
React.js
DevOps
AWS
CI/CD
Docker
Jenkins
OpenShift
Apply
Staff Data Engineer 6 hours ago
$160k – $200k per year • In office • Full-Time • 6+ years exp • Chicago
Python
SQL
Databases
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Arize Phoenix
AutoGen
AWS Bedrock AgentCore
CrewAI
dbt
Fine-tuning
Function Calling
LangChain
LangGraph
LangSmith
LLM
LLM Evaluation
LLM Guardrails
Model Context Protocol
Prefect
Prompt Engineering
RAG
Semantic Kernel
Semantic Search
Semantic Search
Weights & Biases
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
Kubernetes
Vector
Analytics
ETL/ELT
Apply
Cloud Engineer (AWS) 6 hours ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Dalian
Python
DevOps
AWS
CI/CD
CloudFormation
Docker
FinOps
GitHub Actions
Kubernetes
Terraform
GitHub
Apply
$87k – $131k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tysons
Java
Mobile
JUnit
DevOps
AWS
Azure
Azure DevOps
CI/CD
Datadog
Docker
Dynatrace
GCP
Git
Jenkins
Kubernetes
New Relic
Shift-Left
Splunk
Cybersecurity
Shift-Left Security
SonarQube
Management
Jira
QA
Cucumber
JMeter
Postman
Rest-Assured
Selenium
TestNG
Apply
$90k – $150k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
LLM
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
HPC
Apply
$90k – $150k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
LLM
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
Apply
$170k – $240k per year • In office • Full-Time • 3+ years exp • Los Angeles
C++
Python
SQL
Databases
PostgreSQL
Redis
AI/ML
AI Agents
DevOps
AWS
Docker
Rest API
Terraform
Apply
$200k – $320k per year • In office • Full-Time • 10+ years exp • San Francisco • Los Angeles
AI/ML
CUDA Toolkit
Apply
$180k – $300k per year • In office • Full-Time • Los Angeles • San Francisco
C++
Python
Databases
PostgreSQL
Redis
AI/ML
LLM
DevOps
AWS
Docker
gRPC
KVM
Terraform
Cybersecurity
Zero Trust
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$94k – $294k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Portland • Milwaukee • Dallas • Columbus
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$150k – $185k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Los Angeles
AI/ML
Human-in-the-Loop
Apply
$228k – $363k per year • Equity • In office • Full-Time • 5+ years exp • Boston • New York • Los Angeles
Ruby
SQL
JavaScript
Ruby
Ruby on Rails
AI/ML
AI Agents
Frontend
React.js
Mobile
React Native
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.