368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$180k – $300k per year
Location
In office (Los Angeles, San Francisco)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Vast.ai is a cloud computing marketplace company headquartered in San Francisco, California, and founded in 2018. The company operates a platform where owners of idle GPU hardware, from individual hosts to full data centers, rent that capacity to customers who need it for machine learning training and inference. It uses a real time bidding and search model to undercut traditional cloud providers, and is used mainly by AI researchers, startups, and independent developers.

About Us

Vast.ai ’s cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing-reshaping our future for the benefit of humanity.

We are a growing and highly motivated team dedicated to an ambitious technical plan. Our structure is flat, our ambitions are out-sized, and leadership is earned by shipping excellence.

We seek engineers with strong intrinsic drive, a true passion for advancing the state of the art, and a mix of architecture, coding, and communication skills.

LOCATION: On-site at our office in San Francisco or Westwood, Los Angeles.

About the Role

As a Senior Infrastructure Engineer, you will help design and scale the core systems that power Vast.ai’s global GPU marketplace.

You’ll work closely with our founders and core engineering team to extend the underlying compute infrastructure - from GPU provisioning and scheduling to billing, orchestration, and marketplace dynamics.

We’re looking for someone who has previously built large-scale infrastructure platforms - systems with similarities to Vast.ai, or distributed compute orchestration frameworks.

Full-time · On-site at either our SF or LA offices

Tech Stack

Python, C++, PostgreSQL, Linux, Docker, KVM, Redis, Terraform, AWS, REST/gRPC APIs

Ideal Experience

  • Distributed Systems: Experience building high-throughput backend systems or compute clouds

  • Compute Orchestration: Familiarity with Docker, or custom scheduling frameworks

  • GPU Infrastructure: Understanding of GPU provisioning, driver management, and workload scheduling

  • Billing & Metering: Implemented or integrated usage-based billing and account credit systems

  • Marketplace Dynamics: Knowledge of dynamic pricing, spot instances, or supply-demand balancing mechanisms

  • Security & Multi-Tenancy: Experience designing secure, multi-tenant systems in cloud environments

  • Programming: Strong programming skills in Python and C++; ability to write performant, maintainable, well-architected code

  • Database Expertise: Comfortable designing schemas and queries for large-scale data systems (PostgreSQL preferred)

Bonus points for:

  • Experience with GPU security, virtualization, or zero-trust compute isolation

  • Prior startup experience or end-to-end product ownership

Key Responsibilities

  • Improve the backend systems that power Vast.ai’s compute marketplace

  • Integrate GPU provider onboarding, usage tracking, billing, and orchestration APIs

  • Develop scalable infrastructure for workload scheduling and resource management

  • Optimize pricing and marketplace logic for efficiency and transparency

  • Benchmark, profile, and harden systems for performance, reliability, and fault tolerance

  • Collaborate with product and infrastructure teams to shape the future of decentralized compute

Interview Process (≈ 1 week)

After submitting your application, our technical team reviews your credentials. If selected, you’ll proceed through the following stages:

  • 15 min - Initial screening with member of your future team (virtual)

  • 40 min - Systems and architectures (virtual)

  • 1 hour - LLM-assisted coding assessment (virtual)

  • 2 hours - Meet and greet with coding assessment (on-site)

Benefits

  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Los Angeles
$17k – $43k per year (Estimated) • Remote • Moscow
C++
Java
Python
Java
Maven
DevOps
Ansible
CI/CD
Docker
Git
Graylog
HAProxy
Jenkins
Nginx
Prometheus
Zabbix
Apply
$143k – $173k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Wiesbaden
Python
DevOps
Ansible
Terraform
Cybersecurity
Defense in Depth
Wireshark
Zero Trust
Apply
Network Engineer 5 hours ago
$83k – $113k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Raleigh
JavaScript
Perl
Python
DevOps
Splunk
Apply
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
$128k – $173k per year • In office • Full-Time • 7+ years exp • United States
Python
SQL
Databases
Databricks
Snowflake
DevOps
AWS
SLI/SLO/SLA
Analytics
Power BI
Tableau
Management
Confluence
Jira
Apply
$90k – $150k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
LLM
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
HPC
Apply
$90k – $150k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
LLM
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
Apply
$170k – $240k per year • In office • Full-Time • 3+ years exp • Los Angeles
C++
Python
SQL
Databases
PostgreSQL
Redis
AI/ML
AI Agents
DevOps
AWS
Docker
Rest API
Terraform
Apply
$90k – $130k per year • In office • Full-Time • Los Angeles
Bash
Python
AI/ML
CUDA Toolkit
PyTorch
TensorFlow
DevOps
CentOS Stream
Debian
Docker
Docker Compose
Grafana
KVM
Prometheus
Proxmox VE
Ubuntu
VMWare
Apply
$200k – $320k per year • In office • Full-Time • 10+ years exp • San Francisco • Los Angeles
AI/ML
CUDA Toolkit
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$94k – $294k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Portland • Milwaukee • Dallas • Columbus
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$150k – $185k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Los Angeles
AI/ML
Human-in-the-Loop
Apply
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.