1,180,240open jobs
66,362companies
211,333added this week
Browse all
Salary
$175k – $250k per year
Location
Remote (United States)
Seniority
Senior
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 3, 2026. First seen by Alion on Aug 3, 2026.

Overview
Company
Impact
Profile match
Boundless is the inference partner that helps AI-native companies scale their AI usage, lower costs, and keep quality high.

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads - a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.

Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.

Bare-Metal & GPU Optimization: Go deep on GPU performance - PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology - to push throughput per node.

Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.

Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling - without sacrificing reliability.

Requirements

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action

Nice to Have

  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations

Additional Requirements

  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,180,240 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
≈ $135k – $240k per year (Estimated) • Remote (United States) • 7+ years exp • Mountain View • San Francisco • Seattle
Python
Go
JavaScript
TypeScript
AI/ML
Claude
ChatGPT
Model Context Protocol
Embeddings
AI Agents
Hallucination
Reranking
Semantic Search
Hybrid Search
A2A
Semantic Search
Edge AI
Frontend
GraphQL
React.js
DevOps
Rest API
CI/CD
Apply
≈ $116k – $255k per year (Estimated) • Remote (AMER) • Full-Time • 5+ years exp
Rust
C++
AI/ML
vLLM
Quantization
Knowledge Distillation
SGLang
TensorRT
TensorRT-LLM
Triton
Speculative Decoding
KV Cache
Model Distillation
Machine Learning
DevOps
CI/CD
Apply
≈ $140k – $250k per year (Estimated) • Remote (United States) • Seattle
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
NLP
LLM
RAG
Hallucination
Reranking
Context Engineering
Tool Use
Machine Learning
Apply
≈ $139k – $248k per year (Estimated) • Remote (United States) • Seattle • San Francisco • Austin • New York
AI/ML
AI Agents
PyTorch
RAG
Reranking
Semantic Search
Triton
Semantic Search
LLM Evaluation
Machine Learning
DevOps
Platform Engineering
FinOps
SLI/SLO/SLA
Management
Confluence
Jira
Apply
≈ $137k – $245k per year (Estimated) • Remote (United States) • 5+ years exp • Bachelor's Degree • Seattle • San Francisco • Mountain View
Python
Java
Kotlin
TypeScript
SQL
Databases
Databricks
AI/ML
Spark
Fine-tuning
AI Agents
Feature Store
Machine Learning
DevOps
AWS
Management
Agile
Apply
$31k – $47k per year (gross) • In office • Full-Time • 4+ years exp • Hyderabad
Python
JavaScript
PHP
TypeScript
SQL
Node JS
Databases
MySQL
Redis
AI/ML
Copilot
Cursor
Claude
ChatGPT
Frontend
Angular
DevOps
Rest API
Azure
CI/CD
Git
AWS
Linux
QA
Playwright
Postman
Apply
≈ $58k – $114k per year (Estimated) • Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • New York
Python
SQL
Analytics
Power BI
Management
ServiceNow
Apply
$94k – $179k per year • Remote (United States) • Full-Time • 6+ years exp • Bachelor's Degree
Python
JavaScript
Java
C++
Java
Apache Camel
Databases
Apache Kafka
DevOps
Rest API
gRPC
GCP
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Amazon S3
Analytics
Apache NiFi
Management
Agile
Apply
$84k – $98k per year • Hybrid • Contractor • Toronto
Java
TypeScript
Java
Spring Boot
DevOps
GitHub Actions
CI/CD
Jenkins
Platform Engineering
QA
WebDriverIO
Apply
≈ $104k – $187k per year (Estimated) • Hybrid • Full-Time • Morocco
Python
Java
Kotlin
SQL
C#
DevOps
Rest API
CI/CD
AWS
Docker
Kubernetes
GitLab
Cybersecurity
Keycloak
Apply
Applied AI/ML Engineer 2 months ago
$175k – $250k per year • Remote (United States) • Full-Time
Python
AI/ML
vLLM
CUDA Toolkit
Reinforcement Learning
Quantization
SGLang
TensorRT
SkyPilot
TensorRT-LLM
PyTorch
LLM
Ray
CUDA
DPO
SFT
PPO
GRPO
Post-training
Megatron-LM
FSDP
LLM Evaluation
Speculative Decoding
KV Cache
DevOps
SLURM
Kubernetes
GitHub
Apply
$100k – $150k per year • Remote (United States) • Full-Time • San Francisco
AI/ML
vLLM
SGLang
Post-training
Apply
$100k – $150k per year • Remote (United States) • Full-Time
AI/ML
Post-training
Apply
See all jobs
This is one of many
1,180,240 more open roles from verified company boards, updated every day.