1,175,636open jobs
66,247companies
209,092added this week
Browse all
Salary
$175k – $250k per year
Location
Remote (United States)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 3, 2026. First seen by Alion on Aug 3, 2026.

Overview
Company
Impact
Profile match
Boundless is the inference partner that helps AI-native companies scale their AI usage, lower costs, and keep quality high.

Boundless is coordinating GPU compute at scale and building toward becoming a leader in AI. As anApplied AI/ML Engineer, you'll ship AI-powered products end-to-end on top of our growing GPU inference fleet - owning everything from serving low-latency inference to standing up reinforcement-learning post-training pipelines. This is a builder's role: you take an idea from prototype to production, tune it for throughput and cost on real GPUs, and iterate fast on customer and internal feedback.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

End-to-End AI Product Delivery: Own AI features and products from prototype through production - model selection, serving, evaluation, and iteration - shipping working software rather than research artifacts.

Inference Serving: Deploy and optimize LLM inference across the fleet using vLLM and SGLang. Tune continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput and minimize latency and cost per token.

RL & Post-Training Harnesses: Build and operate reinforcement-learning and post-training pipelines using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers). This includes reward and verifier design, rollout orchestration, weight synchronization, and keeping long-running training stable.

Evaluation & Iteration: Build eval harnesses and benchmarks that measure quality, throughput, and cost together, and use them to drive fast, data-informed iteration.

Work Across the Stack: Partner with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next and why.

Requirements

  • 3+ years shipping ML/AI systems to production
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
  • Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
  • Strong Python and PyTorch
  • Working understanding of GPU execution: batching, memory, and basic CUDA concepts
  • Comfort operating in ambiguity with a strong bias for action

Nice to Have

  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
  • Distributed training experience (FSDP, TP/PP/DP parallelism)
  • Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
  • Experience with verifiable inference or large-scale distributed systems
  • Kubernetes and container-based deployment
  • Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)

Additional Requirements

  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,175,636 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
$21k – $31k per year (gross) • Remote (India) • 3+ years exp • Gurgaon
Python
Python
FastAPI
AI/ML
OpenCV
Triton Inference Server
Fine-tuning
Quantization
Multimodal AI
Computer Vision
TensorRT
TensorFlow
PaddlePaddle
PyTorch
PaddleOCR
Synthetic Data
TorchServe
OCR
ONNX Runtime
Machine Learning
DevOps
Docker
Apply
$76k – $126k per year • Hybrid • Secret • 1+ year exp • Bachelor's Degree • McLean
Python
JavaScript
TypeScript
Node JS
Python
FastAPI
AI/ML
Hadoop
Fine-tuning
RAG
OpenAI
Anthropic
Frontend
Vue.js
Svelte
Next.js
React.js
Remix
React Router
DevOps
Vercel
Azure
AWS
Analytics
ETL/ELT
Apply
$66k – $108k per year • Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Indianapolis
AI/ML
AI Agents
Machine Learning
Apply
Remote (Nigeria) • Full-Time • Nigeria
Python
SQL
Python
Flask
FastAPI
Databases
Weaviate
Pinecone
FAISS
AI/ML
LangChain
Scikit-learn
TensorFlow
Pandas
NumPy
PyTorch
Hugging Face
Machine Learning
DevOps
GCP
Azure
Git
AWS
Docker
GitHub
Apply
Lead AI Engineer 5 hours ago
$93k – $170k per year • In office • 6+ years exp • Bachelor's Degree • Dallas
AI/ML
Copilot
LangGraph
LangChain
Claude Code
LlamaIndex
Prompt Engineering
AI Agents
Gemini
LLM
RAG
OpenAI Codex
AWS Strands Agents
Agentic Workflows
Multi-Agent Systems
DevOps
Azure
AWS
Apply
Storage AI Specialist 5 hours ago
≈ $30k – $62k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Bengaluru
Python
Java
SQL
AI/ML
LangChain
Claude
MLFlow
Fine-tuning
Scikit-learn
Prompt Engineering
AI Agents
Llama
PyTorch
RAG
LLMOps
Machine Learning
DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
≈ $132k – $261k per year (Estimated) • In office • Dallas
Python
TypeScript
SQL
AI/ML
Embeddings
AI Agents
DevOps
Azure
AWS
AWS Lambda
AWS Step Functions
API Gateway
Analytics
ETL/ELT
Apply
≈ $112k – $220k per year (Estimated) • In office • 15+ years exp • Dallas
Python
SQL
Databases
Microsoft Fabric
DevOps
Rest API
Azure
Analytics
Power BI
Apply
$93k – $170k per year • In office • 8+ years exp • Bachelor's Degree • Dallas
Python
Java
TypeScript
SQL
C#
AI/ML
Fine-tuning
AI Agents
Machine Learning
Cybersecurity
OWASP
Management
Kanban
Apply
$80k – $137k per year • In office • 5+ years exp • Bachelor's Degree • Dallas
Python
Java
TypeScript
SQL
C#
AI/ML
AI Agents
Machine Learning
DevOps
CI/CD
Cybersecurity
OWASP
Management
Agile
Kanban
Apply
$175k – $250k per year • Remote (United States) • Full-Time
Python
Rust
TypeScript
Bash
AI/ML
CUDA Toolkit
SkyPilot
Ray
CUDA
DevOps
Terraform
Ansible
K3s
Pulumi
SLURM
Docker
Kubernetes
GitHub
Linux
Cybersecurity
Teleport
Apply
$100k – $150k per year • Remote (United States) • Full-Time • San Francisco
AI/ML
vLLM
SGLang
Post-training
Apply
$100k – $150k per year • Remote (United States) • Full-Time
AI/ML
Post-training
Apply
See all jobs
This is one of many
1,175,636 more open roles from verified company boards, updated every day.