577,120open jobs
24,628companies
79,541added this week
Browse all
Salary
$145k – $295k per year (Estimated)
Location
Remote/Hybrid (Mountain View, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match

About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and built the environment infrastructure that frontier labs and enterprises use to make their agents reliable.

Bespoke is uniquely positioned to capture a large share of data and RL environment curation.

About the Role

We're looking for an Infrastructure Engineer to own the execution layer beneath our RL environments: the systems that let an agent operate inside a realistic, multi-tool world coherently for hours or days.

This is a hard systems problem disguised as an AI job. As the tasks agents can complete keep lengthening, the environments that train them have to stay coherent across far longer horizons than anything that exists today. That means sandboxing and isolation you can trust, execution that's fast and cheap enough to run at training scale, and the ability to snapshot, restore, inspect, and branch a running environment instead of treating every rollout as one-shot. You'll build the platform that makes all of this possible.

You'll work closely with our research and data teams, and directly with frontier labs and enterprise customers, to turn environment designs into infrastructure that runs reliably in production.

What You'll Do

  • Environment Execution & Sandboxing:

    • Design and own the sandboxing and execution layer that environments run inside. Build systems to snapshot and restore environment state (disk, process, and where relevant memory and accelerator state) so runs can be paused, resumed, inspected, and branched rather than executed once.

    • Develop the machinery to detect failure modes early in a rollout (reward hacks, infra faults, fairness issues) and to revert to a known-good state, patch, and continue.

    • Extend execution to long-horizon and multi-node environments, where an agent operates across many tools and services over hours or days.

  • Performance & Scale

    • Own the performance characteristics of the platform: throughput, latency, and cost-per-rollout at scale.

    • Drive utilization and scheduling so we can run far more environment rollouts per dollar without sacrificing reliability.

    • Profile and remove bottlenecks across the stack, from container startup to environment teardown.

    • Build the observability that lets us understand what's happening inside thousands of concurrent, long-running rollouts.

  • Environment Platform

    • Build and maintain the framework for specifying, packaging, and deploying RL environments which is used by both humans and agents authoring environments internally.

    • Create the tooling that lets researchers and environment authors debug a specific failure across hundreds of long agent traces.

  • Collaboration & Production Excellence

    • Scale prototypes into production systems with reproducible workflows and high engineering standards.

    • Write the documentation and tools that let internal teams and external users build on the platform.

What We're Looking For

  • Systems & Infrastructure

    • Strong track record building production systems or research infrastructure at scale: distributed systems, execution engines, container/sandboxing infrastructure, or similar.

    • Deep comfort with the systems layer: containers and isolation (e.g. namespaces, cgroups, VMs, gVisor/Firecracker-style sandboxing), filesystems, process and state management.

    • Experience making systems fast and cheap - profiling, scheduling, resource utilization, and cost optimization at scale.

    • Proficiency with cloud platforms (GCP, AWS) and distributed computing.

    • Strong engineering fundamentals and a systematic approach to testing, validation, and reliability.

  • Execution & Ownership

    • Comfort operating in ambiguity.

    • Strong Python skills; comfort in a systems language (Rust, Go, or C++) is a plus.

    • Ability to use modern tools such as Claude Code effectively.

  • Collaboration & Communication

    • Excellent communication skills for working with research teams and enterprise customers.

    • Ability to translate between research needs and infrastructure requirements.

    • Comfortable presenting technical work to diverse audiences.

Nice to Have

Experience with RL training or evaluation infrastructure, or the execution layer for agent rollouts.

Experience with checkpoint/snapshot-restore systems, CRIU, or distributed state management.

Background in high-throughput, low-latency execution systems.

Contributions to widely-used infrastructure, datasets, benchmarks, or open-source systems.

Previous experience in a research engineering or infrastructure role at an AI or systems-heavy company.

Logistics

Location: Mountain View, CA

Compensation: Competitive salary and equity

Benefits: Health coverage, and the opportunity to work directly with the world's leading AI research labs

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
577,120 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Mountain View
$16k – $46k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
Terraform
Ansible
GCP
Azure
AWS
Kubernetes
Service Mesh
OpenStack
Web3
Layer 2
Apply
$109k – $223k per year (Estimated) • Remote • 5+ years exp
Python
AI/ML
AI Agents
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
SLI/SLO/SLA
Cybersecurity
FedRAMP
CVSS
EPSS
KEV
BeyondTrust
Apply
$155k – $221k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree
AI/ML
AI Agents
DevOps
GCP
Azure
AWS
Cybersecurity
Zscaler
Zero Trust
Management
Agile
Apply
$95k – $120k per year • In office • Confidential • Full-Time • 2+ years exp • Austin
Python
MATLAB
SAS
Apply
$4k – $14k per year • Equity • In office • 6+ years exp • Bachelor's Degree • Oakland
Python
MATLAB
Design
SolidWorks
Apply
$250k – $300k per year • In office • Full-Time • Mountain View
AI/ML
Reinforcement Learning
AI Agents
Post-training
Tool Use
DevOps
CI/CD
Apply
$16k – $47k per year (Estimated) • In office • Full-Time • 4+ years exp • Bengaluru
Cybersecurity
Okta
Management
Google Workspace
Apply
$123k – $248k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Mountain View
Python
JavaScript
TypeScript
Databases
PostgreSQL
ElasticSearch
AI/ML
AI Agents
LLM
Agentic Workflows
Frontend
Next.js
React.js
DevOps
GCP
AWS
Apply
$23k – $57k per year (Estimated) • In office • Full-Time • 4+ years exp • Bengaluru
Python
JavaScript
TypeScript
Databases
PostgreSQL
ElasticSearch
AI/ML
AI Agents
LLM
Agentic Workflows
Frontend
Next.js
React.js
DevOps
GCP
AWS
Apply
Engagement Manager 4 months ago
$95k – $196k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Mountain View
AI/ML
Reinforcement Learning
Post-training
Tool Use
Apply
In office • Full-Time • 12+ years exp • Associate's Degree • Atlanta • Milwaukee • Dallas • Columbus • Kirkland
Python
C++
AI/ML
Synthetic Data
Physical AI
Robotics
ROS
Gazebo
Isaac Sim
SLAM
Sensor Fusion
Motion Planning
Imitation Learning
Digital Twin
IoT
MQTT
OPC UA
Apply
$94k – $266k per year • Remote • Full-Time • 5+ years exp • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Apply
$97k – $197k per year (Estimated) • In office • Full-Time • 12+ years exp • Associate's Degree • Atlanta • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
Claude
Vertex AI
AI Agents
RAG
OpenAI
Anthropic
Context Engineering
Agentic Workflows
Apply
$91k – $185k per year (Estimated) • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$133k – $338k per year • In office • Full-Time • 8+ years exp • San Francisco • Milwaukee • Dallas • Columbus • Cincinnati
Management
Agile
Apply
See all jobs
This is one of many
577,120 more open roles from verified company boards, updated every day.