368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$61k – $133k per year (Estimated)
Location
Remote/Hybrid (Yokohama, Japan)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview

The Platform team turns raw heterogeneous compute into a serving platform. ai& owns its data centers and runs AMD, NVIDIA, and Tenstorrent silicon side by side. Your job is everything between the bare metal and the inference engines: cluster orchestration, node lifecycle, scaling, networking, observability, and reliability.

This is not cloud consumption. When capacity is short, you add nodes we own. When a fabric misbehaves, you debug it down to the switch. The platform must let a small team operate hundreds of nodes across multiple sites without heroics, and it must scale by an order of magnitude over the next two years as new sites come online.

You will work directly with the inference team, which owns the engines and serving gateway, and the data center team, which owns power, cooling, and physical deployment. You own the layer that makes their work composable.

Responsibilities

  • Compute orchestration Run Kubernetes across GPU clusters in ai&-owned data centers. Own node lifecycle from bring-up and burn-in through drain and repair, across multiple accelerator vendors.

  • Scaling Build the capacity and scheduling machinery that places inference workloads across heterogeneous silicon and multiple sites, and that lets us bring a new site from empty racks to serving traffic on a predictable timeline.

  • Reliability Define and hold SLOs for the platform. Build the observability stack (metrics, logs, tracing, alerting) and the failure isolation that keeps one bad node or one bad rollout from becoming an incident.

  • Networking and data Operate high-bandwidth fabrics for multi-node inference. Solve model weight distribution: getting hundreds of gigabytes onto the right nodes fast, every time a model ships.

  • Deployment machinery Own CI/CD and GitOps for the fleet. Infrastructure as code, reproducible node images, safe rollouts.

You may be a fit if you have the following skills

  • Production Kubernetes at scale You have operated large multi-cluster Kubernetes environments, ideally with GPU scheduling, device plugins, and topology-aware placement.

  • Systems depth Strong Linux fundamentals. You can reason about NUMA, PCIe, NICs, and storage, and you debug from symptoms to root cause without guessing.

  • Networking fundamentals You understand L2/L3, and ideally RDMA fabrics (InfiniBand or RoCE) in production.

  • Infrastructure as code Terraform or equivalent, GitOps workflows, and the discipline to keep the fleet reproducible.

  • Ownership under load You have carried a pager for systems that matter and you build so the pager stays quiet.

  • Relevant tooling Go or Python, Prometheus-family observability, and comfort automating anything you do twice.

  • Great team spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokohama
$58k – $136k per year (Estimated) • Remote/Hybrid • Full-Time • Guadalajara
Bash
Node JS
Python
JavaScript
Python
FastAPI
Databases
RabbitMQ
Frontend
Next.js
React.js
DevOps
ArgoCD
AWS
CI/CD
Datadog
Docker
Git
GitHub
GitHub Actions
GitOps
Grafana
Helm
Jenkins
Karpenter
KEDA
Kubernetes
OpenTelemetry
Prometheus
Terraform
Apply
AI Engineer 2 hours ago
$120k – $130k per year • In office • Full-Time • Texas
Python
SQL
Databases
Apache Kafka
Snowflake
AI/ML
Amazon SageMaker
Hadoop
DevOps
CI/CD
Git
GitLab
Analytics
ETL/ELT
Apply
ARG SSR Data Engineer 2 hours ago
$38k – $95k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Argentina
Python
SQL
Databases
Trino
AI/ML
Airflow
DevOps
Amazon S3
AWS
Azure
CI/CD
GCP
Git
Analytics
ETL/ELT
Apply
$66k – $123k per year • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Corvallis
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
dbt
DevOps
AWS
GitHub
Analytics
Power BI
Tableau
Apply
$169k – $253k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Boulder
JavaScript
TypeScript
AI/ML
AI Agents
Frontend
GraphQL
React.js
DevOps
AWS
Azure
GCP
Google GKE
Helm
Kubernetes
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
CUDA
CUDA Toolkit
NVLink
Apply
In office • Full-Time
AI/ML
Together AI
Apply
$59k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
DevOps
HPC
Apply
$59k – $129k per year (Estimated) • In office • Full-Time • 10+ years exp • Yokohama
DevOps
HPC
Apply
$33k – $71k per year (Estimated) • In office • PhD • Yokohama
AI/ML
Text-to-Speech
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vercel
Apply
$45k – $92k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokohama
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
CUDA
CUDA Toolkit
NVLink
Apply
$59k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
DevOps
HPC
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.