368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$63k – $138k per year (Estimated)
Location
Remote/Hybrid (Yokohama, Japan)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview

As a Systems Engineer at ai&, you are responsible for the physical and software foundation that everything else runs on. You will plan, configure, and manage the bare-metal infrastructure that powers our data centers - from OS tuning and driver management to rack-scale GPU system provisioning. You are the person who makes sure the hardware is running at its full potential before the software teams ever touch it.

This is a hands-on role. You will work on some of the most advanced compute hardware available, including NVL72 and AMD Helios rack-scale systems, and you will be responsible for keeping them running at maximum efficiency. You think carefully about system configuration, firmware, and the low-level software decisions that compound into real performance differences at scale.

Responsibilities

  • Bare-Metal Infrastructure Management Configure and manage bare-metal servers end to end. Own OS tuning, driver management, firmware upgrades, and CUDA configuration across the fleet.

  • Rack-Scale GPU System Operations Lead the installation, provisioning, and continuous operation of high-density, liquid-cooled rack-scale GPU systems including NVL72 and AMD Helios deployments.

  • System Architecture & Planning Plan and architect the next generation of system configurations including compute, storage, networking interconnects, routers, and switches. Make decisions that scale.

  • Performance Optimization Tune system-level configurations to maximize hardware utilization and minimize overhead. Work closely with the kernel and inference teams to ensure software and hardware are fully aligned.

  • Cross-Team Collaboration Work closely with the network, storage, and data center teams to ensure the physical infrastructure operates as a unified, high-performance system.

You may be a fit if you have the following skills

  • Bare-Metal Operations Experience Deep hands-on experience managing large-scale bare-metal server environments. You have configured OS, drivers, firmware, and CUDA at scale and you know the failure modes.

  • GPU System Expertise Experience provisioning and operating high-density GPU systems. Familiarity with NVIDIA NVLink, NVSwitch, and AMD MI-series architectures is a strong signal.

  • Low-Level Systems Knowledge Strong understanding of Linux internals, kernel parameters, NUMA topology, PCIe configurations, and how these interact with AI workloads.

  • Infrastructure Judgment You make system configuration decisions that hold up at scale. You think about maintainability, reproducibility, and failure recovery from the start.

  • Great Team Spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokohama
$69k – $139k per year (Estimated) • Remote/Hybrid • Internship • 5+ years exp • Berlin
Bash
Python
AI/ML
CUDA Toolkit
KServe
LLM
Quantization
TensorRT
TensorRT-LLM
vLLM
CUDA
Triton
InfiniBand
NCCL
NVLink
DevOps
Ansible
Kubernetes
Red Hat
Terraform
Ubuntu
CI/CD
Git
Cybersecurity
ISO 27001
Apply
Data Scientist 1 day ago
$25k – $56k per year (Estimated) • Remote • Bachelor's Degree • Moscow
C++
Python
SQL
AI/ML
Computer Vision
CUDA
CUDA Toolkit
TensorRT
DevOps
Docker
Git
Kubernetes
Apply
In office • Full-Time
Python
AI/ML
CUDA
CUDA Toolkit
DevOps
Docker
Kubernetes
PagerDuty
SLI/SLO/SLA
SLURM
HPC
Management
Linear
Marketing
Zendesk
Apply
$87k – $157k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Huntsville
C++
MATLAB
Verilog
VHDL
AI/ML
CUDA
CUDA Toolkit
DevOps
CI/CD
Apply
$31k – $66k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Moscow
C++
AI/ML
CUDA
CUDA Toolkit
DevOps
Docker
Git
Game Dev
Godot
Robotics
ROS2
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
$59k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
DevOps
HPC
Apply
$59k – $129k per year (Estimated) • In office • Full-Time • 10+ years exp • Yokohama
DevOps
HPC
Apply
$61k – $133k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
InfiniBand
DevOps
CI/CD
GitOps
Kubernetes
Prometheus
Terraform
Apply
Software Engineer 29 days ago
$32k – $57k per year • Remote • Full-Time • 3+ years exp • Bachelor's Degree • Tokyo
DevOps
AWS
CI/CD
Docker
Kubernetes
Apply
$33k – $71k per year (Estimated) • In office • PhD • Yokohama
AI/ML
Text-to-Speech
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vercel
Apply
$45k – $92k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokohama
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
$59k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
DevOps
HPC
Apply
$59k – $129k per year (Estimated) • In office • Full-Time • 10+ years exp • Yokohama
DevOps
HPC
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.