660,302open jobs
38,440companies
97,349added this week
Browse all
Salary
$71k – $162k per year (Estimated)
Location
Remote/Hybrid (Yokohama, Japan)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

ai&

Ai& runs frontier open-weight models - Kimi, DeepSeek and GLM - on GPU infrastructure it owns and operates in Japan. One API compatible with the OpenAI and Anthropic SDKs, up to 80% lower cost, sub-50ms latency in Japan and zero cross-border data egress.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview

As a Systems Engineer at ai&, you are responsible for the physical and software foundation that everything else runs on. You will plan, configure, and manage the bare-metal infrastructure that powers our data centers - from OS tuning and driver management to rack-scale GPU system provisioning. You are the person who makes sure the hardware is running at its full potential before the software teams ever touch it.

This is a hands-on role. You will work on some of the most advanced compute hardware available, including NVL72 and AMD Helios rack-scale systems, and you will be responsible for keeping them running at maximum efficiency. You think carefully about system configuration, firmware, and the low-level software decisions that compound into real performance differences at scale.

Responsibilities

  • Bare-Metal Infrastructure Management Configure and manage bare-metal servers end to end. Own OS tuning, driver management, firmware upgrades, and CUDA configuration across the fleet.

  • Rack-Scale GPU System Operations Lead the installation, provisioning, and continuous operation of high-density, liquid-cooled rack-scale GPU systems including NVL72 and AMD Helios deployments.

  • System Architecture & Planning Plan and architect the next generation of system configurations including compute, storage, networking interconnects, routers, and switches. Make decisions that scale.

  • Performance Optimization Tune system-level configurations to maximize hardware utilization and minimize overhead. Work closely with the kernel and inference teams to ensure software and hardware are fully aligned.

  • Cross-Team Collaboration Work closely with the network, storage, and data center teams to ensure the physical infrastructure operates as a unified, high-performance system.

You may be a fit if you have the following skills

  • Bare-Metal Operations Experience Deep hands-on experience managing large-scale bare-metal server environments. You have configured OS, drivers, firmware, and CUDA at scale and you know the failure modes.

  • GPU System Expertise Experience provisioning and operating high-density GPU systems. Familiarity with NVIDIA NVLink, NVSwitch, and AMD MI-series architectures is a strong signal.

  • Low-Level Systems Knowledge Strong understanding of Linux internals, kernel parameters, NUMA topology, PCIe configurations, and how these interact with AI workloads.

  • Infrastructure Judgment You make system configuration decisions that hold up at scale. You think about maintainability, reproducibility, and failure recovery from the start.

  • Great Team Spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
660,302 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokohama
$87k – $198k per year • In office • TS/SCI • Full-Time • 8+ years exp • Bachelor's Degree • Dayton
Python
JavaScript
C++
AI/ML
CUDA Toolkit
AI Agents
Anomaly Detection
CUDA
DevOps
CI/CD
Jenkins
Docker
Kubernetes
HPC
Robotics
Sensor Fusion
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
CUDA Toolkit
JAX
PyTorch
RAG
CUDA
DevOps
Docker
Kubernetes
HPC
Apply
$224k – $357k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • New York • Santa Clara
AI/ML
CUDA Toolkit
CUDA
Apply
$82k – $198k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree
AI/ML
vLLM
CUDA Toolkit
Reinforcement Learning
JAX
Multimodal AI
SGLang
TensorRT
TensorRT-LLM
Transformers
PyTorch
CUDA
Post-training
NVIDIA NeMo
NCCL
CUTLASS
World Models
Robotics
Reinforcement Learning
Apply
$51k – $128k per year (Estimated) • Remote • 8+ years exp • Islamabad
JavaScript
Node JS
Prolog
AI/ML
vLLM
CUDA Toolkit
SGLang
TensorRT
TensorRT-LLM
Kubeflow
Tokenization
CUDA
NCCL
InfiniBand
Frontend
Bootstrap
React.js
DevOps
Cilium
Loki
Prometheus
SLURM
GitOps
Kubernetes
Grafana
kubeadm
KVM
QEMU
KubeVirt
HPC
Web3
Bitcoin
Management
Telegram
WhatsApp
Apply
$68k – $155k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
InfiniBand
DevOps
Terraform
Prometheus
CI/CD
GitOps
Kubernetes
Apply
$20k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Yokohama
AI/ML
Copilot
ChatGPT
Gemini
Management
Notion
Google Workspace
Apply
$18k – $42k per year (Estimated) • In office • Full-Time • 1+ year exp • Yokohama
AI/ML
Copilot
Claude
ChatGPT
Edge AI
Management
Notion
Google Workspace
Marketing
Salesforce
Apply
Remote/Hybrid • Full-Time • Yokohama
Apply
Developer Relations 5 months ago
$47k – $100k per year (Estimated) • In office • Full-Time • Yokohama
Python
DevOps
GitHub
Management
Discord
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Yokohama
Apply
$57k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
Post-training
Apply
Clinician 1 day ago
In office • Full-Time • Yokohama
Apply
Process Engineer 1 day ago
In office • Full-Time • Yokohama
Apply
In office • Full-Time • Yokohama
Apply
See all jobs
This is one of many
660,302 more open roles from verified company boards, updated every day.