659,077open jobs
38,403companies
95,422added this week
Browse all
Salary
$68k – $155k per year (Estimated)
Location
Remote/Hybrid (Yokohama, Japan)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

ai&

Ai& runs frontier open-weight models - Kimi, DeepSeek and GLM - on GPU infrastructure it owns and operates in Japan. One API compatible with the OpenAI and Anthropic SDKs, up to 80% lower cost, sub-50ms latency in Japan and zero cross-border data egress.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview

The Platform team turns raw heterogeneous compute into a serving platform. ai& owns its data centers and runs AMD, NVIDIA, and Tenstorrent silicon side by side. Your job is everything between the bare metal and the inference engines: cluster orchestration, node lifecycle, scaling, networking, observability, and reliability.

This is not cloud consumption. When capacity is short, you add nodes we own. When a fabric misbehaves, you debug it down to the switch. The platform must let a small team operate hundreds of nodes across multiple sites without heroics, and it must scale by an order of magnitude over the next two years as new sites come online.

You will work directly with the inference team, which owns the engines and serving gateway, and the data center team, which owns power, cooling, and physical deployment. You own the layer that makes their work composable.

Responsibilities

  • Compute orchestration Run Kubernetes across GPU clusters in ai&-owned data centers. Own node lifecycle from bring-up and burn-in through drain and repair, across multiple accelerator vendors.

  • Scaling Build the capacity and scheduling machinery that places inference workloads across heterogeneous silicon and multiple sites, and that lets us bring a new site from empty racks to serving traffic on a predictable timeline.

  • Reliability Define and hold SLOs for the platform. Build the observability stack (metrics, logs, tracing, alerting) and the failure isolation that keeps one bad node or one bad rollout from becoming an incident.

  • Networking and data Operate high-bandwidth fabrics for multi-node inference. Solve model weight distribution: getting hundreds of gigabytes onto the right nodes fast, every time a model ships.

  • Deployment machinery Own CI/CD and GitOps for the fleet. Infrastructure as code, reproducible node images, safe rollouts.

You may be a fit if you have the following skills

  • Production Kubernetes at scale You have operated large multi-cluster Kubernetes environments, ideally with GPU scheduling, device plugins, and topology-aware placement.

  • Systems depth Strong Linux fundamentals. You can reason about NUMA, PCIe, NICs, and storage, and you debug from symptoms to root cause without guessing.

  • Networking fundamentals You understand L2/L3, and ideally RDMA fabrics (InfiniBand or RoCE) in production.

  • Infrastructure as code Terraform or equivalent, GitOps workflows, and the discipline to keep the fleet reproducible.

  • Ownership under load You have carried a pager for systems that matter and you build so the pager stays quiet.

  • Relevant tooling Go or Python, Prometheus-family observability, and comfort automating anything you do twice.

  • Great team spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
659,077 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokohama
$140k – $210k per year • Equity • In office • Full-Time • 5+ years exp • Sunnyvale
Python
PowerShell
Bash
DevOps
Terraform
GitHub Actions
CI/CD
Jenkins
AWS
Cybersecurity
Metasploit
ISO 27001
Wiz
SOC 2
Management
Jira
Apply
DevSecOps Engineer 4 hours ago
$79k – $160k per year • In office • Top Secret • 4+ years exp • Washington
Python
JavaScript
Java
PHP
SQL
PowerShell
Node JS
Bash
Java
Apache Tomcat
PHP
Drupal
AI/ML
Claude
Claude Code
Frontend
React.js
DevOps
GitHub Actions
CloudFormation
GitLab CI
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Nginx
Amazon EKS
GitLab
Amazon ECS
Apply
$104k – $158k per year • In office • Full-Time • 10+ years exp • Charlotte • Jersey City • Plano
Python
Java
SQL
Scala
Python
pySpark
Databases
Db2
Oracle
Apache Iceberg
Apache Kafka
AI/ML
Hadoop
Spark
Flink
DevOps
GCP
Azure
CI/CD
AWS
Bitbucket
Management
Jira
ServiceNow
Agile
Apply
$210k – $330k per year • Equity • In office • Full-Time • 15+ years exp • Bachelor's Degree • Berkeley Heights
Python
Python
pySpark
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
Amazon Redshift
Microsoft Fabric
BigQuery
AI/ML
Spark
RAG
Analytics
Tableau
Power BI
Cognos
Collibra
Apply
$50k – $126k per year (Estimated) • Remote • Contractor
JavaScript
Mobile
JUnit
DevOps
Azure DevOps
GitLab CI
Azure
CI/CD
Jenkins
Git
GitLab
Management
Agile
Scrum
Kanban
QA
TestNG
Selenium
Cucumber
Apply
$57k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
Post-training
Apply
$20k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Yokohama
AI/ML
Copilot
ChatGPT
Gemini
Management
Notion
Google Workspace
Apply
$18k – $42k per year (Estimated) • In office • Full-Time • 1+ year exp • Yokohama
AI/ML
Copilot
Claude
ChatGPT
Edge AI
Management
Notion
Google Workspace
Marketing
Salesforce
Apply
Remote/Hybrid • Full-Time • Yokohama
Apply
Developer Relations 5 months ago
$47k – $100k per year (Estimated) • In office • Full-Time • Yokohama
Python
DevOps
GitHub
Management
Discord
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Yokohama
Apply
$57k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
Post-training
Apply
Clinician 1 day ago
In office • Full-Time • Yokohama
Apply
Process Engineer 1 day ago
In office • Full-Time • Yokohama
Apply
In office • Full-Time • Yokohama
Apply
See all jobs
This is one of many
659,077 more open roles from verified company boards, updated every day.