660,302open jobs
38,440companies
97,349added this week
Browse all
Location
In office
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

ai&

Ai& runs frontier open-weight models - Kimi, DeepSeek and GLM - on GPU infrastructure it owns and operates in Japan. One API compatible with the OpenAI and Anthropic SDKs, up to 80% lower cost, sub-50ms latency in Japan and zero cross-border data egress.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview

As a Network Engineer at ai&, you are the domain expert on the lossless networking fabrics that tie our GPU fleet together. AI at scale lives and dies on the network. Collective communication operations, AllReduce, AllGather, ReduceScatter, are on the critical path of every distributed training and inference workload we run. Your job is to make sure the fabric is fast, lossless, and never the bottleneck.

You will work across RoCE v2 and InfiniBand fabrics, tune NCCL and network interfaces, and own the end-to-end network performance of our compute clusters. You will work closely with the systems, kernel, and inference teams to ensure that what gets built at the physical layer translates directly into performance at the workload layer.

Responsibilities

  • Lossless Fabric Design & Operations Design, deploy, and operate lossless networking fabrics across our data centers. Own RoCE v2 and InfiniBand (NDR/XDR) deployments end to end.

  • NCCL & Interface Tuning Tune NCCL, NICs, and DPUs to guarantee maximum bandwidth and zero packet loss for distributed AI workloads. Own the performance of collective communication operations across the fleet.

  • Network Architecture Design the network architecture for new data center deployments. Make topology, switch, and cabling decisions that scale from current clusters to future multi-site deployments.

  • Performance Monitoring & Optimization Instrument the network for observability. Proactively identify and eliminate bottlenecks before they affect workloads. Own network performance benchmarks and drive continuous improvement.

  • Cross-Team Collaboration Work closely with the systems, storage, and ML infrastructure teams to ensure the network fabric supports the demands of distributed training and inference at every scale.

You may be a fit if you have the following skills

  • AI Networking Expertise Deep experience designing and operating lossless AI networking fabrics. You have worked with InfiniBand and RoCE v2 at scale and you understand the trade-offs between them.

  • NCCL & Collective Communications Hands-on experience tuning NCCL for distributed AI workloads. You understand how collective communication patterns interact with network topology and you know how to optimize for both bandwidth and latency.

  • NIC & DPU Proficiency Experience configuring and tuning high-performance NICs and DPUs from vendors including NVIDIA ConnectX and Bluefield series.

  • Network Architecture Judgment You make network design decisions that hold up at scale. Fat-tree topologies, rail-optimized designs, congestion control - you have an informed view on all of it.

  • Great Team Spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
660,302 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$48k – $120k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree
JavaScript
Node JS
Prolog
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
TensorRT
TensorRT-LLM
Kubeflow
CUDA
NCCL
InfiniBand
Frontend
Bootstrap
DevOps
Cilium
Loki
Prometheus
SLURM
GitOps
Kubernetes
Grafana
kubeadm
KVM
QEMU
KubeVirt
HPC
Apply
$74k – $187k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree
JavaScript
Node JS
Prolog
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
TensorRT
TensorRT-LLM
Kubeflow
CUDA
NCCL
InfiniBand
Frontend
Bootstrap
DevOps
Cilium
Loki
Prometheus
SLURM
GitOps
Kubernetes
Grafana
kubeadm
KVM
QEMU
KubeVirt
HPC
Apply
$85k – $213k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree
JavaScript
Node JS
Prolog
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
TensorRT
TensorRT-LLM
Kubeflow
CUDA
NCCL
InfiniBand
Frontend
Bootstrap
DevOps
Cilium
Loki
Prometheus
SLURM
GitOps
Kubernetes
Grafana
kubeadm
KVM
QEMU
KubeVirt
HPC
Apply
$168k – $270k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
vLLM
CUDA Toolkit
SGLang
TensorRT
TensorRT-LLM
PyTorch
CUDA
NCCL
DevOps
Splunk
Terraform
Puppet
Ansible
GCP
OpenTelemetry
Chef
Prometheus
Azure
AWS
Kubernetes
Grafana
Argo Workflows
KubeVirt
Incident Management
AWS Step Functions
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara • Westford • Austin • Durham
Python
AI/ML
InfiniBand
DevOps
Cilium
Envoy
SLURM
Kubernetes
Cloudflare
HPC
Cybersecurity
Calico
Apply
$57k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
Post-training
Apply
$68k – $155k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
InfiniBand
DevOps
Terraform
Prometheus
CI/CD
GitOps
Kubernetes
Apply
$20k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Yokohama
AI/ML
Copilot
ChatGPT
Gemini
Management
Notion
Google Workspace
Apply
$18k – $42k per year (Estimated) • In office • Full-Time • 1+ year exp • Yokohama
AI/ML
Copilot
Claude
ChatGPT
Edge AI
Management
Notion
Google Workspace
Marketing
Salesforce
Apply
Remote/Hybrid • Full-Time • Yokohama
Apply
See all jobs
This is one of many
660,302 more open roles from verified company boards, updated every day.