368,530open jobs
9,432companies
50,439added this week
Browse all
Location
In office
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

Role overview

As a Network Engineer at ai&, you are the domain expert on the lossless networking fabrics that tie our GPU fleet together. AI at scale lives and dies on the network. Collective communication operations, AllReduce, AllGather, ReduceScatter, are on the critical path of every distributed training and inference workload we run. Your job is to make sure the fabric is fast, lossless, and never the bottleneck.

You will work across RoCE v2 and InfiniBand fabrics, tune NCCL and network interfaces, and own the end-to-end network performance of our compute clusters. You will work closely with the systems, kernel, and inference teams to ensure that what gets built at the physical layer translates directly into performance at the workload layer.

Responsibilities

  • Lossless Fabric Design & Operations Design, deploy, and operate lossless networking fabrics across our data centers. Own RoCE v2 and InfiniBand (NDR/XDR) deployments end to end.

  • NCCL & Interface Tuning Tune NCCL, NICs, and DPUs to guarantee maximum bandwidth and zero packet loss for distributed AI workloads. Own the performance of collective communication operations across the fleet.

  • Network Architecture Design the network architecture for new data center deployments. Make topology, switch, and cabling decisions that scale from current clusters to future multi-site deployments.

  • Performance Monitoring & Optimization Instrument the network for observability. Proactively identify and eliminate bottlenecks before they affect workloads. Own network performance benchmarks and drive continuous improvement.

  • Cross-Team Collaboration Work closely with the systems, storage, and ML infrastructure teams to ensure the network fabric supports the demands of distributed training and inference at every scale.

You may be a fit if you have the following skills

  • AI Networking Expertise Deep experience designing and operating lossless AI networking fabrics. You have worked with InfiniBand and RoCE v2 at scale and you understand the trade-offs between them.

  • NCCL & Collective Communications Hands-on experience tuning NCCL for distributed AI workloads. You understand how collective communication patterns interact with network topology and you know how to optimize for both bandwidth and latency.

  • NIC & DPU Proficiency Experience configuring and tuning high-performance NICs and DPUs from vendors including NVIDIA ConnectX and Bluefield series.

  • Network Architecture Judgment You make network design decisions that hold up at scale. Fat-tree topologies, rail-optimized designs, congestion control - you have an informed view on all of it.

  • Great Team Spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$30k – $78k per year (Estimated) • In office • 5+ years exp
Go
Python
Rust
TypeScript
Databases
Apache Kafka
NATS
AI/ML
AI Agents
LLM
RAG
Semantic Search
Together AI
Function Calling
Knowledge Graph
Semantic Search
DevOps
ArgoCD
GitOps
Grafana
Incident Management
Kubernetes
Prometheus
Management
Slack
Apply
Data Engineer 9 days ago
In office • Contractor • Christchurch
Python
SQL
Databases
ClickHouse
DynamoDB
AI/ML
Dagster
dbt
Together AI
DevOps
AWS
AWS Lambda
CI/CD
Git
Amazon EventBridge
Amazon S3
Analytics
Power BI
Tableau
Apply
$117k – $257k per year (Estimated) • In office • Full-Time • London
Python
TypeScript
AI/ML
AI Agents
Fine-tuning
LLM
Prompt Engineering
RAG
Together AI
Edge AI
LLM Evaluation
Text-to-Speech
Function Calling
DevOps
Kubernetes
Apply
$200k – $250k per year • Remote • Full-Time • 5+ years exp • San Francisco
Python
SQL
AI/ML
Together AI
InfiniBand
DevOps
HPC
Apply
$38k – $85k per year (Estimated) • In office • 10+ years exp • Master's Degree • Bengaluru
Python
Python
FastAPI
Pydantic
pySpark
Databases
OpenSearch
pgvector
PostgreSQL
AI/ML
AI Agents
AWS Bedrock
Claude
Fine-tuning
LangChain
LangGraph
LLM
NLP
Polars
PyTorch
RAG
Together AI
Spark
Amazon SageMaker
Anthropic
Hugging Face
OpenAI
DevOps
AWS
AWS Lambda
CI/CD
GitLab CI
Terraform
Vector
Amazon EventBridge
GitLab
Apply
$61k – $133k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
InfiniBand
DevOps
CI/CD
GitOps
Kubernetes
Prometheus
Terraform
Apply
Software Engineer 29 days ago
$32k – $57k per year • Remote • Full-Time • 3+ years exp • Bachelor's Degree • Tokyo
DevOps
AWS
CI/CD
Docker
Kubernetes
Apply
$46k – $133k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
LLM
AI Agents
Apply
$47k – $111k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
InfiniBand
NVLink
DevOps
HPC
Apply
$59k – $129k per year (Estimated) • In office • Full-Time • 10+ years exp • Yokohama
DevOps
HPC
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.