368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$63k – $139k per year (Estimated)
Location
Remote/Hybrid (Yokohama, Japan)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.

About ai&

ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI. Our vision is twofold: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider. We are building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services. We believe that the most effective way to build and scale AI is to own the stack from top to bottom.

At ai&, we empower small teams with the autonomy needed to tackle significant challenges. Our approach is to deconstruct large problems into manageable components and solve complex issues collaboratively. We seek highly motivated, mission-driven individuals who demonstrate strong personal agency. We value curiosity as the foundation of talent, and we are looking for people eager to develop alongside our evolving technology and expanding business.

We are actively hiring worldwide, with presence in Tokyo, SF, Austin, and Toronto. We are more than happy to meet exceptional talent where they are.

As an inference & serving engineer, your objective is to build a high-performance, multi-tenant serving stack that squeezes maximum utilization out of heterogeneous hardware. This involves navigating the trade-offs between various state-of-the-art inference frameworks and engines, selecting and optimizing the right runtime for the right workload. The scope of work is not limited to Large Language Models; it extends to the frontier of Generative AI, including high-throughput Video generation and complex Multimodal systems where memory pressure and compute requirements are significantly more demanding.

Beyond just deploying models at scale, this role is responsible for building a robust system that bridges the gap between boutique, high-performance clusters and massive, multi-node deployments as the company grows. This requires a deep understanding of the "Inference Triangle"-constantly tuning the stack to find the optimal equilibrium between low-latency (TTFT/ITL), high-throughput, and inference quality (Precision/Quantization). The ideal candidate is a hands-on engineer who views the entire GPU fleet as a single, programmable compute fabric and is eager to get their hands dirty at every level of the stack.

Responsibilities:

  • Runtime Selection & Deep Optimization: Lead the evaluation, integration, and continuous tuning of diverse inference frameworks to ensure best-in-class performance across LLM, Video, and Multimodal workloads.

  • Latency & Throughput Engineering: Own the end-to-end performance profile of the model lifecycle, implementing advanced strategies such as disaggregated prefill/decode, speculative decoding, and continuous batching to minimize TTFT and maximize tokens-per-second.

  • Scalable Systems Evolution: Design and implement serving architectures that function seamlessly on small experimental clusters while providing a clear, robust path to massive-scale, multi-node deployments.

  • Advanced Memory & Cache Orchestration: Implement and optimize memory management techniques to maximize KV-cache reuse and minimize redundant computations in multi-turn or high-concurrency scenarios.

  • Day 0 Model Support: Working with the ecosystem, craft a Day 0 model support strategy ensuring our stack provides stable, high-performance support for new models when they are released.

  • Cross-Stack Integration: Collaborate with the Backend/Gateway and Compute Orchestration teams to ensure the inference engine’s telemetry, failure domains, and lifecycle management are perfectly aligned with the global load balancer and API layers.

  • Hands-on Technical Leadership: Maintain a high level of personal agency by writing production code, debugging complex distributed system "hangs," and contributing to architectural decisions in a flat, fast-moving team environment.

  • Collaborative Communication: Function as a primary technical peer to engineering leads, translating complex hardware and model constraints into clear product and infrastructure strategies.

  • Inference Strategy & Trade-offs: Define path forward when balancing model precision and quantization against the physical limits of HBM bandwidth and compute throughput

You may be a fit if you have the following skills:

  • Inference Engine: Deep experience with the internals of modern runtimes. You are a prominent contributor to inference engine ecosystems, including but not limited to OSS projects or proprietary engines at top-tier AI labs.

  • Multimodal Domain Knowledge: Understanding of the specific challenges involved in serving Large Language Models alongside Video and Vision-based generative models.

  • Scale-First Engineering: A track record of building and managing distributed systems that have evolved from small-scale proofs-of-concept to large-scale production deployments.

  • Great Team Spirit: A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokohama
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
up to $63k per year (gross) • In office • Full-Time • 5+ years exp • Moscow
SQL
Databases
Apache Kafka
AI/ML
LLM
RAG
DevOps
CI/CD
Git
gRPC
WebSockets
QA
Postman
Swagger
Apply
$43k – $103k per year (Estimated) • In office • Full-Time • 5+ years exp • Moscow
AI/ML
LLM
RAG
Apply
$32k – $77k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Moscow
Python
Databases
FAISS
AI/ML
Hadoop
Hugging Face
LangChain
LLM
NLP
PyTorch
smolagents
Spark
Apply
$62k – $142k per year (Estimated) • In office • Full-Time • 6+ years exp • PhD • Madrid
Apex
JavaScript
Python
TypeScript
Databases
Databricks
Google BigQuery
Snowflake
AI/ML
Agentforce
AI Agents
Claude
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
Marketing
Salesforce
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
CUDA
CUDA Toolkit
NVLink
Apply
$59k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
DevOps
HPC
Apply
$59k – $129k per year (Estimated) • In office • Full-Time • 10+ years exp • Yokohama
DevOps
HPC
Apply
$61k – $133k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
InfiniBand
DevOps
CI/CD
GitOps
Kubernetes
Prometheus
Terraform
Apply
$33k – $71k per year (Estimated) • In office • PhD • Yokohama
AI/ML
Text-to-Speech
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vercel
Apply
$45k – $92k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokohama
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
Python
AI/ML
DeepSpeed
LLM
PyTorch
Reinforcement Learning
Synthetic Data
vLLM
FSDP
Post-training
SFT
Apply
$63k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
AI/ML
CUDA
CUDA Toolkit
NVLink
Apply
$59k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • Yokohama
DevOps
HPC
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.