826,646open jobs
53,236companies
135,873added this week
Browse all
Salary
≈ $46k – $122k per year (Estimated)
Location
In office (Tokyo)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Aug 26, 2026. Rakuten scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Headquartered in Tokyo, Japan, Rakuten is a major global technology conglomerate that operates a diverse ecosystem of internet and digital services. The company is widely known for its massive e-commerce platforms, alongside a robust portfolio spanning financial technology and digital content. Furthermore, it has established itself as an innovative force in the telecommunications industry through the development of cloud-native mobile network infrastructure.

Job Description:

Business Overview

AI & Data Division (AIDD) spearheads data science & AI initiatives by leveraging data from Rakuten Group. We build a platform for large-scale field experimentations using cutting-edge technologies to provide critical insights that enable faster and better and faster contribution for our business. Our division boasts an international culture created by talented employees from around the world. Following the strategic vision “Rakuten as a data-driven membership company”, AIDD is expanding its data & AI related activities across multiple Rakuten Group companies.

Department Overview

GPU Optimization Department (GPUOD)is responsible for the strategic management, optimization, and governance of Rakuten's company-wide AI infrastructure, ensuring high-performance, cost-efficient utilization of compute resources for machine learning workloads. We oversee a large-scale hybrid infrastructure spanning thousands of accelerators, including the latest Hopper and upcoming Blackwell architectures.

As a central enabler for AI innovation, we:

- Optimize compute resource allocation across on-premises and multi-cloud environments, maximizing efficiency for training and inference workloads.

- Manage hybrid orchestration of diverse accelerator resources, ensuring seamless scalability and cost-effective deployment.

- Develop and enhance frameworks for large-scale distributed training, with special focus on LLMs and generative AI.

- Optimize inference performance through model optimization techniques and system-level acceleration.

- Collaborate with internal teams to deliver scalable, high-availability inference services tailored to business needs.

- Continuously evaluate next-generation hardware solutions, including specialized AI chips optimized for LLM workloads.

- By effectively managing both conventional and specialized compute resources across on-premises and cloud environments, our team ensures Rakuten's AI ecosystem remains at the forefront of performance, reliability, and cost-efficiency.

Position:

Why We Hire

- Work on cutting-edge LLM training & inference optimization at scale.

- Directly impact Rakuten’s AI infrastructure by improving efficiency and reducing costs.

- Collaborate with global AI/ML teams on high-impact challenges.

- Opportunity to research and implement state-of-the-art GPU optimizations.

Position Details

As a GPU Training & Inference Optimization Engineer, you will focus on maximizing the performance, efficiency, and scalability of LLM training and inference workloads on Rakuten’s GPU clusters. You will deeply optimize training frameworks (e.g., PyTorch, DeepSpeed, FSDP) and inference engines (e.g., vLLM, TensorRT-LLM, Triton, SGLang), ensuring Rakuten’s AI models run at peak efficiency.

This role requires strong expertise in GPU-accelerated ML frameworks, distributed training, and inference optimization, with a focus on reducing training time, improving GPU utilization, and minimizing inference latency.

Key Responsibilities

- Optimize LLM training frameworks (e.g., PyTorch, DeepSpeed, Megatron-LM, FSDP) to maximize GPU utilization and reduce training time.

- Profile and optimize distributed training bottlenecks (e.g., NCCL issues, CUDA kernel efficiency, communication overhead).

- Implement and tune inference optimizations (e.g., quantization, dynamic batching, KV caching) for low-latency, high-throughput LLM serving (vLLM, TensorRT-LLM, Triton, SGLang).

- Collaborate with infrastructure teams to improve GPU cluster scheduling, resource allocation, and fault tolerance for large-scale training jobs.

- Develop benchmarking tools to measure and improve training throughput, memory efficiency, and inference latency.

- Research and apply cutting-edge techniques (e.g., mixture-of-experts, speculative decoding) to optimize LLM performance.

Mandatory Qualifications:

- 3+ years of hands-on experience in GPU-accelerated ML training & inference optimization, preferably for LLMs or large-scale deep learning models.

- Deep expertise in PyTorch, DeepSpeed, FSDP, or Megatron-LM, with experience in distributed training optimizations.

- Strong knowledge of LLM inference optimizations (e.g., quantization, pruning, KV caching, continuous batching).

- Bachelor’s or higher degree in Computer Science, Engineering, or related field.

Desired Qualifications:

- Proficiency in CUDA, Triton kernel, NVIDIA tools (Nsight, NCCL), and performance profiling (e.g., PyTorch Profiler, TensorBoard).

- Experience with LLM-specific optimizations (e.g., FlashAttention, PagedAttention, LoRA, speculative decoding).

- Familiarity with Kubernetes (K8s) for GPU workloads (e.g., KubeFlow, Volcano).

- Contributions to open-source ML frameworks (e.g., PyTorch, DeepSpeed, vLLM).

- Experience with inference serving frameworks (e.g., vLLM, TensorRT-LLM, Triton, Hugging Face TGI).

#engineer #applicationsengineer #AI #aianddatadiv

Languages:

English (Overall - 3 - Advanced)
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
826,646 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Tokyo
$32k – $45k per year • Remote (Japan) • Full-Time
C++
AI/ML
AWS Bedrock
DevOps
Azure
AWS
Management
Slack
Apply
$26k – $32k per year • Remote (Japan) • Full-Time
Python
C++
AI/ML
AWS Bedrock
OpenAI
Amazon SageMaker
DevOps
AWS
Docker
GitHub
Management
Slack
Apply
$29k – $39k per year • Remote (Japan) • Full-Time
C++
DevOps
Azure
AWS
Management
Slack
Apply
$39k – $65k per year • Remote (Japan) • Full-Time
C++
DevOps
Azure
AWS
Management
Slack
Apply
$34k – $52k per year • Remote (Japan) • Full-Time
JavaScript
Java
C#
C++
Frontend
Vue.js
Apply
≈ $105k – $272k per year (Estimated) • In office • Full-Time • Master's Degree • London
Python
AI/ML
DeepSpeed
LoRA
RLHF
PEFT
PyTorch
DPO
SFT
Post-training
FSDP
NCCL
Machine Learning
Apply
$200k per year • Hybrid • 3+ years exp • New York
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
CUDA Toolkit
Reinforcement Learning
JAX
PyTorch
Ray
Time Series Forecasting
CUDA
Triton
Megatron-LM
FSDP
NCCL
InfiniBand
NVLink
cuDNN
XLA
CUTLASS
Machine Learning
DevOps
SLURM
Kubernetes
HPC
Apply
$84k – $120k per year • In office • Internship • Master's Degree • Menlo Park • Bellevue
Python
Go
Java
SQL
Databases
Snowflake
AI/ML
LangGraph
LangChain
CUDA Toolkit
Multimodal AI
AI Agents
NLP
PyTorch
LLM
RAG
Federated Learning
CUDA
Feature Store
Agentic Workflows
Machine Learning
DevOps
Prometheus
Cortex
Apply
DSP SW Engineer 3 days ago
$100k – $200k per year • In office • Full-Time • Santa Clara
Python
C++
SystemC
MATLAB
AI/ML
Quantization
Apply
$87k – $157k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Tewksbury
SQL
C++
C++
OpenGL
Qt
AI/ML
CUDA Toolkit
CUDA
DevOps
Red Hat
CI/CD
Linux
Management
Agile
Scrum
Apply
≈ $51k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Tokyo
Python
AI/ML
Machine Learning
DevOps
Azure
AWS
Linux
Analytics
Tableau
Power BI
ETL/ELT
Apply
≈ $43k – $108k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Tokyo
Apply
≈ $40k – $108k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tokyo
Apply
≈ $48k – $102k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Tokyo
JavaScript
Java
Kotlin
TypeScript
Frontend
Angular
DevOps
GCP
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Linux
Unix
Apply
≈ $18k – $38k per year (Estimated) • In office • Internship • Tokyo
Analytics
Microsoft Excel
Apply
≈ $38k – $86k per year (Estimated) • In office • Full-Time • Tokyo
Apply
≈ $43k – $90k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tokyo
Apply
≈ $19k – $36k per year (Estimated) • In office • Internship • Bachelor's Degree • Tokyo
Apply
In office • Full-Time • 5+ years exp • Tokyo
Apply
See all jobs
This is one of many
826,646 more open roles from verified company boards, updated every day.