743,556open jobs
44,602companies
107,147added this week
Browse all
Salary
≈ $52k – $129k per year (Estimated)
Location
Hybrid (Tokyo, Japan)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026.

Overview
Company
Impact
Profile match
Preferred Networks is a Japanese artificial intelligence company founded in Tokyo in 2014 and long regarded as the country's most technically ambitious AI startup. It created Chainer, one of the first define-by-run deep learning frameworks and an influence on the design of modern autograd systems, and has since built a full vertical stack from its own MN-Core accelerator silicon up to foundation models. The company applies that stack to manufacturing, robotics, materials science through the Matlantis simulator and healthcare, working closely with Japanese industrial partners including Toyota, Fanuc and Hitachi.

Job Description

Preferred Networks (PFN) では、大規模言語モデル (LLM) を中心としたマルチモーダルな基盤モデルの開発を進めています。基盤モデル開発に加え、エンタメ、科学計算、ロボットといった自社事業領域を強化する形での導入を目指した応用開発や、多様な産業への応用研究も行っています。

推論最適化チームのミッションは、最先端のモデルと現実世界のプロダクトを結ぶ「架け橋」となることです。自社開発のLLMのポテンシャルを極限まで引き出し、実運用環境で求められる厳しい品質・速度・コスト要件を満たす形で最適化し提供することを通じて、PFNのプロダクトやソリューション事業の拡大を技術の出口から支援します。代表的な業務内容の例は以下の通りです。

- **LLMモデルの推論性能向上**: 量子化・推論エンジンの改善・パフォーマンスチューニングなどを通じて、モデルの出力品質を担保しながら、ターゲット環境におけるスループットの向上やレイテンシ/メモリ使用量の削減を行う。

- **LLM推論エンジンの開発**: SaaS環境やオンプレミス環境でのサービングに必要となる推論エンジンと、その周辺機能の開発およびインテグレーションを行う。

- **LLMのカスタマイズ・デプロイの支援**: 案件ごとのニーズに応じてLLMモデルのカスタマイズを行うためのワークフローや、デプロイのパイプラインを整備する。

- **OSSコミュニティ活動**: vLLMをはじめとするLLM推論エンジンのエコシステムで、新機能の実装・バグフィックスのアップストリーム対応を行う。また、必要に応じて、関連イベントへの登壇や情報発信を行う。

入社後に実際に担当いただく業務内容は、専門・ご経験を考慮のうえ決定します。

***

Preferred Networks (PFN) is developing multimodal foundation models centered around Large Language Models (LLMs) and is expanding its research into applied development and industrial application research to strengthen its business domains in entertainment, scientific computing, robotics, and others.

The mission of the Inference Optimization Team is to serve as a "bridge" connecting cutting-edge models with real-world products. By fully realizing the potential of PFN's LLMs and optimizing them to meet the stringent quality, performance, and cost requirements of production environments, we support the expansion of PFN's product and solution business from the technical implementation perspective. Typical responsibilities include:

- **Improving Performance of LLM Inference**: Increasing throughput and reducing latency/memory usage in target environments while ensuring the model output quality, through techniques such as quantization, inference engine optimization, and performance tuning.

- **Developing LLM Inference Engines**: Create and integrate the necessary inference engines and supporting functions required for serving LLMs in SaaS environments and on-premise deployments.

- **Supporting LLM Customization and Deployment**: Establish workflows and deployment pipelines to enable customized LLM models according to each project's specific requirements.

- **Contributing to OSS Communities**: Participate in the development of the LLM inference engine ecosystem, particularly for vLLM, by implementing new features and upstream bug fixes. Additionally, when needed, present at relevant events and disseminate information.

The specific responsibilities you will actually handle after joining will be determined based on your expertise and experience.

Qualifications

- コンピュータサイエンスの知識や関心

- コンピューターサイエンスのすべての分野への精通を目指し、常に最先端の技術を追いかけ続けていること

- Pythonを使ったソフトウェア開発経験

- コンピューターアーキテクチャーを理解し、ソフトウェアの実効効率や、計算量を意識したプログラムの作成が出来ること

- 商用環境・オンプレミス環境など、品質要件の高いソフトウェアの開発経験

- LLM推論フレームワークの利用経験

- LLM推論最適化に関する各種技術に対する理解

- LLM推論に関連する最先端の技術動向を主体的にキャッチアップし、専門性を深める意欲

- CUDAアプリケーションの実装・性能改善・デバッグの経験

- チーム内外のメンバーと協働しながら課題解決ができること

- 既存フレームワークのコードや英語のドキュメントが抵抗なく読めること

***

- Knowledge and interest in Computer Science

- Aim to be familiar with all areas of computer science and continually pursue the latest technology

- Especially, research or practical experience and achievements in machine learning or natural language processing

- Experience in software development using Python

- Understand computer architecture and be able to create programs considering the actual efficiency of software and computational complexity

- Experience in developing software with high quality requirements in commercial environments or on-premises environments

- Experience working with LLM inference frameworks

- Understanding of optimization techniques for LLM inference

- Proactive ability to keep up with cutting-edge developments in LLM inference technology and deepen specialized knowledge

- Experience implementing, optimizing, and debugging CUDA applications

- Ability to collaborate effectively with team members both within and outside the organization

- Ability to understand existing framework code and English documentation

Preferred Qualifications

- LLM推論エンジンのコア技術(演算のスケジュール制御やKV cacheの管理機構など)の開発経験

- 商用環境や実際のユースケースにおけるLLM推論の性能測定・最適化の経験

- LLM推論に関連するOSSや、データサイエンス関連のOSSへのコントリビューション経験

- C++またはRustでのソフトウェア開発経験

- GPUクラスタ(Kubernetes等)の利用経験

- MLの学習・評価を行うワークフローフレームワーク(Argo, MLFlow等)の利用経験

- ビジネスレベルの英語・日本語コミュニケーション能力

***

- Development experience in core LLM inference engine technologies (e.g., scheduling control and KV cache management mechanisms)

- Experience in measuring and optimizing LLM inference performance in commercial environments and real-world use cases

- Experience contributing to open-source projects related to LLM inference or data science

- Software development experience in C++ and/or Rust

- Experience using GPU clusters (such as Kubernetes)

- Experience with workflow frameworks for ML training and evaluation (such as Argo and MLFlow)

- Business-level English and Japanese communication skills

Salary

経験、業績、能力、貢献に応じて、当社規定により優遇

Experience, performance, skills, contribution are taken into consideration.

Location

東京都千代田区大手町1-6-1 大手町ビル / Otemachi Bldg., 1-6-1 Otemachi, Chiyoda-ku, Tokyo, Japan 100-0004

Work style / 勤務形態

専門労働型裁量労働制(みなし労働時間:8時間)もしくはフレックス制

Discretionary-work (deemed work hours: 8 hours) or Flex-time system

ハイブリッド勤務(オフィス出社と在宅勤務を組み合わせての勤務)

Hybrid work (Combination of working from the office and working from home)

Salary increase & bonus / 昇給・賞与

年2回の人事評価及び会社業績に基づいて決定

Based on the result of a individual performance review (twice a year) and company’s performance

Allowances / 諸手当

通勤手当、在宅勤務手当

Commutation allowances / teleworking allowances

Holidays / 休日・休暇

休日:土曜日、日曜日、国民の祝日、国民の休日、年末年始

当社規定による年次有給休暇制度(入社時26日付与)

育児休暇、慶弔休暇など

Holiday: Saturdays and Sundays, public holidays, Year-end and new-year

Annual paid leave based on company regulations (26 days granted upon hire)

Parental leave, conguratulation / condolence leave etc.

Welfare / 福利厚生

社会保険完備(厚生年金保険、健康保険、雇用保険、労災保険)

確定拠出年金制度

ラップトップPC購入補助

定期健康診断実施

Various social insurance programs: pension insurance, health insurance, employment insurance, workers’ compensation

Defined contribution pension

Allowance for purchasing a laptop PC

Regular health checks

Employment Status / 雇用形態

正社員(試用期間3ヶ月、本採用と同条件)

Full-time regular employment (3 months of probation period under the same condition as regular employment)

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
743,556 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Tokyo
生成AIエンジニア 11 months ago
$38k – $95k per year • In office • Full-Time • Tokyo
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Databases
MySQL
PostgreSQL
Frontend
Tailwind CSS
Next.js
React.js
DevOps
Terraform
AWS
GitHub
Management
Slack
Apply
$374k – $575k per year • In office • Osaka
Python
TypeScript
SQL
Python
FastAPI
AI/ML
Cursor
Claude
ChatGPT
Gemini
LLM
RAG
OpenAI Codex
DevOps
Terraform
GCP
Docker
Google Cloud Run
Apply
$34k – $53k per year • In office • Full-Time • Tokyo
Python
Python
FastAPI
Databases
DynamoDB
AI/ML
Copilot
Claude
ChatGPT
Claude Code
Gemini
LLM
OpenAI Codex
DevOps
Git
AWS
AWS Lambda
GitHub
Management
Slack
Google Workspace
Discord
Apply
AI エンジニア 6 months ago
$64k – $114k per year • In office • Tokyo
JavaScript
Node JS
Node JS
Nest.JS
Databases
PostgreSQL
Amazon Aurora
AI/ML
Cursor
Claude Code
AWS Bedrock
LLM
Frontend
Next.js
React.js
DevOps
GitHub Actions
AWS CDK
CI/CD
AWS
Docker
AWS Fargate
Amazon ECS
Amazon EventBridge
QA
Playwright
Apply
AI Engineer 8 hours ago
Remote (India) • Full-Time • Pune
Python
Python
Flask
FastAPI
Databases
Supabase
AI/ML
LangGraph
AutoGen
LangChain
Claude
DSPy
LlamaIndex
LoRA
Model Context Protocol
vLLM
Fine-tuning
Prompt Engineering
Multimodal AI
Chain-of-Thought
AI Agents
OpenAI SDK
PEFT
Transformers
TensorFlow
PyTorch
CrewAI
Gemini
LLM
RAG
LLMOps
GPT-4
Agentic Workflows
Multi-Agent Systems
Vercel AI SDK
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Management
n8n
Apply
≈ $53k – $132k per year (Estimated) • Hybrid • Full-Time • Tokyo
Rust
C++
AI/ML
llama.cpp
vLLM
Quantization
SGLang
LLM
KV Cache
Apply
≈ $38k – $104k per year (Estimated) • Hybrid • Full-Time • Tokyo
Python
AI/ML
DeepSpeed
Multimodal AI
Computer Vision
VLM
LLM
Synthetic Data
SFT
Post-training
Pre-training
FSDP
arXiv
Machine Learning
DevOps
AWS
Kubernetes
Amazon EKS
Amazon EC2
Amazon S3
IAM
Amazon CloudWatch
Apply
≈ $52k – $131k per year (Estimated) • Hybrid • Full-Time • Tokyo
Python
AI/ML
DeepSpeed
Multimodal AI
AI Agents
TRL
Transformers
LLM
DPO
SFT
GRPO
Post-training
Pre-training
FSDP
Machine Learning
Apply
≈ $53k – $133k per year (Estimated) • Hybrid • Full-Time • Tokyo
Python
AI/ML
Reinforcement Learning
Multimodal AI
LLM
Post-training
Pre-training
Machine Learning
Robotics
Reinforcement Learning
Apply
$457k – $915k per year • In office • Contractor • Master's Degree • Tokyo
Python
AI/ML
LLM
Machine Learning
Apply
≈ $50k – $125k per year (Estimated) • Hybrid • Full-Time • Tokyo
Python
Ruby
AI/ML
LLM
OpenAI
DevOps
FinOps
Apply
≈ $53k – $132k per year (Estimated) • Hybrid • Full-Time • Tokyo
Rust
C++
AI/ML
llama.cpp
vLLM
Quantization
SGLang
LLM
KV Cache
Apply
≈ $49k – $124k per year (Estimated) • Hybrid • Full-Time • Tokyo
C#
C#
.NET
AI/ML
LLM
DevOps
Kubernetes
eBPF
HPC
Linux
BGP
Apply
≈ $49k – $122k per year (Estimated) • Hybrid • Full-Time • Tokyo
Python
TypeScript
C#
C#
.NET
AI/ML
LLM
DevOps
CI/CD
AWS
Kubernetes
Linux
Apply
新卒採用 4 hours ago
In office • Full-Time • Tokyo
Apply
In office • Full-Time • Tokyo
Apply
In office • Full-Time • Tokyo
Java
PHP
C++
Apply
In office • Full-Time • Tokyo
Game Dev
Unreal Engine
Houdini
Substance 3D Designer
Design
ZBrush
Substance 3D Painter
Apply
≈ $35k – $94k per year (Estimated) • In office • Full-Time • Tokyo
Apply
See all jobs
This is one of many
743,556 more open roles from verified company boards, updated every day.