{"id":1177407,"url":"https://alion.io/job/referred-etworks-lm-erving-ngine-ngineer-lmsaabinguenjinenjinia","title":"LLM Serving Engine Engineer / LLMサービングエンジンエンジニア","company":{"id":130,"name":"Preferred Networks","domain":"preferred.jp","url":"https://alion.io/company/preferred-networks","size_band":"201-500","is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Talentio","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Tokyo, Japan"],"countries":["JP"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":53000,"max_usd":132000,"period":"year","method":"role_country_seniority_unknown","sample_n":23},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"KV Cache","optional":false},{"name":"llama.cpp","optional":false},{"name":"LLM","optional":false},{"name":"Quantization","optional":false},{"name":"SGLang","optional":false},{"name":"vLLM","optional":false},{"name":"C++","optional":true},{"name":"Rust","optional":true}],"status":"live","first_seen_at":"2026-09-24T12:06:17Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-24T16:10:48Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Job Description / 職務内容\nWe are looking for an engineer to join the development of the LLM inference serving engine for the MN-Core L series.\nAt PFN, we are developing and utilizing the MN-Core™ architecture along with its recent generations MN-Core L series. The MN-Core architecture employs compiler-based pre-scheduling to optimize the majority of processor control, meaning the compiler's quality directly determines the performance of the accelerator. The L series leverages novel 3D stacked DRAM technology to achieve high bandwidth required for LLM inference. The serving engine aims to enable ultra-low latency LLM inference by tightly integrating the compiler-generated programs, runtime environment, and data processing on DRAM.\nIn this position, you will be involved in developing this serving engine. Your specific software development responsibilities will include:\n- Implementing execution control and request scheduling for LLM inference processing. \n- Optimizing KV cache management to support useful features like prefix caching while fully utilizing the high DRAM bandwidth. \n- Working closely with compiler, runtime, and hardware developers. \n- Designing benchmarks and simulations to ascertain the LLM inference performance on specific workloads and use cases. \n- Conducting surveys of existing inference engines like vLLM and SGLang.\nIn addition to these core responsibilities, depending on your interests and aptitudes, you may also participate in broader projects related to high-performance computing system design and utilization, compression such as quantization, testing with downstream tasks, and architectural considerations for the MN-Core itself.\nMN-Core Lシリーズ向けLLMサービングエンジンの開発チームに加わるエンジニアを募集しています。\nPFNでは、最新世代のMN-Core™アーキテクチャおよびMN-Core Lシリーズの開発・活用を進めています。MN-Coreアーキテクチャはコンパイラによる事前スケジューリングを採用しており、プロセッサ制御の大部分を最適化しています。このため、コンパイラの品質がアクセラレータの性能に直接影響します。Lシリーズでは、LLM推論に必要な高帯域幅を実現するため、3D積層DRAM技術を採用しています。本LLMサービングエンジンは、コンパイラが生成したプログラム、ランタイム環境、およびDRAM上でのデータ処理を緊密に協調させることで、超低遅延なLLM推論を実現することを目的としています。\n本ポジションでは、このサービスエンジンの開発に携わっていただきます。具体的なソフトウェア開発業務としては以下が含まれます：\n* LLM推論処理における実行制御とリクエストスケジューリングの実装 \n* 高帯域幅DRAMを最大限に活用しつつ、プレフィックスキャッシュなどの有用な機能をサポートするためのKVキャッシュ管理の最適化 \n* コンパイラ・ランタイム・ハードウェア開発者との緊密な連携 \n* 特定のワークロードやユースケースにおけるLLM推論性能を評価するためのベンチマークおよびシミュレーションの設計 \n* vLLMやSGLangなどの既存推論エンジンに関する調査\nこれらの主要業務に加え、ご自身の興味や適性に応じて、高性能コンピューティングシステムの設計・活用、量子化などの圧縮技術、下流タスクとの連携テスト、MN-Coreアーキテクチャ自体の設計検討など、より広範なプロジェクトにも関与していただくことも可能です。\n**What makes this position appealing**\nAt PFN, you'll have the opportunity to work on the development of the serving engine for cutting-edge LLM inference accelerators. You'll work in an environment where you can collaborate closely with users as well as hardware developers, making this an ideal opportunity for those passionate about applying world-class computing technology in real-world applications.\nPFNでは、最先端のLLM推論アクセラレータ向けソフトウェアの開発に携わることができます。ハードウェア開発者と密接に連携しながら開発を進めることができる環境であり、世界最先端のコンピューティング技術を実社会で活用することに情熱を持つ方にとって、理想的な環境です。\n**Portrait of a person**\n- Individuals with broad interests and a desire to acquire knowledge in new technical domains \n- Colleagues who can respect and work well together, regardless of job role or background \n- Those who can leverage their strengths and support team members \n- Individuals who can approach problem-solving as their own responsibility, regardless of ownership \n- People who can absorb new knowledge and enjoy working in environments with diverse expertise\n- 様々な分野への関心、新たな技術領域の知見獲得の意欲のある方 \n- 同職種・他職種に関わらずリスペクトして一緒に楽しく働ける方 \n- 強みを活かして、チームメンバと助け合える方 \n- 周りの課題に対しても自分事として捉え課題解決を推進できる方 \n- 様々な専門性を持つ人がいる環境で新しいことを吸収し、楽しめる方\n**Reference Links**\n[About the inference chip MN-Core L1000](https://mn-core.com/)\n[推論チップ MN-Core L1000について](https://mn-core.com/ja)\nQualifications / 応募資格（必須）\n- Basic understanding of computer science at university undergraduate level \n- Interest in emerging LLM inference accelerators. \n- Theoretical and practical familiarity with LLM inference. \n- Understand how hardware specifications influence performance. \n- Understanding of different workload scenarios and how they relate to performance. \n- Understanding of the different tradeoffs involved in inference optimization. \n- Familiar with open source LLM inference engines (such as vLLM, SGLang, dynamo, llama.cpp). \n- Openness to work in an English-Japanese-mixed environment\n- 大学学部レベルのコンピュータサイエンスについての基礎的な理解\n- 新興のLLM推論アクセラレータ技術に対する関心 \n- LLM推論に関する理論的・実践的な知識を有していること \n- ハードウェア仕様が性能に与える影響を理解できること \n- 様々なワークロードシナリオとそれらが性能に及ぼす影響についての理解\n- 推論最適化に伴う各種トレードオフについての理解\n- オープンソースのLLM推論エンジン（vLLM、SGLang、dynamo、llama.cppなど）に精通していること\n- 英語と日本語が混在する環境での業務に柔軟に対応できること\nPreferred Qualifications / 応募資格（歓迎）\n- Contributions to open source LLM inference engines (vLLM, SGLang, llama.cpp, dynamo, …). \n- Experience with LLM serving (either running locally or at scale). \n- Familiar with core parts of LLM serving (such as request scheduling, KV cache management or prefix caching). \n- Knowledge about LLM serving APIs (/chat/completion, /responses, /messages, …). \n- Interest in LLM applications (such as coding agents). \n- Knowledge about ascertaining LLM inference performance through benchmarks and simulations. \n- System engineering experience (understanding of performance optimization, profiling, memory management, storage, concurrency, etc.). \n- Experience with system programming languages (C/C++, Rust, Go, …). \n- Experience with emerging LLM inference accelerators. \n- Ability to read technical documentation/discussions in Japanese.\n- OSSのLLM推論エンジン（vLLM、SGLang、llama.cpp、dynamoなど）への貢献経験 \n- LLMのサービス提供に関する実務経験（ローカル環境および大規模環境での運用双方） \n- LLMサービス提供の主要コンポーネントに関する知識（リクエストスケジューリング、KVキャッシュ管理、プレフィックスキャッシュ処理など） \n- LLMサービス提供APIに関する知見（/chat/completion、/responses、/messagesなど）\n- LLM応用技術への関心（コーディング支援エージェントなどの事例） \n- ベンチマークテストやシミュレーションを通じたLLM推論性能評価に関する知識 \n- システムエンジニアリングの実務経験（パフォーマンス最適化、プロファイリング、メモリ管理、ストレージ、並行処理などの理解）\n- システムプログラミング言語の使用経験（C/C++、Rust、Goなど）\n- 最新のLLM推論アクセラレータ技術に関する知見 \n- 日本語の技術文書/ディスカッションを読み解く能力\nSalary /賃金\n経験、業績、能力、貢献に応じて、当社規定により優遇\nExperience, performance, skills, contribution are taken into consideration.\nLocation / 勤務地\n東京都千代田区大手町１-６-１ 大手町ビル / Otemachi Bldg., 1-6-1 Otemachi, Chiyoda-ku, Tokyo, Japan 100-0004\nリモート勤務制度あり （日本国内に限る) / Remote work system available (limited to work in Japan)\nWork style / 勤務形態\n専門労働型裁量労働制（みなし労働時間：8時間）もしくはフレックス制\nDiscretionary-work (deemed work hours: 8 hours) or Flex-time system\nハイブリッド勤務（オフィス出社と在宅勤務を組み合わせての勤務）\nHybrid work (Combination of working from the office and working from home)\nSalary increase & bonus / 昇給・賞与\n年2回の人事評価及び会社業績に基づいて決定\nBased on the result of a individual performance review (twice a year) and company’s performance\nAllowances / 諸手当\n通勤手当、在宅勤務手当\nCommutation allowances / teleworking allowances\nHolidays / 休日・休暇\n休日：土曜日、日曜日、国民の祝日、国民の休日、年末年始\n当社規定による年次有給休暇制度（入社時26日付与）\n育児休暇、慶弔休暇など\nHoliday: Saturdays and Sundays, public holidays, Year-end and new-year\nAnnual paid leave based on company regulations (26 days granted upon hire)\nParental leave, conguratulation / condolence leave etc.\nWelfare / 福利厚生\n社会保険完備（厚生年金保険、健康保険、雇用保険、労災保険）\n確定拠出年金制度\nラップトップPC購入補助\n定期健康診断実施\nVarious social insurance programs: pension insurance, health insurance, employment insurance, workers’ compensation\nDefined contribution pension\nAllowance for purchasing a laptop PC\nRegular health checks\nEmployment Status / 雇用形態\n正社員（試用期間3ヶ月、本採用と同条件）\nFull-time regular employment (3 months of probation period under the same condition as regular employment)","description_format":"text","description_chars":7415,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"Japanese","level":"Advanced (C1)","optional":false}]},"benefits":["Health insurance","Hybrid work","Parental leave"],"hiring_locations":[{"name":"Japan","iso":"JP","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Robotics AI","Foundation Models","AI Chips & Accelerators","AI for Science"],"lifecycle":[{"event":"open","at":"2026-09-24T12:06:17Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":33,"reasons":["conf:3","win:early"],"computed_at":"2026-09-24T20:09:08Z"},"pay":null,"html_url":"https://alion.io/job/referred-etworks-lm-erving-ngine-ngineer-lmsaabinguenjinenjinia","json_url":"https://alion.io/job/referred-etworks-lm-erving-ngine-ngineer-lmsaabinguenjinenjinia.json","meta":{"generated_at":"2026-09-24T20:09:08Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}