823,995open jobs
52,997companies
135,021added this week
Browse all
Location
In office (Singapore)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Aug 26, 2025.

Overview
Company
Impact
Profile match
TikTok is a short-form video platform owned by the Chinese technology group ByteDance and launched internationally in 2016, growing further after the 2018 merger with the lip-sync app Musical.ly. Its recommendation feed surfaces content by engagement signals rather than social graph, which lets new creators reach large audiences quickly and made the app one of the most downloaded in the world. The company runs dual operating hubs in Singapore and Los Angeles and has built an advertising business, a live streaming economy and the TikTok Shop commerce marketplace on top of the video feed.

岗位职责 / Responsibilities

Team Introduction

The mission of our AML team is to push next-generation machine learning algorithms and platforms for the recommendation system, ads ranking and search ranking in our company. We also drive substantial impact on core businesses of the company.

Responsibilities:

1. Resource Efficiency Optimization in Distributed Orchestration and Scheduling:

- Develop and extend distributed orchestration frameworks within the Kubernetes/Godel ecosystem. Select appropriate frameworks based on different business scenarios, and optimize cluster utilization and load balancing strategies according to the specific characteristics of each scenario;

- Integrate and expand AutoScaling and automatic parallelization capabilities for various models and tasks. Employ load modeling and analytic methods for different models to automatically optimize resource requests, achieving large-scale improvements in resource usage efficiency and global optimality;

- Responsible for preemption and re-scheduling mechanisms for services with different prioritties, and manage automatic resource multiplexing across different clusters and resource types; handle scheduling and load adaptation across multi-datacenter, multi-region, and multi-cloud environments.

2. Building Training System Architecture for Next-Generation Ultra-Large and Ultra-Deep Recommendation Models:

- Develop a flexible, elastic and robust distributed training runtime focused on hyper-scaled embeddings and large-scale GPU training;

- Design and optimize distributed computing APIs and runtimes geared towards future recommendation and ads model paradigms (e.g., reinforcement learning, fine-tuning and/or distillation);

- Collaborate with platform teams to enhance the diagnosability and usability of distributed training systems.

3. Constructing Online Orchestration Architecture for Next-Generation Recommendation Systems:

- Build a robust distributed model inference architecture for online learning scenarios involving hyper-scaled embeddings;

- Optimize the usability of online recommendation and ads model architectures and MLops workflows.

任职要求 / Requirements

Minimum Qualifications

- Bachelor's degree or above, majoring in Computer Science, Engineering or related fields.

- Strong programming and coding experience with at least one modern language such as Golang, Python.

- Experience contributing to the large scale distributed systems, multi-tenant systems (architecture, reliability and scaling).

- Strong analytical abilities and problem solving.

- Good communication, self-motivation, engineering practice, documentation, etc.

- At least 3 years of relevant experience.

Preferred Qualifications

- Familiar with large-scale distributed scheduling systems like Kubernetes, Yarn, Flink and/or Spark

- Familiar with opensourced orchestration frameworks like VeRL, vLLM, Ray or TFX, etc.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
823,995 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Singapore
≈ $77k – $203k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Singapore
AI/ML
AI Agents
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Singapore
Java
C++
AI/ML
Machine Learning
Apply
In office • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Java
C++
Scala
Databases
HBase
Delta Lake
AI/ML
Flink
Recommender Systems
Apply
≈ $74k – $197k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Singapore
Python
Go
Java
C++
DevOps
Linux
Apply
≈ $72k – $190k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Singapore
Python
Apply
$100k – $180k per year • Remote (United States) • 10+ years exp • Bachelor's Degree
Python
Go
JavaScript
Java
Node JS
Bash
Node JS
Commander.js
DevOps
GCP
Istio
OpenTelemetry
Consul
Datadog
Linkerd
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Chaos Engineering
Service Mesh
Linux
Apply
$160k – $180k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Hadoop
Spark
MLFlow
Reinforcement Learning
Scikit-learn
Computer Vision
Kubeflow
TensorFlow
Pandas
NumPy
PyTorch
Explainable AI
Feature Store
Recommender Systems
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Management
Agile
Apply
$80k – $100k per year • Remote (United States) • 6+ years exp • Master's Degree
Python
AI/ML
RLHF
Reinforcement Learning
AI Agents
Reward Modeling
Machine Learning
Robotics
Reinforcement Learning
Apply
≈ $32k – $86k per year (Estimated) • In office
Python
Java
C#
C++
C#
.NET
Apply
≈ $28k – $60k per year (Estimated) • In office • Full-Time • 6+ years exp • India
Python
JavaScript
TypeScript
C#
Node JS
C#
.NET
AI/ML
Anomaly Detection
Frontend
React.js
DevOps
Splunk
GCP
New Relic
OpenTelemetry
Datadog
Dynatrace
Azure
CI/CD
AWS
Grafana
Self-Healing
AppDynamics
AIOps
Incident Management
Apply
In office • Internship • Singapore
Python
Go
Java
C++
AI/ML
Machine Learning
Apply
In office • Full-Time • Bachelor's Degree • Singapore
Python
Java
SQL
C++
Databases
ClickHouse
Presto
Apache Kafka
Trino
AI/ML
Spark
Model Context Protocol
dbt
Function Calling
Flink
RAG
LLM Evaluation
Recommender Systems
Analytics
Dimensional Modeling
Apply
In office • Full-Time • 2+ years exp • Bachelor's Degree • Sydney
Python
Go
C++
AI/ML
vLLM
CUDA Toolkit
Quantization
Multimodal AI
SGLang
LLM
CUDA
CUTLASS
Apply
≈ $140k – $259k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Jose
Apply
In office • Internship • Bachelor's Degree • Sydney
Python
PHP
Apply
≈ $58k – $125k per year (Estimated) • In office • Full-Time • 3+ years exp • Singapore
DevOps
PagerDuty
Incident Management
Management
ServiceNow
Apply
≈ $86k – $214k per year (Estimated) • In office • Full-Time • Master's Degree • Singapore
Python
Java
SQL
C++
AI/ML
Hadoop
Spark
AI Agents
Flink
Machine Learning
Apply
≈ $65k – $159k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Singapore
Apply
≈ $86k – $232k per year (Estimated) • In office • Internship • PhD • Singapore
Python
C++
AI/ML
Multimodal AI
Machine Learning
Apply
≈ $88k – $237k per year (Estimated) • In office • Internship • PhD • Singapore
AI/ML
Reinforcement Learning
Multimodal AI
Computer Vision
NLP
LLM
Pre-training
Recommender Systems
Machine Learning
Apply
See all jobs
This is one of many
823,995 more open roles from verified company boards, updated every day.