814,841open jobs
52,433companies
131,214added this week
Browse all
Salary
≈ $129k – $281k per year (Estimated)
Location
In office (Singapore)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on May 27, 2025.

Overview
Company
Impact
Profile match
ByteDance is a Chinese internet technology company founded in Beijing in 2012 that operates some of the world's largest content and commerce platforms. Its portfolio includes the short-video apps TikTok and Douyin, the news aggregator Toutiao, the video editor CapCut, the workplace suite Lark and the Doubao family of AI assistants and models. Recommendation systems trained on user behaviour sit at the centre of every product, and the company has grown into one of the highest-revenue private technology businesses in the world.

岗位职责 / Responsibilities

Content Security Algorithm Research Team:

The International Content Safety Algorithm Research Team is dedicated to maintaining a safe and trustworthy environment for users of ByteDance's international products. We develop and iterate on machine learning models and information systems to identify risks earlier, respond to incidents faster, and monitor potential threats more effectively. The team also leads the development of foundational large models for products. In the R&D process, we tackle key challenges such as data compliance, model reasoning capability, and multilingual performance optimization. Our goal is to build secure, compliant, and high-performance models that empower various business scenarios across the platform, including content moderation, search, and recommendation.

Research Project Background:

In recent years, Large Language Models (LLMs) have achieved remarkable progress across various domains of natural language processing (NLP) and artificial intelligence. These models have demonstrated impressive capabilities in tasks such as language generation, question answering, and text translation. However, reasoning remains a key area for further improvement. Current approaches to enhancing reasoning abilities often rely on large amounts of Supervised Fine-Tuning (SFT) data. However, acquiring such high-quality SFT data is expensive and poses a significant barrier to scalable model development and deployment.

To address this, OpenAI's o1 series of models have made progress by increasing the length of the Chain-of-Thought (CoT) reasoning process. While this technique has proven effective, how to efficiently scale this approach in practical testing remains an open question. Recent research has explored alternative methods such as Process-based Reward Model (PRM), Reinforcement Learning (RL), and Monte Carlo Tree Search (MCTS) to improve reasoning. However, these approaches still fall short of the general reasoning performance achieved by OpenAI's o1 series of models. Notably, the recent DeepSeek R1 paper suggests that pure RL methods can enable LLM to autonomously develop reasoning skills without relying on the expensive SFT data, revealing the substantial potential of RL in advancing LLM capabilities.

Project Challenges:

1. Design of Reward Models: In the RL process, designing an effective reward model is crucial. It must accurately reflect the effectiveness of the reasoning process and guide the model to iteratively improve its reasoning ability. This involves not only setting appropriate evaluation criteria across different tasks, but also ensuring the reward model to adapt dynamically during training to match the evolving model performance.

2. Stability of the Training Process: In the absence of high-quality SFT data, ensuring stable training in RL becomes a major challenge. RL often involves extensive exploration and trial-and-error, which may lead to unstable training or even performance degradation. Developing robust training strategies is essential to ensure the reliability and effectiveness of the training process for models.

3. Expanding from Mathematics and Code Tasks to Natural Language Tasks: Current RL reasoning methods are primarily applied to mathematics and code tasks, where CoT data is more abundant. However, natural language tasks are more open and complex. Expanding from successful RL strategies to natural language processing tasks requires in-depth research and innovation in both data design and RL methodology to enable cross-task general reasoning capabilities.

4. Improving Reasoning Efficiency: While maintaining high reasoning quality, improving reasoning efficiency is another critical challenge. Efficient reasoning directly impacts the model's practicality and cost-effectiveness in real-world applications. Approaches such as knowledge distillation (transferring knowledge from complex models to smaller models) can be explored to reduce computational resource consumption, or the use of Long Chain-of-Thought (Long-CoT) techniques to improve Short-CoT models to balance reasoning accuracy with computational efficiency.

任职要求 / Requirements

1. Got PhD degree in Computer Science, Electronics, or other related fields.

2. Extensive experience in ML/CV/NLP/Recommendation Systems, including but not limited to:

a. Participation in competitions or industry projects in ML, Data Mining, CV, NLP, or Multimodal.

b. Publications in conferences in ML, data mining, AI, or large models (e.g., KDD, WWW, NIPS, ICML, CVPR, ACL, AAAI etc).

Bonus points:

1) Research experience or innovation in large models or RL.

2) Strong hands-on skills with contributions to large model projects in the open-source community.

3) Practical experience in deploying large models in real-world business scenarios.

4. Strong programming skills and proficient in Python/C++ or other relevant programming languages.

5. Outstanding problem-solving and analytical skills, with a passion for tackling challenging problems.

6. Strong enthusiasm for technology, with excellent communication skills and collaborative mindset.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
814,841 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Singapore
≈ $86k – $214k per year (Estimated) • In office • Full-Time • Master's Degree • Singapore
Python
Java
SQL
C++
AI/ML
Hadoop
Spark
AI Agents
Flink
Machine Learning
Apply
≈ $86k – $232k per year (Estimated) • In office • Internship • PhD • Singapore
Python
C++
AI/ML
Multimodal AI
Machine Learning
Apply
≈ $88k – $237k per year (Estimated) • In office • Internship • PhD • Singapore
AI/ML
Reinforcement Learning
Multimodal AI
Computer Vision
NLP
LLM
Pre-training
Recommender Systems
Machine Learning
Apply
≈ $91k – $245k per year (Estimated) • In office • Internship • PhD • Singapore
AI/ML
Reinforcement Learning
Multimodal AI
NLP
SGLang
Tokenization
Time Series Forecasting
Post-training
Knowledge Graph
Recommender Systems
World Models
Machine Learning
Apply
≈ $86k – $231k per year (Estimated) • In office • Internship • PhD • Singapore
Python
C++
AI/ML
Multimodal AI
RAG
Apply
≈ $76k – $143k per year (Estimated) • In office • Full-Time • 8+ years exp • PhD • Boston
AI/ML
Multimodal AI
Human-in-the-Loop
Physical AI
Machine Learning
Robotics
Digital Twin
Apply
IT Project Lead 1 day ago
$87k – $198k per year • In office • Top Secret • Full-Time • 12+ years exp • Bachelor's Degree • Chantilly
AI/ML
Multimodal AI
Management
Agile
Apply
≈ $104k – $207k per year (Estimated) • In office • 7+ years exp • Sunnyvale
AI/ML
Multimodal AI
Apply
$55k – $60k per year • In office • Full-Time • PhD • Vancouver
Python
AI/ML
Multimodal AI
Machine Learning
Apply
$179k – $205k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • McLean • New York • Plano
Python
SQL
Scala
AI/ML
Reinforcement Learning
Transformers
TensorFlow
PyTorch
Time Series Forecasting
Sentiment Analysis
Recommender Systems
Machine Learning
DevOps
AWS
Apply
≈ $92k – $247k per year (Estimated) • In office • Internship • Bachelor's Degree • Singapore
AI/ML
Multimodal AI
NLP
Post-training
Machine Learning
Apply
≈ $85k – $229k per year (Estimated) • In office • Internship • Bachelor's Degree • Singapore
Python
Java
Rust
C++
AI/ML
Model Context Protocol
vLLM
RLHF
AI Agents
NLP
TensorRT
TensorRT-LLM
LLM
RAG
Mixture of Experts
DPO
SFT
Post-training
Context Engineering
Prompt Caching
Tool Use
RLAIF
Reward Modeling
Apply
≈ $87k – $234k per year (Estimated) • In office • Internship • PhD • Singapore
AI/ML
Stable Diffusion
Multimodal AI
Function Calling
NLP
VLM
ControlNet
TensorRT
TensorFlow
PyTorch
LLM
RAG
Anomaly Detection
Tool Use
Apply
≈ $87k – $232k per year (Estimated) • In office • Internship • PhD • Singapore
Python
Java
AI/ML
Reinforcement Learning
AI Agents
NLP
TensorFlow
PyTorch
LLM
Time Series Forecasting
SFT
Pre-training
Knowledge Graph
Recommender Systems
Machine Learning
Apply
≈ $128k – $279k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Singapore
Python
C++
C++
TensorFlow C++
PyTorch C++
Databases
Redis
RocksDB
AI/ML
TensorFlow
PyTorch
KV Cache
Machine Learning
DevOps
AIOps
Linux
Apply
≈ $86k – $214k per year (Estimated) • In office • Full-Time • Master's Degree • Singapore
Python
Java
SQL
C++
AI/ML
Hadoop
Spark
AI Agents
Flink
Machine Learning
Apply
≈ $86k – $232k per year (Estimated) • In office • Internship • PhD • Singapore
Python
C++
AI/ML
Multimodal AI
Machine Learning
Apply
≈ $88k – $237k per year (Estimated) • In office • Internship • PhD • Singapore
AI/ML
Reinforcement Learning
Multimodal AI
Computer Vision
NLP
LLM
Pre-training
Recommender Systems
Machine Learning
Apply
≈ $91k – $245k per year (Estimated) • In office • Internship • PhD • Singapore
AI/ML
Reinforcement Learning
Multimodal AI
NLP
SGLang
Tokenization
Time Series Forecasting
Post-training
Knowledge Graph
Recommender Systems
World Models
Machine Learning
Apply
≈ $86k – $231k per year (Estimated) • In office • Internship • PhD • Singapore
Python
C++
AI/ML
Multimodal AI
RAG
Apply
See all jobs
This is one of many
814,841 more open roles from verified company boards, updated every day.