Salary
≈ $57k – $122k per year (Estimated)
Location
Remote (Australia, Taiwan, Hong Kong)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Binance is the world's largest cryptocurrency exchange by trading volume, founded in 2017 by Changpeng Zhao and Yi He. The platform offers spot, margin and derivatives trading across hundreds of digital assets, alongside staking, savings products, payments, an institutional custody arm and a self-custodial Web3 wallet. The group also created BNB Chain, one of the most used smart contract networks, and now operates under a licensed regional structure after a 2023 settlement with United States authorities that installed new leadership and compliance oversight.
About the Role
In the AI era, large language models are reshaping core business scenarios such as dialogue and trading. Model capability iteration relies on a scientific and trustworthy evaluation system - the "ruler" that measures model quality and guides R&D direction. We are seeking an evaluation expert with an algorithmic background to build LLM evaluation capabilities covering dialogue, financial trading, and other scenarios, using professional evaluation methods to quantify model performance, pinpoint issues, and drive continuous model improvement.
Responsibilities
- Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.
- Lead the design and construction of evaluation datasets. Define evaluation dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.
- Analyze model capability boundaries and failure modes based on evaluation results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.
- Drive the automation and scaling of evaluation workflows. Build sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.
- Collaborate with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.
Requirements
- Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.
- Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.
- Familiar with mainstream evaluation methods (human evaluation, model-based automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.
- Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.
- Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.
- Strong business understanding and communication skills, able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration.
Bonus Qualifications
- Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs.
- Experience building high-quality AI training/evaluation data or data annotation systems.
- Familiarity with RLHF, reward models, or preference data-related work.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
666,674 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Hong Kong
≈ $59k – $145k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Denmark
Python
AI/ML
Copilot
AI Agents
RAG
Human-in-the-Loop
LLM Guardrails
Apply
Senior Engineer - Analog Simulation CAD
13 min ago
In office • Full-Time • Hyderabad
Python
Perl
AI/ML
Copilot
AI Agents
Chips/EDA
Cadence Virtuoso
Cadence Xcelium
Synopsys StarRC
Cadence Quantus
Synopsys PrimeSim
Synopsys FineSim
Synopsys Custom Compiler
Apply
Senior Data Scientist
3 min ago
≈ $36k – $88k per year (Estimated) • Remote/Hybrid • Full-Time • Athens
Python
SQL
Python
FastAPI
Databases
Databricks
AI/ML
CatBoost
MLFlow
XGBoost
Reinforcement Learning
Scikit-learn
LightGBM
PyTorch
DevOps
Azure
AWS
Docker
Apply
$83k – $149k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Boston • San Francisco • Toronto
Python
SQL
AI/ML
Physical AI
Analytics
Tableau
Power BI
Apply
(Senior) Data Engineer
6 min ago
≈ $32k – $78k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Belgrade • Prague • Brno
Python
Java
PowerShell
Scala
Databases
PostgreSQL
Redis
Snowflake
Databricks
DynamoDB
MS SQL
ElasticSearch
Apache Kafka
Google BigQuery
Amazon Aurora
BigQuery
AI/ML
dbt
DevOps
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
AWS
AWS Lambda
Amazon Kinesis
AWS Step Functions
Analytics
Tableau
Power BI
ETL/ELT
AWS Glue
Looker
Management
Confluence
Jira
Agile
Apply
≈ $43k – $116k per year (Estimated) • Remote • Full-Time • Master's Degree • Hong Kong
Python
AI/ML
Reinforcement Learning
Time Series Forecasting
Web3
DeFi
Apply
Strategy & Operations Lead - Africa
1 day ago
≈ $20k – $52k per year (Estimated) • Remote • Full-Time • 8+ years exp • Cape Town
Apply
Compliance Analyst
1 day ago
≈ $14k – $37k per year (Estimated) • In office • Full-Time • 3+ years exp • Astana
Web3
Chainalysis
Apply
Regional Branding Manager - GC
2 days ago
≈ $44k – $111k per year (Estimated) • Remote • 5+ years exp • Taipei
Apply
Legal Admin Executive - 12 months contract
2 days ago
≈ $84k – $217k per year (Estimated) • Remote • Contractor • 5+ years exp
Apply
Project Management, Equities - VP
1 day ago
≈ $47k – $98k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Hong Kong
Management
SharePoint
Agile
Apply
Client Support Analyst (APAC)
1 day ago
≈ $21k – $50k per year (Estimated) • Remote/Hybrid • 1+ year exp • Hong Kong
DevOps
SLI/SLO/SLA
Apply
Customer Success Manager - APAC
1 day ago
≈ $23k – $54k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Hong Kong
Marketing
Salesforce
HubSpot
Apply
Apply
Marketing and Communications Manager
1 day ago
≈ $57k – $127k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Hong Kong
Marketing
X (Twitter)
LinkedIn
Instagram
Apply
This is one of many
666,674 more open roles from verified company boards, updated every day.

