About the Role
This is a senior individual contributor and team lead role sitting at the heart of an AI/ML platform company focused on reinforcement learning environments and post-training data for frontier agents. You will own the strategy and systems that measure, improve, and scale training data quality, directly shaping what makes agent data rigorous and useful in production.
What You'll Do
Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
Define data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs.
Develop new methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
Turn qualitative research insights into production systems: internal tools, dashboards, validation pipelines, and feedback loops.
Build internal research intuition around what makes agent training data realistic, learnable, diverse, reliable, and genuinely useful.
Mentor other research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
What We're Looking For
5+ years of experience in research or data quality engineering, specifically building systems for AI/ML data evaluation.
Demonstrated track record leading technical projects or teams on ambiguous problems, from definition through implementation and iteration.
Advanced proficiency in Python, Docker, and Linux environments.
Background as an ML researcher or engineer specializing in quality and evaluation of AI training data at scale, not traditional data science or analytics.
Deep intuition for what makes training tasks high quality: realistic, learnable, diverse, reliable, and useful for agent training.
Experience designing metrics, experiments, and QA/QC processes, not just executing them.
Experience translating research insights into production systems and data pipelines.
Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems.
Strong written communication skills, with the ability to explain methodology clearly to researchers, engineers, and external audiences.
Research-oriented understanding of AI evals and post-training, beyond surface-level agent tooling.
Comfort navigating complex systems involving domain experts, vendors, model outputs, graders, and infrastructure.
Early-stage startup experience and the ability to work independently at pace.
Compensation & Benefits
Salary: USD 150,000 to 250,000 annually
Visa sponsorship: Available
Location
On-site in Singapore. Candidates must be available to work from the Singapore office.

