About the Role
This is a senior individual-contributor-and-lead role at an early-stage AI infrastructure startup building the tooling that powers reinforcement learning environments and post-training data for frontier AI models. You will own the strategy, systems, and culture around data quality, shaping what "good" training data looks like at scale and making that judgment reproducible across the team.
What You'll Do
Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
Define and enforce data quality strategy by building QC systems, setting standards, and designing experiments to grade agent outputs.
Develop new methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
Turn qualitative research insights into production systems: internal tools, dashboards, validation pipelines, and feedback loops.
Build internal research intuition around what makes agent training data realistic, learnable, diverse, reliable, and useful.
Mentor other research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
What We're Looking For
5+ years of engineering or research experience, specifically building systems for AI/ML data evaluation or data quality.
Demonstrated track record leading technical projects or teams on ambiguous problems, from definition through implementation and iteration.
Advanced proficiency in Python, Docker, and Linux environments.
Background as an ML researcher or engineer focused on quality and evaluation of AI training data at scale, not traditional data science or analytics.
Deep intuition for what makes AI agent training data high-quality: realistic, learnable, diverse, and reliably scored.
Experience designing metrics, experiments, and QA/QC processes, not just executing them.
Experience collaborating with subject-matter experts, vendors, and domain specialists to capture judgment and convert it into scalable review or generation systems.
Research-oriented understanding of AI evals and post-training pipelines.
Strong written communication skills, with the ability to explain methodology clearly to mixed audiences.
Early-stage startup experience and comfort working independently in fast-paced environments.
Detail-oriented mindset with a strong eye for subtle inconsistencies or edge cases in data.
Compensation & Benefits
Salary range: $150,000 to $180,000 USD annually. Visa sponsorship is available.
Location
On-site in Singapore.

