Anyone AI Labs - Anyone AI’s Human Data Division Owns: the Human Data platform (RL / coding / STEM evaluation environments, pipelines & project taxonomies, and the admin/ops surface like payments, task approve/reject, roles, revenue & cost)
Location: Remote / LatAm / US
The role Development of the platform that runs Anyone AI’s Human Data work end to end. That means standing up and evolving environments for RL, coding, STEM, and related evals; leading pipeline and project-taxonomy design; and building the admin systems that keep production moving, payments, task approval and rejection, roles, and visibility into project revenue and cost. This is a high-autonomy seat: you lead the platform initiative. You need enough fluency in human data / evaluation workflows to make the right product and engineering calls without constant hand-holding.
Responsibilities
- Evaluation environments. Design, stand up, and harden RL, coding, STEM, and other eval environments that contributors and internal teams can actually run against (reliably, repeatably, and at the quality bar labs expect).
- Pipelines & taxonomies. Lead platform development for how projects are structured: pipelines from brief → tasks → QC → delivery, plus project taxonomies that stay coherent as we add domains and clients.
- Admin & operations surface. support the operational layer: payments / payouts, task approve/reject flows, RBAC and roles management, and practical revenue & cost visibility per project.
- Lead the initiative. Set technical direction for the platform, prioritize ruthlessly between env work, pipeline, and admin firefighting, and leave the system more instrumented and operable than you found it. Experience
- Owned a multi-sided platform (operators + contributors + internal stakeholders), not only feature slices.
- Built or deeply operated systems in human data, RLHF, labeling, or model evaluation - or very close: RL/eval harnesses, annotation pipelines, coding/STEM eval environments.
- Shipped admin or back-office workflows: approvals, roles/permissions, payments or payouts, and operational metrics.
- Led a technical initiative with high ambiguity on a small team.
Qualifications
- Senior software engineer (4+ years, or equivalent ownership): backend or full-stack, comfortable owning a production web product and the services behind it.
- Proven autonomy, you can take a messy domain (human data / evals) and turn it into a roadmap and shipped systems.
- Working fluency with how frontier human-data and evaluation work actually runs (tasks, QC, environments, expert workflows) enough to design for it, not only implement specs.
- Nice to have: env orchestration (containers, workers, sandboxes), prior work at or adjacent to Scale AI / Surge / Handshake AI / Labelbox / Other platforms, modern web/data stack familiarity.
- Fluent English. Spanish is a nice-to-have.

