665,268open jobs
38,876companies
100,271added this week
Browse all
Salary
$130k per year
Location
Remote (Argentina, Brazil, Colombia, Mexico, Spain, Chile, Ecuador, Portugal, Uruguay)
Seniority
Middle · 3+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match
Anyone AI is an AI talent company that trains software developers, mainly from Latin America, for careers in artificial intelligence and supplies experts who create coding and STEM training data for frontier AI models. Its founders were members of the founding team of Deep Vision AI, and it is venture-backed by investors including Global Founders Capital, Canvas Ventures and Latitud. Its openings are remote, project-based contract roles for full-stack and Python developers and physics experts who design and review tasks for leading AI labs, recruited in countries across Europe and Asia.

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

  • Experiment design and model selection

  • Small and synthetic datasets

  • Data quality and preprocessing

  • Distribution shift and data contamination

  • Label noise and feature leakage

  • Model evaluation and metric selection

  • Hyperparameter tuning

  • Train / validation / test methodology

  • Reproducibility and deterministic pipelines

  • Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For

  • 3+ years of hands-on applied machine learning experience

  • Strong experience with:

    • ML experiment design

    • Model selection

    • Hyperparameter tuning

    • Model evaluation

    • Data preprocessing and validation

  • Strong understanding of train, validation, and test splits

  • Ability to identify:

    • Data leakage

    • Label noise

    • Distribution shift

    • Spurious correlations

    • Feature leakage

    • Data contamination

  • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations

  • Strong understanding of ML evaluation metrics and when different metrics are appropriate

  • Experience debugging ML workloads across CPU and GPU environments

  • Ability to analyze technical problems and provide clear written feedback

Nice to Have

  • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions

  • Experience designing benchmark datasets or ML challenges

  • Background in data-centric AI or dataset quality

  • Experience with synthetic data generation and validation

  • Familiarity with statistical testing, confidence intervals, and effect sizes

  • Experience with ML evaluation pipelines, RLHF, or AI model evaluation

  • Experience developing ML curricula or technical assessments

  • Understanding of common ML failure modes such as:

    • Shortcut learning

    • Spurious correlations

    • Goodhart’s Law

    • Simpson’s paradox

    • Metric gaming

What You’ll Be Responsible For

  • Reviewing ML challenges and determining whether they are well designed and technically solvable

  • Evaluating whether datasets contain meaningful and learnable signals

  • Identifying unintended shortcuts or artifacts in synthetic datasets

  • Determining whether tasks require genuine diagnosis of the underlying ML problem

  • Reviewing evaluation metrics and improvement thresholds

  • Detecting metric gaming, data leakage, and evaluation flaws

  • Verifying reproducibility across the complete data → model → evaluation pipeline

  • Assessing whether challenge difficulty is appropriately calibrated

  • Providing clear recommendations for improving, recalibrating, or excluding problematic tasks

Engagement

Work Type: Remote

Engagement: Part-time, project-based consulting

Focus: Applied machine learning, experiment design, data quality, and model evaluation

This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
665,268 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$25k – $53k per year (Estimated) • Remote/Hybrid • 3+ years exp • Saint Petersburg
Python
Python
FastAPI
AI/ML
LoRA
Fine-tuning
RLHF
AI Agents
NLP
PEFT
Transformers
LLM
Apply
$130k per year • Remote • Contractor • 3+ years exp
Python
SQL
AI/ML
RLHF
DevOps
Docker Compose
CI/CD
Docker
Cybersecurity
CVE
CWE
CVSS
pwntools
Apply
$130k per year • Remote • Contractor • 3+ years exp
AI/ML
CUDA Toolkit
RLHF
JAX
CUDA
Triton
TPU
cuDNN
AWS Trainium
MLIR
XLA
DevOps
AWS
Apply
$130k per year • Remote • Contractor • 2+ years exp
AI/ML
CUDA Toolkit
RLHF
CUDA
Triton
AWS Trainium
XLA
DevOps
AWS
Apply
$25k – $60k per year (Estimated) • Remote • Full-Time • 8+ years exp
Go
TypeScript
SQL
Go
Temporal
AI/ML
AI Agents
Arize Phoenix
DeepEval
Langfuse
LangSmith
Promptfoo
LLM
RAG
Synthetic Data
Human-in-the-Loop
DevOps
GitHub Actions
OpenTelemetry
CI/CD
Cybersecurity
SOC 2
GDPR
QA
Playwright
k6
Locust
Apply
$130k per year • Remote • Contractor • 3+ years exp
Python
SQL
AI/ML
RLHF
DevOps
Docker Compose
CI/CD
Docker
Cybersecurity
CVE
CWE
CVSS
pwntools
Apply
$130k per year • Remote • Contractor • 3+ years exp
AI/ML
CUDA Toolkit
RLHF
JAX
CUDA
Triton
TPU
cuDNN
AWS Trainium
MLIR
XLA
DevOps
AWS
Apply
$130k per year • Remote • Contractor • 2+ years exp
AI/ML
CUDA Toolkit
RLHF
CUDA
Triton
AWS Trainium
XLA
DevOps
AWS
Apply
$90k – $160k per year • Remote • Contractor • 2+ years exp
Python
Go
JavaScript
TypeScript
C#
Mobile
JUnit
QA
Jest
Pytest
Apply
$90k – $160k per year • Remote • Contractor • 2+ years exp
Python
Go
JavaScript
TypeScript
C#
Mobile
JUnit
QA
Jest
Pytest
Apply
See all jobs
This is one of many
665,268 more open roles from verified company boards, updated every day.