369,078open jobs
9,456companies
47,986added this week
Browse all
Salary
up to $70k per year
Location
Remote (United States)
Employment
Freelance
Overview
Company
Impact
Profile match
Mindrift is an expert-sourcing and AI training platform owned by Toloka (a data services enterprise owned by Nebius Group, Nasdaq: NBIS), operating primarily out of Amsterdam, Netherlands. Launched by Toloka CEO and founder Olga Megorskaya to elevate crowd-sourced data generation into high-level human feedback, Mindrift operates on a B2B data-as-a-service (DaaS), custom enterprise AI training contract, and freelance expert marketplace model.

Please submit your CV in English and indicate your level of English proficiency.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

About the Role

You’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

Responsibilities:

  • Invent a realistic developer scenario - a real bug, a broken ETL, a missing feature - not a toy problem.
  • Build a reproducible Docker environment with pinned dependencies.
  • Write a pytest that verifies outcomes, not specific commands - deterministic, non-flaky, and does not leak the fix.
  • Write an instruction.md that reads like a Jira ticket a developer would receive.
  • Write a reference solve.sh proving the task is solvable.
  • Calibrate difficulty so current state-of-the-art agents solve the task 20-60% of the time.
  • Iterate based on feedback from expert QA reviewers.
  • Later: review other authors’ tasks as a QA reviewer.

Not in scope

  • Data labeling, prompt engineering.
  • Production code to ship - you design problems and verification for AI agents.
  • Leetcode puzzles - scenarios must look like real developer work.
  • Not every candidate task ships - quality over quantity.

Requirements

  • 3+ years of production software development in one backend stack - Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
  • Python + pytest fluency - required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
  • Docker authoring - reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
  • Linux & Bash - comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
  • AI coding agent experience - Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
  • English - B2+ written.

Not a fit

  • Data Science, ML, or Computer Vision engineers without backend-engineering output.
  • Manual QA testers without automation or test authoring.
  • Frontend-only, low-code / no-code, IT Support, or Business Analysts.
  • Engineers who have never written pytest from scratch.
  • Junior, intern, or assistant as the most recent role.

Preferred qualifications

  • Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
  • Modern Python tooling (uv, poetry, pyproject.toml).
  • Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov).
  • Fuzzing or property-based testing (Hypothesis).
  • Prior contribution to agent-evaluation benchmarks or related frameworks.

Process

Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.

Time commitment

  • Onboarding: ~10 hours per first task.
  • Steady state: ~5 hours per task, 2-4 parallel tasks per author.
  • Realistic weekly load: 8-20 hours. Higher volume available for top performers.
  • You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.

Compensation:

  • Paid contributions, rates up to $35/hour *.
  • Task-based compensation equivalent to hourly rate, depending on performance and volume.
  • Some projects include incentive payments.

*Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.

Apply

Submit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist - Terminal Bench), and include a GitHub profile link if available.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
369,078 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$87k – $166k per year (Estimated) • Equity • Remote • Full-Time • PhD
Python
Ruby
SQL
JavaScript
Ruby
RSpec
Ruby on Rails
Databases
PostgreSQL
AI/ML
Claude
Claude Code
AI Agents
OpenAI Codex
Frontend
GraphQL
Vue.js
DevOps
GitLab
Management
Slack
Apply
$47k – $106k per year (Estimated) • In office • Full-Time • Moscow
Python
AI/ML
AutoGen
Chain-of-Thought
Claude
CrewAI
Gemini
LangChain
LangGraph
LLM
Prompt Engineering
RAG
GPT-4
AI Agents
Function Calling
Apply
$28k – $65k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Bash
PowerShell
Python
Node JS
JavaScript
Node JS
Commander.js
AI/ML
AI Agents
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
Kubernetes
Amazon ECS
IAM
Cybersecurity
Crowdstrike
Zero Trust
Apply
$16k – $36k per year (Estimated) • Remote • Full-Time • Moscow
AI/ML
ChatGPT
Claude
Claude Code
Cursor
LLM
Model Context Protocol
Apply
$23k – $54k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Hyderabad
AI/ML
AI Agents
Agentforce
DevOps
SLI/SLO/SLA
Analytics
Power BI
Tableau
Marketing
Salesforce
Apply
up to $90k per year • Remote • Freelance
JavaScript
Python
Python
Beautiful Soup
AI/ML
LangChain
LLM
OpenRouter
DevOps
AWS
Docker
GitHub
QA
Selenium
Apply
up to $50k per year • Remote • Freelance • Buenos Aires
JavaScript
Python
Python
Beautiful Soup
AI/ML
LangChain
LLM
OpenRouter
DevOps
AWS
Docker
GitHub
QA
Selenium
Apply
up to $80k per year • Remote • Freelance • Bucharest
JavaScript
Python
Python
Beautiful Soup
AI/ML
LangChain
LLM
OpenRouter
DevOps
AWS
Docker
GitHub
QA
Selenium
Apply
up to $80k per year • Remote • Freelance • Warsaw
JavaScript
Python
Python
Beautiful Soup
AI/ML
LangChain
LLM
OpenRouter
DevOps
AWS
Docker
GitHub
QA
Selenium
Apply
up to $50k per year • Remote • Freelance • São Paulo
JavaScript
Python
Python
Beautiful Soup
AI/ML
LangChain
LLM
OpenRouter
DevOps
AWS
Docker
GitHub
QA
Selenium
Apply
See all jobs
This is one of many
369,078 more open roles from verified company boards, updated every day.