542,849open jobs
19,763companies
74,611added this week
Browse all
Location
Remote (Serbia)
Seniority
Senior · 4+ years exp
Overview
Company
Impact
Profile match
Toloka is a data company founded in 2014 and headquartered in Amsterdam that supplies human-generated training, fine-tuning and evaluation data for artificial intelligence systems. It began as a large-scale crowdsourcing platform and has repositioned around expert work, combining domain specialists, machine learning tooling and quality control pipelines to produce reasoning traces, preference rankings and agentic task data. Operating within the Nebius group and backed by Bezos Expeditions since 2025, the company serves frontier model laboratories and enterprises building their own systems.

About Toloka

AtToloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety. We have experts in more than 50 different domains-from doctors and lawyers to physicists and engineers-and boast one of the most diverse global crowds,representing ove r100 countries and speaking 40+ languages. We are a well-funded startup with an enviable portfolio of clients includingAnthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.

Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation fromMikhail Parakhin, CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board. Our remote-first team is globally distributed around the world: USA, UK, the Netherlands, Serbia, and more.

About the Team

We are the ML team inside Toloka - we build the machine-learning products that power the platform itself, so every project running on Toloka is faster, cheaper, and more reliable.

A few examples of what we own:

  • LLM QA - the core technology behind Toloka's automated quality-check mechanism. Every annotation flowing through Self-Service is reviewed by an LLM agent we design, train, and operate.
  • Model distillation and fine-tuning - adapting frontier and open-source models to Toloka's tasks to hit the right quality at the right cost.
  • Evaluation, benchmarking, cost modeling, and model selection across providers.

We own the full chain. The same team designs the ML solution, ships it to production, keeps it running 24/7, analyzes the results coming back from real projects, and feeds that signal into the next iteration. No hand-off between research, engineering, and operations - it's all us.

About the Position

You will own Toloka’s end-to-end fine-tuning, RL, and evaluation stack, bridging applied research and product engineering. In this role, you will spearhead greenfield post-training initiatives (such as GRPO and reward modeling), transform complex ML experiments into scalable platform features, and occasionally author technical write-ups on your findings for the AI community.

What you’ll do

  • Own end-to-end fine-tuning pipelines: data prep, SFT/LoRA training, distillation from frontier models to smaller ones, evaluation, and serving handoff - both as self-serve platform capabilities and in hands-on client engagements.
  • Extend our post-training stack beyond SFT into RL (RFT/GRPO-style methods, reward modeling, LLM-judge-based rewards) and help design the user-facing RL flow on the platform - this part is greenfield.
  • Build and calibrate evaluation harnesses: LLM-as-judge setups calibrated against human labels, golden datasets, regression evals for optimization runs.
  • Improve the platform's guiding agent: prompt and tool design, eval-driven improvement loops, stress-testing scenarios and fixing what breaks.
  • Run experiments for client projects (e.g. prompt compression vs. fine-tuning trade-off studies) and turn the results into repeatable platform features.
  • Work closely with platform engineers on the SDK/API surface so that training and eval jobs are callable from a developer's existing workflow.

What we're looking for

  • 4+ years in ML engineering or applied research, with at least 1-2 years hands-on with LLMs in production or research settings.
  • Practical experience fine-tuning open-weight models (LoRA/full FT), including data curation and knowing when fine-tuning is the wrong answer.
  • Solid grasp of LLM evaluation: building evals from scratch, LLM-as-judge pitfalls, calibration against human judgments.
  • Strong Python engineering: you write code others can run, not just notebooks; comfortable with the training/inference stack (PyTorch, HF ecosystem, vLLM or similar).
  • Product mindset: you'll often be the ML person closest to a client problem, so you need to reason about what's worth building, not just what's possible.
  • Comfortable with ambiguity - priorities shift as we learn from pilots.
  • Language: Fluency in English (B2 or above).

Nice to have

  • Hands-on RL for LLMs: GRPO/PPO-style post-training, reward modeling, RLHF/RLAIF pipelines.
  • Prompt optimization frameworks (DSPy/GEPA or similar) or prompt-compression research (gisting).
  • Experience with distillation and quantization for cost/latency optimization.
  • Experience building agentic systems (tool use, multi-step workflows) or shipping ML features in a self-serve product.

What we can offer

  • You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
  • Competitive compensation package including base salary, bonus, and ESOP.
  • Paid PTO and benefits will vary depending on location.
  • We offer a full remote or hybrid model (if you are based in NL or Serbia).
  • IT setup and home office allowances.

Equal Opportunity Employer:

Toloka is committed to providing equal opportunity and fostering an inclusive environment. We welcome applications from all qualified individuals and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, marital status, veteran status, disability, or any other characteristic protected by applicable law. Selection decisions are made based on qualifications, merit, and business need.

[Important Notice] Scam Alert Regarding Fake Job Postings

It has come to our attention that an individual or group is fraudulently impersonating Toloka to post fake jobs and solicit personal information from applicants. Please be aware:

  • Official Communication: Our recruiting team will only contact you from an official "toloka.ai" email address. We will NEVER use Gmail, Yahoo, Tolokainc, toloka.inc, or other personal or seemingly business email accounts.
  • Our Process: We will never ask for your bank account details, credit card number, or any fees as part of the application or interview process.
  • Official Listings: All legitimate job openings are posted on our official careers page: https://toloka.ai/careers#job-list

What to do: If you see a suspicious job posting or have been contacted by someone you suspect is a scammer, please do not provide any personal information. Instead, report the incident to us directly at [email protected] and report the profile/post to LinkedIn.We are taking this matter very seriously and are working with the appropriate parties to resolve it.

Thank you for your vigilance!

To learn how we collect, use, disclose, and store personal data, check out our Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
542,849 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
In office • Internship • Singapore
Python
SQL
AI/ML
Anomaly Detection
Analytics
Tableau
Power BI
Apply
DevOps Engineer 2 hours ago
$16k – $45k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Bash
DevOps
Terraform
GCP
Helm
Azure DevOps
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
SRE
Platform Engineering
IAM
Cybersecurity
ISO 27001
SOC 2
Apply
Research Assistant II 2 hours ago
$69k – $153k per year (Estimated) • In office • Full-Time • 1+ year exp • Associate's Degree • United States
Python
Apply
Data Analyst II 2 hours ago
$74k – $153k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • United States
Python
SAS
SPSS
Stata
Apply
Firmware Engineer I 3 hours ago
$28k – $60k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Ho Chi Minh City
Python
Go
Rust
C++
Go
Chi
AI/ML
AI Agents
DevOps
RTOS
Management
WhatsApp
Apply
$36k – $67k per year (Estimated) • Remote • Freelance • 1+ year exp
AI/ML
Anthropic
Management
Gmail
Google Drive
Apply
$42k – $111k per year (Estimated) • Remote • Freelance
AI/ML
Anthropic
Management
Gmail
Marketing
Shopify
LinkedIn
Apply
$115k – $256k per year (Estimated) • Remote
SQL
AI/ML
Anthropic
Analytics
Tableau
Power BI
Microsoft Excel
Management
Gmail
Google Sheets
Apply
Remote • Freelance
AI/ML
Claude
AI Agents
Anthropic
Agentic Workflows
Management
Gmail
Marketing
Shopify
LinkedIn
Apply
Remote • Freelance
AI/ML
Claude
AI Agents
Anthropic
Agentic Workflows
Management
Gmail
Marketing
Shopify
LinkedIn
Apply
See all jobs
This is one of many
542,849 more open roles from verified company boards, updated every day.