1,368,003open jobs
79,527companies
208,804added this week
Browse all
Salary
≈ $40k – $120k per year (Estimated)
Location
Hybrid (Amsterdam, Netherlands)
Visa
Sponsorship offered in the posting · Recognised sponsor of the Dutch immigration service (IND)

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Jul 14, 2026.

Overview
Company
Impact
Profile match
Kaizo gives support leaders full visibility across every conversation. Automated QA, AI coaching, and performance insights.

Do you want to work out how you actually measure whether an AI system is doing a good job, and then make it better? We're looking for a sharp, analytical data scientist to own the evaluation and quality loop of our AutoQA product. You can be early in your career (recent graduates with strong LLM fundamentals are welcome), we have room and a clear path for you to grow.

In a nutshell

  • Join a fast-growing SaaS company in an international environment (steep learning curve guaranteed).

  • Own a high-impact problem: making LLM-powered quality assurance measurably accurate at scale.

  • Sit at the intersection of AI and CX, working directly with enterprise customers and their real-world QA rubrics.

  • Grow into a Senior Data Scientist, AI Engineer, or ML Engineer role. We invest in progression.

  • Enjoy the perks: flexible hours, open holiday policy, an office in the heart of Amsterdam with hybrid flexibility, visa sponsorship, great gear, workations, and team events.

About Kaizo

At Kaizo, we build a performance development and quality platform for customer support teams. Our AutoQA product uses LLMs to review support conversations against each customer's own quality rubric, automatically and at scale. Behind it sits a microservices-based stream processing platform handling over 200 million events per day (Kafka, Kubernetes on Google Cloud, ElasticSearch, MongoDB, BigQuery), and an LLMOps stack built around LangSmith for experimentation, prompt management, and tracing.

The hard part isn't calling an LLM. It's knowing, with evidence, how well the system performs on every customer's unique rubric, and having a reliable, repeatable way to improve it. That's where you come in.

What you'll focus on

  • Translate customer rubrics into AutoQA instructions. Work with real customer quality criteria and turn them into precise, testable instructions that LLMs can score reliably.

  • Run experiments that move accuracy. Design and execute evaluation experiments on large, representative datasets using LangSmith and BigQuery, and track quality with our performance metrics.

  • Build the datasets that make evaluation possible. Curate raw production data into golden datasets with balanced coverage, and generate synthetic data to cover the rare cases that matter most. QA is a discipline of rare occurrences: distributions are skewed, positives are scarce, and resourcefulness beats volume.

  • Build LLM-as-a-judge pipelines to assess system quality internally and make evaluation repeatable.

  • Make results actionable. Your experiments should end in a decision: change a prompt, adjust which tools the system uses, surface context the AI is missing, or flag where new capabilities are needed. You'll work with the AI team to ship those decisions.

  • Evaluate across the full stack. Beyond scoring quality, you'll help validate retrieval (RAG/IR), tool calling, and speech pipelines (transcription and diarization quality).

  • Join customer calls with the team to understand how QA leaders define quality, and feed what you learn back into the product.

What you'll grow into

  • Shaping how customers monitor quality themselves, catch drift, and keep their AutoQA setup improving over time.

  • Smarter categorization and routing of conversations to power analytics and get the right tickets to the right evaluation.

Requirements

  • A solid understanding of how LLMs work and hands-on experience prompting them for accuracy (coursework, thesis, internships, or side projects all count; production experience is a bonus).

  • Good applied statistics: experiment design, classifier evaluation, precision/recall trade-offs, and working with heavily imbalanced data.

  • Strong Python skills and fluency with the standard data toolkit (Pandas, NumPy, Jupyter). SQL is a plus.

  • An analytical, evidence-first mindset: you'd rather measure than assume.

  • Product sense and empathy for end users. You'll be building for QA managers and support agents, not just for benchmarks.

  • Excellent written and verbal communication. You'll present findings to the team and join customer conversations.

  • 0 to 2 years of industry experience. Recent graduates with strong relevant work are encouraged to apply.

  • A team player who's comfortable wearing multiple hats. We're an early-stage company and things move fast.

Bonus points for:

  • Experience with LangSmith or similar LLMOps/evaluation tooling

  • Google Cloud Platform, BigQuery, Docker, or Kubernetes

  • Speech/audio processing or ASR evaluation

  • Fine-tuning or building synthetic datasets for LLMs

Who you'll work with

You'll join our AI team (two data scientists, an AI engineer, and an ML engineer) and collaborate closely with our data engineers, frontend engineers, designer, product manager, and CX teams. You'll have mentorship from day one and real ownership fast.

What's in it for you?

  • An office in the heart of Amsterdam, with the flexibility of hybrid working

  • Visa sponsorship available for eligible candidates

  • Great office gear: MacBook, tools, desk, chair, whatever you need

  • Flexible working schedule and an open holiday policy

  • Fun workations and team events

  • A clear growth path into Senior Data Scientist, AI Engineer, or ML Engineer roles

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,368,003 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Amsterdam
Data Engineer 3 hours ago
≈ $52k – $101k per year (Estimated) • Hybrid • Full-Time • Deventer
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Azure DevOps
Azure
Analytics
Power BI
ETL/ELT
Management
Agile
Scrum
Apply
$60k – $87k per year • Remote (Netherlands)
Python
SQL
Databases
Snowflake
AI/ML
Airflow
dbt
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Docker
Analytics
Fivetran
Apply
≈ $78k – $186k per year (Estimated) • In office • Full-Time • 's-Hertogenbosch
AI/ML
AI Agents
Machine Learning
Apply
Data Scientist 1 day ago
≈ $67k – $113k per year (Estimated) • In office • Poland
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
AI/ML
Spark
Kedro
Scikit-learn
TensorFlow
Pandas
NumPy
PyTorch
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Apply
≈ $65k – $109k per year (Estimated) • In office • Poland
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
AI/ML
Spark
dbt
DevOps
GCP
Azure
AWS
Analytics
Tableau
Power BI
ETL/ELT
Looker
Management
Agile
Apply
DevSecOps Engineer 15 days ago
≈ $41k – $109k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Singapore
DevOps
Terraform
Ansible
GCP
Helm
OpenTelemetry
Azure
CI/CD
Jenkins
Git
Docker
Kubernetes
Canary Release
Shift-Left
Configuration Management
Helmfile
Linux
Cybersecurity
Shift-Left Security
Management
Agile
Apply
≈ $20k – $52k per year (Estimated) • In office • Full-Time • 10+ years exp • Bengaluru
Python
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Polly
gRPC for .NET
Databases
Apache Kafka
AI/ML
ML.NET
Semantic Kernel
Pandas
NumPy
OpenAI
Machine Learning
Frontend
GraphQL
DevOps
Rest API
gRPC
Splunk
Terraform
Ansible
GCP
OpenShift
Azure DevOps
GitHub Actions
CloudFormation
Prometheus
GitLab CI
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Grafana
Amazon EKS
Azure AKS
API Gateway
Cybersecurity
SonarQube
QA
Swagger
Postman
Apply
SDE 2 - Node.js 15 days ago
In office • 8+ years exp • Bengaluru
JavaScript
TypeScript
Node JS
Node JS
Express
Databases
Redis
RabbitMQ
Apache Kafka
Frontend
GraphQL
Next.js
React.js
DevOps
Rest API
Splunk
GCP
New Relic
Datadog
Azure
AWS
Docker
Kubernetes
Management
Agile
QA
Jest
Mocha
Apply
≈ $17k – $41k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • India
Python
SQL
Python
Django
Django REST Framework
Databases
PostgreSQL
AI/ML
LangChain
LLM
OpenAI
Anthropic
DevOps
AWS
Docker
Kubernetes
AWS Lambda
Amazon EC2
Amazon S3
Apply
≈ $14k – $31k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Manila
Python
SQL
Databases
Snowflake
DevOps
GCP
Azure
AWS
Analytics
Tableau
Power BI
Alteryx
Looker
Apply
≈ $106k – $209k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Paris • Budapest • Barcelona • Amsterdam • Dublin
Python
Java
C#
Julia
SAS
Apply
≈ $74k – $131k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Amsterdam
DevOps
GCP
Analytics
Tableau
Looker
Apply
$78k – $90k per year • Hybrid • Full-Time • 3+ years exp • Amsterdam
Apply
≈ $54k – $104k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Madrid • Paris • London • Helsinki • Amsterdam
DevOps
Incident Management
Cybersecurity
GDPR
Apply
≈ $35k – $73k per year (Estimated) • In office • Full-Time • Amsterdam
Apply
See all jobs
This is one of many
1,368,003 more open roles from verified company boards, updated every day.