433,445open jobs
14,773companies
64,942added this week
Browse all
Salary
$67k – $137k per year (Estimated)
Location
In office (Berlin, Germany)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Wieden+Kennedy is an American advertising agency founded in Portland in 1982 by two men who had worked on Nike's account and who wrote the line Just Do It three years later. It remains independently owned rather than part of a holding company, which has allowed it to keep an unusually strong creative culture and to turn down work that conflicts with existing clients. The agency works for Nike, Coca-Cola, Ford, McDonald's and other large brands from offices in Portland, New York, London, Amsterdam, Tokyo, Shanghai, Delhi, Mexico City and Sao Paulo.

About Us

There are hundreds of thousands of lawyers across Europe, and at Libra we're

transforming how they work. Our AI platform combines deep legal reasoning

with cutting-edge generative technology, fundamentally changing how lawyers

research, draft, and deliver legal work.

We operate with the speed and ownership of a startup, while being backed by

the scale and stability of an established leader. Libra is an independent

business unit within the Legal & Regulatory division of Wolters Kluwer, a leading

global provider of information, software, and services for professionals. For 180

years, Wolters Kluwer has supported and simplified the work of experts and

organizations through innovative solutions, relying today on more than 20,000

colleagues worldwide to bring that vision to life.

About The Role

As an Evals Engineer (m/w/d) at Libra, you'll be the first dedicated hire on a new

team that owns how we measure quality: the datasets, rubrics, judges and

harnesses that decide whether an AI feature is good enough to ship, and the

automated optimization that runs against them.

Optimizing a system against a target is newly automatable: LLM optimizers like

GEPA now propose and test the candidates themselves. Nobody has to invent

the experiments any more, and the system will improve in whatever direction

the evals point, whether or not that's where you meant to go. That makes

defining what good means the highest-leverage work we do, and it has to be

done task by task, from scratch, for a legal AI platform at European scale.

You'll work in a lean, agile environment with AI Engineers, Legal Engineers and

Product, based at the vibrant Merantix AI Campus in Berlin, surrounded by a

community of AI innovators.

What you'll do

  • Build and own the eval platform: datasets, judges, harnesses, regression gates, cost and quality in one view, on Python, FastAPI and Langfuse. Make it self-service, so AI Engineers can evaluate and tune their own features without going through you, and eval-driven development becomes the most effective path rather than a tax.

  • Map the quality landscape: good means something different for research, drafting, summarisation and retrieval, and again per jurisdiction. Work out what a defensible measure looks like for each, going first on the ones nobody has evaluated before and then making them repeatable without you.

  • Design the rubrics and set the standard for LLM-as-judge: turn Legal Engineers' and subject-matter experts' judgment into version-controlled criteria, keep judges recalibrated as models and jurisdictions change, and make authoring cheap enough to do at volume on privileged material.

  • Deep-dive results and traces until you can say why something failed, then find the lever that moves it and automate the fix.

  • Unhobble the optimizer: instrument the app so prompts, hyperparameters and harness architecture become levers a search can safely pull, then run automated optimization (GEPA, DSPy) over them against a fitness function you trust, widening that surface as you go.

  • Make cheaper models win: treat quality per euro as a first-class metric, instrumented per call, and find the configuration where a smaller model matches or beats an expensive one.

  • Own guardrails: ungrounded advice, invented or misattributed citations, jurisdiction and language leakage, prompt injection from ingested documents. Design them, red-team them, prove they hold.

What you'll bring

Education

  • Bachelor's degree or equivalent in a relevant technical field (e.g. Computer Science, Software Engineering, Statistics, Data Science); advanced degree is a plus.

Experience

  • Minimum 5 years in software engineering,at least 1 building LLM-powered products in production.

  • StrongPython: FastAPI, modern tooling, and the data stack (pandas, numpy, notebooks), because much of this job is analysis.

  • Hands-on experience designing evaluations: datasets, rubrics, LLM-as-judge, benchmarking, human labelling.

  • Solid security and data-privacy practice. You'll handle traces, documents and datasets derived from privileged legal material.

  • AI coding agents (Claude Code, Codex, Cursor) in your daily workflow.

  • Genuine interest in the legal domain.

  • Bonus: automated prompt or pipeline optimization (GEPA, DSPy or similar), and a broader data science toolkit, e.g. embedding clustering to check dataset coverage.

Skills

  • A strong engineer who hasn't hand-written code in months. The architecture is yours, the typing isn't. You ship more working software than you ever did alone, reject code that runs but is shaped wrong, and leave less rework and cognitive debt behind you.

  • You think in systems and expect them to be gamed. The app, the evals, the optimizer and the people using them are one loop. Once evals gate releases, everything optimizes toward them, so some judgments stay human.

  • Hard to fool by a single number. "Could these all be within variance?" comes before any ranking, and a judge's reliability before you trust its grades. You report bounds, and retire a result that doesn't hold up, including your own.

  • You get expertise out of people who have no time to give you any, and build tools they choose to use. An eval nobody runs is worth nothing.

  • Runs on macro-management. You ask the sharp questions up front, then come back with options, their trade-offs and the assumptions behind each, and help pick. Good to think out loud with. Entrepreneurial, accountable, pragmatic.

  • Excellent communication in English.

Our Interview Practices

To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we’re getting to know you-not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process.

Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
433,445 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Berlin
$23k – $55k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Pune
Python
SQL
Python
Flask
FastAPI
Django
Databases
PostgreSQL
Redis
Oracle
Cassandra
AI/ML
Copilot
Claude
SciPy
AI Agents
Pandas
NumPy
Devin
DevOps
GCP
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Platform Engineering
Bitbucket
GitHub
GitLab
Cybersecurity
OWASP Top 10
QA
Pytest
Apply
$18k – $52k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Hyderabad • Bengaluru
Python
PowerShell
Databases
PostgreSQL
Snowflake
Amazon Redshift
AI/ML
AI Agents
DevOps
Self-Healing
IAM
Analytics
Tableau
Power BI
Management
Google Workspace
ServiceNow
Apply
$14k per year • In office • Full-Time • 2+ years exp • Ufa
Python
Python
Django
Databases
PostgreSQL
DevOps
Rest API
Management
Telegram
Apply
$24k – $67k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bucharest
Python
Java
C++
Apply
Remote/Hybrid • Full-Time • Gurgaon
Python
SQL
SAS
Python
pySpark
AI/ML
Spark
AI Agents
Apply
$18k – $46k per year (Estimated) • In office • Full-Time • Master's Degree • Beijing • Shanghai
Apply
$16k – $37k per year (Estimated) • In office • Full-Time • 8+ years exp • Pune
SQL
Apply
$94k – $206k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Berlin
AI/ML
AI Agents
LLM
RAG
Human-in-the-Loop
Apply
In office • Full-Time • Bachelor's Degree • London
Management
Outlook
Apply
$53k – $128k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Madrid • Milan
Analytics
Power BI
Marketing
Salesforce
Apply
In office • High School Diploma • Berlin
Web3
Smart Contracts
Apply
$78k – $202k per year (Estimated) • Remote • Berlin
AI/ML
EU AI Act
DevOps
Platform Engineering
Cybersecurity
ISO 27001
Apply
$34k – $76k per year (Estimated) • Remote/Hybrid • Internship • Berlin
Management
Slack
Google Sheets
Apply
$53k – $121k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Berlin
Python
Java
SQL
Databases
MySQL
PostgreSQL
AI/ML
Hadoop
Spark
Airflow
DevOps
GCP
Cybersecurity
GDPR
Analytics
ETL/ELT
Apply
$87k – $105k per year • In office • Full-Time • 5+ years exp • Berlin
JavaScript
Apply
See all jobs
This is one of many
433,445 more open roles from verified company boards, updated every day.