At Pencil, we are driving innovation in advertising technology through our state-of-the-art SaaS product, which harnesses Generative AI to redefine content creation. Our mission is to make AI the default in advertising without replacing creative people. To achieve this, we need to make sure that our technology isn't just in the hands of big brands - we need to try to help small businesses and creative individuals too.
The role:
Pencil's agents produce advertising creative at scale - video, adaptation and multiple formats, orchestrated by our agentic system, Scribble. The product only works if the output is right first time. This role owns measuring quality and improving it.
You'll own evals end to end. Internally, that means the systems that measure whether agents, skills and workflows produce work a client would ship. Externally, it means how we demonstrate quality to customers and the market: the evidence behind client RFP responses, creative enablement packages, and adevals.ai, our public evals site.
Measuring creative quality is hard because there is no single right answer. Your job is to build reliable measurement anyway: rubrics that define what good means, calibration between human reviewers and automated judges, and a golden dataset that stays representative as the product and client base evolve. You'll then make those metrics central to how the business operates and sells.
Key responsibilities:
Own the charter for Evals & Quality: the durable problem, success metrics, and an explicit in/out scope list.
Define the core quality metrics - first-pass-right, coverage, regression detection - and drive them up.
Own the judging methodology: rubrics for creative quality, and the calibration process that keeps automated judges and human raters aligned over time.
Design and run the human QC mechanism: where humans review output, how their judgements feed the automated layer, and how it scales.
Own the golden dataset: curation, coverage across formats and client contexts, and versioning.
Build eval loops that other teams use themselves. A team shipping a new video skill should be able to set up evals without your involvement.
Own adevals.ai as a product, including its design quality. The site needs to meet a high design bar to be credible.
Provide the evidence layer for commercial work: client RFP responses and creative enablement packages.
Partner with Agent Architects on client-specific quality bars, delivered through configuration and rubrics rather than custom eval code.
Run the full product lifecycle: Discovery through post-launch audit, with a written hypothesis before every ship and a review 14 days after launch.
Your background:
You might be a great fit if you can walk us through:
A quality, evals or ML measurement problem you owned as a product, with its own roadmap and users.
Measurement you shipped in a subjective domain: how you defined quality, and how you kept the signal reliable as it scaled.
Direct experience with the mechanics: golden datasets, rubric design, judge calibration, inter-rater reliability. Where they failed and what you did.
A product decision you derived from evidenced customer need in a domain you had to learn. If your instinct is to start with a solution, this isn't the role.
A metric you owned that changed decisions outside your own team.
Time spent in front of clients: presenting or defending methodology to buyers.
You'll Thrive Here If You…
Move Fast. You're comfortable making decisions with imperfect information, iterating quickly, and fixing things before they become fires;
Stay Human. You care deeply about users and teammates, and design with empathy for real workflows;
Act Like an Owner. You take responsibility for outcomes, not just outputs. You spot problems early and take initiative to fix them.
Keep It Simple. You fight unnecessary complexity, design reusable patterns, and make hard things feel intuitive.
Love to Surprise. You enjoy finding unexpected ways to delight users, especially in the last mile of polish and usability.
KPIs & Success Measures
We'll measure success through outcomes and quality, not just output:
First-pass-right: share of generations a client would ship without rework. The lead metric.
Eval coverage: share of skills and formats with live eval loops.
Judge-human calibration: agreement between automated judges and human raters, maintained over time.
Golden dataset health: representative, versioned and current across formats and client contexts.
Commercial evidence: adevals.ai live; evals used in RFP wins and enablement packages.
Learning velocity (supporting): every launch reviewed 14 days post-launch against its stated hypothesis.
Benefits:
25 days PTO plus public holidays, although we operate a Flexible Time Off scheme
Health insurance / private medical cover
Monthly stipend towards wellness, fitness, and learning and development
Remote - work from anywhere in your home country
Enhanced parental leave policies, whether you become a parent through birth, adoption or surrogacy
Access to our Pencil office in The Shard, London for our UK employees
Flexible working hours
Deep Dive About Pencil
Pencil uses AI to make ads. Our mission is to make marketing effective and effortless. We want to become the default way ads get made - because AI ads are 10x faster and cheaper to make, and 2x better performing, than making them without AI.
We're called Pencil because we believe AI will be a tool for creative people, not a replacement - it may even be as fundamental a tool in the future as the pencil was in the past.
Pencil was founded in 2018 with a team from Google, Facebook and Uber with backing from Sequoia and Entrepreneur First. We were acquired by The Brandtech Group in 2023 to pursue a shared vision of bringing GenAI to the Fortune 500.
About The Brandtech Group
The Brandtech Group's mission is to be the best company in the world at helping leading global brands drive growth by connecting content, data and media using technology.
It was founded in June 2015 (as You & Mr Jones) by former Havas Global CEO David Jones, with a simple mission to help brands do their marketing better, faster and cheaper using technology. It was renamed The Brandtech Group in January 2022.
Today it generates more than $1BN in revenue and is the largest global digital content partner for many of the world's biggest brands and companies, often using its unique in-housing model. It works with eight of the world's top 10 global advertisers and 49 of the world's top 100. Clients include Banco Itaú, Danone, Google, Intuit, LVMH, Microsoft, Morgan-Stanley, Netflix, Reckitt, Renault-Nissan, PayPal, TikTok, Uber and Unilever.
In addition, the Group invests in cutting-edge technology companies relevant to and with potential in marketing and has been investing in the AI, AR and the metaverse space since 2015. The Group was the first external investor in Niantic (creators of Pokémon Go), alongside investments in AI Foundation (AI digital twins, personalized AI, and deep-fake detection 2018), Automat (AI chatbots, 2016), Elsy (AI media planning, 2016), Jivox (dynamic automated content, 2016), Crossing Minds (deep learning AI recommendation system, 2017), CreativeX (AI-driven content optimization, 2022), Zappar (AR, 2017), Provenance (blockchain, 2022) and GGP (the world's largest gaming fund).

