368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$310k – $420k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Principal
Employment
Full-Time
Overview
Company
Impact
Profile match
Canva is an Australian multinational graphic design and online publishing platform founded in 2013 by Melanie Perkins, Cliff Obrecht, and Cameron Adams. Designed to make visual communication accessible to non-professionals, Canva features an intuitive drag-and-drop editor alongside a vast library of customizable templates, stock media, fonts, and graphics.

Join the team redefining how the world experiences design.

Hey, g'day, mabuhay, kia ora,你好, hallo, vítejte!

Thanks for stopping by. We know job hunting can be a little time consuming and you're probably keen to find out what's on offer, so we'll get straight to the point.

Where and how you can work

Our head office is in Sydney, Australia, but San Francisco is now home to our US operations. The role is listed as hybrid, meaning we are incredibly flexible and empower you to work where you prefer - whether that's at home or at the office.

About the role

Canva's generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack, and it gates everything else. If we cannot measure design quality reliably, we cannot train against it, we cannot tell a real improvement from noise, and we cannot decide what ships.

We are looking for a Principal Research Scientist who defines what evaluation needs to become as the space gets harder, rather than running the playbook we already have. You will own how Canva evaluates generative quality across the whole of Canva Research, including problems we have not framed yet: new modalities, evaluation that reflects real differences between content types, user segments and markets, and a much tighter link between what our metrics say and what users and the business actually experience.

This is a Canva-wide craft leadership role, setting direction across our research groups in Australia, Europe, the US and China. You will be the person others come to when the numbers and the eyes disagree.

What you'll own

The evaluation strategy for Canva Research. Define what great evaluation looks like across design, image, video, audio and agentic workflows, and set the long-term direction for how Canva measures generative quality as the space evolves. You will shape the principles teams use to trade off evaluation compute, human data spend and signal fidelity, and you will defend them.

The science of measurement itself. Auto-raters and MLLM judges are only as good as their correlation with the thing they proxy for. You will treat that correlation as a research problem: validating metrics against human preference and downstream product outcomes, quantifying judge bias, and catching benchmark saturation and contamination before they quietly stop telling us anything.

The link between evaluation and what actually matters. Evaluation should reflect user experience and product impact, not just perform well in isolation. You own closing that gap, including the fact that a good evaluator is not automatically a good reward model for RL. That extends to the full experience rather than the artefact alone, editability included.

One standard across every region. Our teams in London, Vienna, San Francisco, Sydney and China all need to know they are measuring the same thing. You will build the shared evaluation layer that makes results comparable, and you will spot the collaboration opportunities nobody has been chartered to own yet. Taking that initiative is a core expectation of this role, not a bonus.

Focus areas

Human preference at scale. Rubric design, rater guidelines, inter-rater reliability, and calibration across markets and cultural contexts. Aesthetic judgement is not universal, and our evaluation systems need to hold that honestly rather than average it away.

Learned quality models and reward signals. Reward modelling and preference learning that feed post-training (RLHF, RLAIF) and inference-time selection, and being explicit about where a good judge does not translate into a good reward model.

Vision-Language Models for quality understanding. Novel architectures and training approaches for models that understand what makes a design effective, not just well-formed. These become the reward signal and feedback loop for our design generation models, so their failure modes are our failure modes.

Agentic and automated evaluation. MLLM-as-a-judge systems, model-based grading, and the infrastructure to run hundreds of evaluations against live training checkpoints without the results becoming noise.

Multimodal and segment-specific evaluation. Extending rigour into photo AI and video as they mature and start moving faster than traditional user research and marketing testing can cover, with evaluation that can be sliced meaningfully by doc type, user group and locale.

Evaluation in production. Offline evaluation suites and online monitoring, with regression detection that makes a quality drop impossible to miss before it reaches users.

External credibility. Benchmarking Canva's models against the frontier, and building evaluation work others in the industry look to. Publication and public reporting where it serves the mission.

Primary responsibilities

  • Establish the evaluation gates that inform launch decisions, and be accountable for the judgement calls when the signal is ambiguous

  • Diagnose anomalous evaluation results during production training runs, separate model regressions from infrastructure artefacts, and communicate the answer clearly and quickly

  • Partner with Design Generation, Foundation Models and Agents teams so evaluation shapes their training and inference, rather than reporting on it after the fact

  • Work directly with designers, creators and product teams to turn subjective creative judgement into measurable criteria

  • Mentor senior research scientists and engineers, and raise the bar for evaluation craft across the group

  • Represent Canva's evaluation vision and practice to senior leadership and, where valuable, to the broader industry community

You're probably a match if you have

  • A track record of defining evaluation frameworks and standards from the ground up rather than operating inside someone else's, ideally at an organisation pushing the frontier of GenAI evaluation, and measurement systems that changed how a team made decisions rather than papers about metrics

  • Experience linking evaluation metrics to downstream business or user outcomes, and diagnosing why they diverge

  • Experience turning subjective human judgement into reliable, objective evaluation signal through rubric design, human data pipelines and model training

  • Strong grounding in multimodal generative models (diffusion, transformers, VLMs and MLLMs) and their architectures, deep enough to know where evaluation will break

  • Experience with reward modelling, preference learning, or alignment methods involving human feedback

  • Experience setting technical direction across multiple teams or a wide specialty area in a globally distributed organisation, with the instinct to find the gap nobody owns and close it without waiting for a mandate

Nice to have

  • Research background in human perception, psychophysics, aesthetics or HCI

  • Experience evaluating for harm, bias and safety alongside quality

  • Publication record in evaluation, alignment or generative modelling

  • Background or genuine interest in visual arts and graphic design

What's in it for you?

Achieving our crazy big goals motivates us to work hard - and we do - but you'll experience lots of moments of magic, connectivity and fun woven throughout life at Canva, too. We also offer a range of benefits to set you up for every success in and outside of work.

Here's a taste of what's on offer:

  • Equity packages - we want our success to be yours too
  • Health benefits plans to support you and your wellbeing
  • 401(k) retirement plan with company contribution
  • Inclusive parental leave policy that supports all parents & carers
  • An annual Vibe & Thrive allowance to support your wellbeing, social connection, office setup & more
  • Flexible leave options that empower you to be a force for good, take time to recharge and supports you personally

Check out lifeatcanva.com for more information.

Other stuff to know

We make hiring decisions based on your experience, skills, merit and business needs, in compliance with applicable local laws. We celebrate all types of skills and backgrounds at Canva so even if you don’t feel like your skills quite match what’s listed above - we still want to hear from you!

When you apply, please tell us the pronouns you use and any reasonable adjustments you may need during the interview process. Please note that interviews are conducted virtually.

At Canva, we value fairness, and we strive to provide competitive, market-informed compensation whilst ensuring internal equity within the team in each region. The target salary range for this position is $310,000 - $420,000. When calculating offers, we make salary decisions based on market data and candidates' skills and experience.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$20k – $46k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
AI/ML
AI Agents
Amazon SageMaker
Anomaly Detection
AWS Bedrock
Computer Vision
Embeddings
LLM
Multimodal AI
Prompt Engineering
RAG
Time Series Forecasting
DevOps
Amazon S3
AWS
AWS Lambda
CI/CD
Git
Vector
Analytics
Power BI
Tableau
Apply
$20k – $57k per year (Estimated) • In office • Full-Time • 2+ years exp • Chennai • Coimbatore
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
Function Calling
DevOps
Cortex
Prometheus
Apply
Data Engineer 2 hours ago
$22k – $55k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad
Databases
Databricks
AI/ML
AI Agents
Analytics
ETL/ELT
Apply
$63k – $137k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Madrid • Barcelona
AI/ML
AI Agents
Apply
LLM Model Developer 2 hours ago
$28k – $77k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad
AI/ML
AI Agents
Fine-tuning
LLM
Apply
$129k – $287k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • Sydney
AI/ML
LLM
Model Context Protocol
DevOps
Shift-Left
Cybersecurity
Shift-Left Security
Design
Canva
Apply
$179k – $285k per year • Equity • Remote/Hybrid • Full-Time • San Francisco
Python
SQL
Databases
Databricks
Google BigQuery
Snowflake
AI/ML
dbt
Design
Canva
Apply
$81k – $208k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • Sydney
Python
SQL
Analytics
A/B Testing
Design
Canva
Apply
$81k – $230k per year (Estimated) • Equity • In office • Full-Time • Sydney
SQL
Databases
Snowflake
Analytics
A/B Testing
Design
Canva
Apply
$99k – $249k per year (Estimated) • Equity • Remote • Full-Time • Sydney
Databases
ElasticSearch
OpenSearch
DevOps
Vector
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$185k – $385k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
Databases
MySQL
PostgreSQL
AI/ML
OpenAI
Frontend
React.js
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.