Talent.com
Canva
Principal Research Scientist EvaluationsCanva • Sydney, New South Wales, Australia
Principal Research Scientist Evaluations

Principal Research Scientist Evaluations

Canva • Sydney, New South Wales, Australia
6 days ago
Job description

About the role

Canvas generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack and it gates everything else. If we cannot measure design quality reliably we cannot train against it we cannot tell a real improvement from noise and we cannot decide what ships.

We are looking for a Principal Research Scientist who defines what evaluation needs to become as the space gets harder rather than running the playbook we already have. You will own how Canva evaluates generative quality across the whole of Canva Research including problems we have not framed yet: new modalities evaluation that reflects real differences between content types user segments and markets and a much tighter link between what our metrics say and what users and the business actually experience.

This is a Canva-wide craft leadership role setting direction across our research groups in Australia Europe the US and China. You will be the person others come to when the numbers and the eyes disagree.

What youll own

The evaluation strategy for Canva Research. Define what great evaluation looks like across design image video audio and agentic workflows and set the long-term direction for how Canva measures generative quality as the space evolves. You will shape the principles teams use to trade off evaluation compute human data spend and signal fidelity and you will defend them.

The science of measurement itself. Auto-raters and MLLM judges are only as good as their correlation with the thing they proxy for. You will treat that correlation as a research problem: validating metrics against human preference and downstream product outcomes quantifying judge bias and catching benchmark saturation and contamination before they quietly stop telling us anything.

The link between evaluation and what actually matters. Evaluation should reflect user experience and product impact not just perform well in isolation. You own closing that gap including the fact that a good evaluator is not automatically a good reward model for RL. That extends to the full experience rather than the artefact alone editability included.

One standard across every region. Our teams in London Vienna San Francisco Sydney and China all need to know they are measuring the same thing. You will build the shared evaluation layer that makes results comparable and you will spot the collaboration opportunities nobody has been chartered to own yet. Taking that initiative is a core expectation of this role not a bonus.

Focus areas

Human preference at scale. Rubric design rater guidelines inter-rater reliability and calibration across markets and cultural contexts. Aesthetic judgement is not universal and our evaluation systems need to hold that honestly rather than average it away.

Learned quality models and reward signals. Reward modelling and preference learning that feed post-training (RLHF RLAIF) and inference-time selection and being explicit about where a good judge does not translate into a good reward model.

Vision-Language Models for quality understanding. Novel architectures and training approaches for models that understand what makes a design effective not just well-formed. These become the reward signal and feedback loop for our design generation models so their failure modes are our failure modes.

Agentic and automated evaluation. MLLM-as-a-judge systems model-based grading and the infrastructure to run hundreds of evaluations against live training checkpoints without the results becoming noise.

Multimodal and segment-specific evaluation. Extending rigour into photo AI and video as they mature and start moving faster than traditional user research and marketing testing can cover with evaluation that can be sliced meaningfully by doc type user group and locale.

Evaluation in production. Offline evaluation suites and online monitoring with regression detection that makes a quality drop impossible to miss before it reaches users.

External credibility. Benchmarking Canvas models against the frontier and building evaluation work others in the industry look to. Publication and public reporting where it serves the mission.

Primary responsibilities

  • Establish the evaluation gates that inform launch decisions and be accountable for the judgement calls when the signal is ambiguous

  • Diagnose anomalous evaluation results during production training runs separate model regressions from infrastructure artefacts and communicate the answer clearly and quickly

  • Partner with Design Generation Foundation Models and Agents teams so evaluation shapes their training and inference rather than reporting on it after the fact

  • Work directly with designers creators and product teams to turn subjective creative judgement into measurable criteria

  • Mentor senior research scientists and engineers and raise the bar for evaluation craft across the group

  • Represent Canvas evaluation vision and practice to senior leadership and where valuable to the broader industry community

Youre probably a match if you have

  • A track record of defining evaluation frameworks and standards from the ground up rather than operating inside someone elses ideally at an organisation pushing the frontier of GenAI evaluation and measurement systems that changed how a team made decisions rather than papers about metrics

  • Experience linking evaluation metrics to downstream business or user outcomes and diagnosing why they diverge

  • Experience turning subjective human judgement into reliable objective evaluation signal through rubric design human data pipelines and model training

  • Strong grounding in multimodal generative models (diffusion transformers VLMs and MLLMs) and their architectures deep enough to know where evaluation will break

  • Experience with reward modelling preference learning or alignment methods involving human feedback

  • Experience setting technical direction across multiple teams or a wide specialty area in a globally distributed organisation with the instinct to find the gap nobody owns and close it without waiting for a mandate

Nice to have

  • Research background in human perception psychophysics aesthetics or HCI

  • Experience evaluating for harm bias and safety alongside quality

  • Publication record in evaluation alignment or generative modelling

  • Background or genuine interest in visual arts and graphic design


Additional Information :

Dont tick all the boxes Dont worry about that - nobody does!

Wed still love to hear from you! At Canva we know that great engineers come from a variety of backgrounds and we value passion curiosity and a willingness to learn just as much as specific experience. If youre excited about this role but dont tick every box we encourage you to apply you might a great fit in ways you didnt expect!

Whats in it for you

Achieving our crazy big goals motivates us to work hard - and we do - but youll experience lots of moments of magic connectivity and fun woven throughout life at Canva too. We also offer a stack of benefits to set you up for every success in and outside of work.

Heres a taste of whats on offer:

  • Equity packages - we want our success to be yours too
  • Inclusive parental leave policy that supports all parents & carers
  • An annual Vibe & Thrive allowance to support your wellbeing social connection office setup & more
  • Flexible leave options that empower you to be a force for good take time to recharge and supports you personally

Check out for more info.

Other stuff to know

We make hiring decisions based on your experience skills and passion as well as how you can enhance Canva and our culture. When you apply please tell us the pronouns you use and any reasonable adjustments you may need during the interview process.

Please note that interviews are conducted virtually.


Remote Work :

No


Employment Type :

Full-time


Experience: years
Vacancy: 1

Create a job alert for this search

Principal Research Scientist Evaluations • Sydney, New South Wales, Australia

Similar jobs

Market Research Executive

Albion Rye AssociatesSydney, AU

This company helps advertisers, planners, and publishers.End-to-end ownership of projects: from.Flexible hybrid working, with access to the Sydney office and a.Flexible / hybrid working & home offi... Show more

 • Promoted

Postdoctoral Research Associate - Neurodegeneration

University of SydneySydney, AU

Postdoctoral Research Associate - Neurodegeneration.Be among the first 25 applicants.Postdoctoral Research Associate - Neurodegeneration.An exciting opportunity to join the Snow Vision Accelerator.... Show more

 • Promoted

Principal AI Engineer

MYOBSydney, NSW, AU

We’re a leading business management solution with a core purpose: helping more businesses in Australia and New Zealand start, survive and succeed.At MYOB, we believe what’s good for one business is... Show more

Research Engineer

IMC B.V.Sydney, AU

The Research Engineering team develops and deploys the pipelines connecting a research idea all the way to its rollout in production.We exist to ensure that the entire pipeline describing our resea... Show more

 • Promoted

Laboratory Scientist

Australian Red Cross LifebloodSydney, AU

Australian Red Cross Lifeblood.Discover life-giving possibilities.Lifeblood is an organisation focused on life giving donations and life changing outcomes, with a commitment to helping you build a ... Show more

 • Promoted

Clinical Research Associate

IQVIASydney, AU

Look no further than IQVIA, a global CRO with outstanding reputation.Our CRA teams have current openings across various locations in Australia.Visa sponsorship for candidates that meet the requirem... Show more

 • Promoted

Principal Environmental Scientist - Contaminated Sites

ALRA RecruitmentSydney, AU

Principal Environmental Scientist - Contaminated Sites.Principal Environmental Scientist in Contaminated sites to join the thriving Sydney team supporting the NSW Manager as a key strategic senior ... Show more

 • Promoted

Senior Research Director – Insights & Advisory

Resources GroupSydney, AU

Senior Research Director – Insights & Advisory.You’ll be the person who steps into the market as one of the agency’s senior commercial voices.Spot opportunities before others do.Write compelling pr... Show more

 • Promoted

Senior Manager, Clinical Research ANZ — Strategy & Studies

TRESP Recruitment - Medtech and Healthtech ExpertsSydney, AU

TRESP Recruitment is seeking a strategic leader for a pivotal role in clinical research initiatives across multiple high-impact MedTech therapy areas.You will partner with regional and global stake... Show more

 • Promoted

2026 Cushman and Wakefield Research Analyst in Australia

Jamnet TeamSydney, AU

Cushman and Wakefield Research Analyst (Sydney, NSW).Cushman & Wakefield has opened applications for a Paid Research Analyst role based in Sydney, Australia.This position serves as an active gatewa... Show more

 • Promoted

Principal AI Engineer

Myob Group LimitedSydney, AU

We’re a leading business management solution with a core purpose: helping more businesses in Australia and New Zealand start, survive and succeed.At MYOB, we believe what’s good for one business is... Show more

 • Promoted

Junior Market Researcher — Insights & Analysis (Remote)

IpsosSydney, AU
Remote

Ipsos is hiring a Research Executive to begin a career in market research.You will contribute to multiple service lines, support end-to-end project delivery, and work with industry experts in a col... Show more

 • Promoted

Postdoctoral Research Associate

University of SydneySydney, AU

Full time, 2 year fixed term position (with the possibility of extension).Located on the Camperdown Campus, University of Sydney.Base Salary Academic Level A $117,936 - $125,896 + 17% superannuatio... Show more

 • Promoted

Research Analyst

JLLSydney, AU

Do you want to be part of Australia’s largest and most respected dedicated property research team? Do you want to produce market‑leading analysis on the New South Wales property sector for external... Show more

 • Promoted

Senior Researcher – Strategy-led Insights Consultancy

Resources GroupSydney, AU

Senior Researcher – Strategy-led Insights Consultancy.Senior Researcher – Strategy-led Insights Consultancy.Be among the first 25 applicants.This range is provided by Resources Group.Your actual pa... Show more

 • Promoted

Lead Scientist, Single-Cell Genomics & Autoimmune Research

The Garvan Institute of Medical ResearchSydney, AU

The Garvan Institute of Medical Research in Sydney seeks a senior researcher to lead projects using single-cell RNA and whole-genome sequencing to study somatic mutations in autoimmune diseases.You... Show more

 • Promoted

Principal Machine Learning Engineer

Mantech RecruitmentSydney, AU

Get AI-powered advice on this job and more exclusive features.Direct message the job poster from Mantech Recruitment.Principal Consultant | Below Average Golfer.Principal Machine Learning Engineer ... Show more

 • Promoted

Director, Research Services: Lead High-Impact Research

The University of Notre Dame AustraliaSydney, AU

A private Catholic university in Australia is seeking a Director of Research Services to lead their research initiative, ensuring effective services for grants and training.The ideal candidate will... Show more

 • Promoted

Senior Research Assistant

Victor Chang Cardiac Research InstituteSydney, AU

Victor Chang Cardiac Research Institute.Talent Acquisition at Victor Chang Cardiac Research Institute.Employer: Victor Chang Cardiac Research Institute Limited.Location: Darlinghurst, Sydney, NSW.S... Show more

 • Promoted

Marketing Science Partner

MutinexSydney, AU

We're an early-stage B2B SaaS startup with a proven platform, big-name clients, and millions in revenue.We're not chasing unicorn status; we're building a sustainable, long-lasting business (think ... Show more