← Back to work

Case study · AI Evaluation & Product Thinking

Urdu Creative Bench: judging what global AI benchmarks structurally cannot.

Global benchmarks can tell us whether a model performs well. They cannot always tell us whether its outputs are usable, culturally appropriate, or trustworthy in real-world contexts. Through native human evaluation and data-driven quality systems, this project explores one question. Two ways of answering one question: what does "good" mean, and how do you measure it?

Hero · the Urdu Creative Bench leaderboard or product-concept cover

Role

Product & eval designer (solo)

Context

AI Evaluation & Product Thinking

Timeline

5 days

Method

3 models · 18 images · 8 native raters

Overview

The challenge

This project began as an open-ended product thinking exercise with no fixed steps or scoring criteria.Part 1 asked me to design and run one useful text-to-image evaluation for India. Part 2 flipped the lens: judge the quality of human work using data signals instead of human raters. Both come down to the same skill, so I treated the whole thing as one product question.

What I delivered

3 modelsGPT Image 1, Gemini 2.5,and Gemini 3.1, judged blind
8 native raters144 ratings, 48 forced picks
51,783 tasksanalysed for the transcription quality system

Supporting materials: workbooks, ratings & proofs ·

My role

Sole author. I chose the evaluation, wrote the prompts, generated and captured every image, recruited and briefed raters, built the rating instrument, analysed results, designed the product concept and the quality system, and did the full Part 2 data analysis and system design.

Why it's here

It shows the part of design that isn't screens: defining "good," designing how something gets measured, and turning messy evidence into a system a team can act on.

The eval & why it matters

What I chose to test

Can text to image models produce commercially usable Urdu creatives that Kashmiri small businesses currently pay local designers to make: signboards, festival greetings, posters, and wedding cards? Each of six prompts is a real job someone in Srinagar pays for today.

Urdu is a stress test by construction: right-to-left, cursive Nastaliq where letters change shape by position, and users type it in Roman transliteration ("Eid Mubarak") because that is how real people search. 50M+ Indians read Urdu, yet it ranks worst of 14 languages in Josh Talks' own Voice of India speech benchmark. If a model gets Urdu right, the hard case is covered.

Why this evaluation matters

  • Revenue: design tools, WhatsApp Business marketing and print shops are paying use cases no model could serve unsupervised for Urdu users until very recently.
  • Invisible risk: the most dangerous output is not an ugly image, it is a beautiful image with wrong text: shippable by a non-reader, embarrassing to a customer. Only native-reader evaluation catches it.
  • Progress meter: labs need per-release, per-script measurement. This eval produced exactly that number (1.19 to 4.94).

How the evaluation works

  • Fair setup: six prompts, one use case, identical text to every model, each ending with the same "Square 1:1 format" line. Fresh chat per prompt, first image only, no rerolls or cherry-picking.
  • Exact-model discipline: GPT Image 1, Gemini 2.5 Flash Image, and Gemini 3.1 Flash Image Preview were captured with exact model IDs and verified through proof screenshots. No model substitutions or rerolls were allowed.
  • Blind judging: images labelled only A / B / C, mapping never revealed, simplified task lines instead of full prompts to avoid spec-checking bias.
  • The instrument: Each image was rated on four criteria: Urdu text accuracy, legibility, cultural fit, and visual appeal. Participants also made one forced choice selection per set and could optionally transcribe the Urdu text, providing verbatim evidence of what they actually read.
  • Participants:Nine participants responded. One response was excluded, leaving eight consenting native readers aged 23 to 32. One non reader was intentionally retained as a customer control to test Finding 2.
Grid of 18 generated images: three models across six Urdu prompts
Grid of 18 generated images: three models across six Urdu prompts
All 18 outputs, same prompt across models. Beautiful and unreadable sits right next to correct.

All eighteen generations used identical prompts across models and were captured with their exact model IDs.

Supporting materials: Generation screenshots & Proofs · Google Form export

Results & findings

Rank Model Text Legibility Cultural fit Overall / Preferred
1 (C) Gemini 3.1 Flash Image Preview 4.94 4.98 4.83 4.86 / 46 of 48
2 (A) GPT Image 1 2.35 2.94 2.56 2.51 / 2 of 48
3 (B) Gemini 2.5 Flash Image 1.19 1.58 2.10 1.66 / 0 of 48

All 144 ratings and 48 picks live in the ratings workbook (Participant_Ratings, Leaderboard), with participant information removed for privacy.

Finding 01

Google fixed Urdu in one generation

On identical prompts, Gemini 2.5 scored 1.19/5 on text accuracy (fluent readers alone: 1.13); Gemini 3.1 scored 4.94. Raters read the same Eid headline as "inaayirik" (A), "andainaaki" (B), and "Eid Mubarak, Kashmir Crafts" (C).

Finding 02

"Looks right, reads wrong" is the dangerous failure

Gemini 2.5 produced the most photorealistic scenes yet wrote decorative pseudo-Urdu or escaped to English, including a fully hallucinated English wedding invitation. A non-reading owner could ship the pretty fakes; a reading customer would mock them. This is exactly what automated metrics and non-native eyes miss.

Finding 03

Even the winner isn't deployable unsupervised

Gemini 3.1 spelled almost everything correctly but still over generated: uninvited English taglines, an unrequested Bismillah header, invented names, and even a template leak printed into the final artwork. A spelling-only metric would have scored it flawless.

Finding 04

Cultural fit needs local judges

Only native readers could tell whether the outputs felt authentically Kashmiri. They recognised local details such as the vertical tandoor used by kandurs and papier mâché motifs, while also catching failures that text accuracy alone missed. Location phrases such as “for a shop window” caused models to paint the scene around the poster rather than the poster itself, breaking the intended 1:1 artifact.

Native readers overwhelmingly preferred Gemini 3.1, selecting it 46 out of 48 times.
Native readers overwhelmingly preferred Gemini 3.1, selecting it 46 out of 48 times.

From evaluation to product concept

The evaluation was designed as more than a one time study. I translated it into a reusable product concept that a lab could run repeatedly across models, releases, and Indian scripts.

  • Collect: the blind rating UI, whose typed-reading field doubles as evidence collection.
  • Rank: a leaderboard with criterion filters, including a "fluent readers only" view.
  • Diagnose: a quality-flag feed that turns rater comments into a fix-list for labs.
  • Scale:from an 8 rater pilot to a standing benchmark for Indian scripts. Josh Talks AI already operates native language evaluations at much larger scales, making this approach reusable across model releases and languages.
JoshTalks-Dashboardimage1
JoshTalks-Dashboardimage2
JoshTalks-Dashboardimage3
JoshTalks-Dashboardimage4
From pilot to product: native judging, scalable voting, hybrid scoring, and release tracking for Indian scripts.

Question 2: a data-driven quality system

Part 2 asked me to catch low-quality transcribers from data alone. Transcribers are paid per accurate hour of audio; their real cost is time, and listening is the slow part, so all bad work reduces to one behaviour: submitting without truly hearing the audio. That behaviour leaves three fingerprints in the 51,783-task dataset.

The structural fact that shapes everything: 48,875 unique users, of whom 48,865 appear exactly once. Detection has to work on a single task first; per-user history rules apply only to the tiny minority with a track record.

Three warning signs

  • Impossible listening time: time_taken < duration means the clip physically cannot have been heard. Fires on 5,985 tasks (11.6%); a fifth used under half the audio's length.
  • Blindly accepting the AI draft: unedited text alone is weak evidence, but unedited and faster than the audio strongly suggests no listening and no human value added (3,568 tasks).
  • Effort-free submissions: 10 seconds of continuous Hindi holds 25 to 35 words, so a 5-second clip with under 5 characters is speech dismissed, not transcribed.
Signal Tasks Share What it means
Impossible listening time 5,985 11.6% Submitted faster than the audio duration itself.
Impossible speed + unedited 3,568 6.9% Accepted the AI draft without edits and submitted faster than the audio could be heard.
Effort-free submissions 1,091 2.1% Long audio clips (>5 seconds) submitted with fewer than 5 characters.

Supporting documentation: Detailed analyses for all behavioral signals are available here.

The blocking system: behaviour flags, verified evidence blocks

Because these accounts get blocked, the logic has to be one I'm sure of. Speed and edit stats are circumstantial, so they only flag; conviction comes from gold-standard clips: audio with a known-correct transcription injected invisibly into a suspect's queue. No honest worker fails gold repeatedly; every cheater does.

  • Tier 0: listen ratio < 1 AND unedited, or over-5s audio with under-5 characters, then withhold pay, requeue, 1 strike.
  • Tier 1: 3 strikes or ≥20% of the last 20 tasks flagged, then inject 5 unannounced gold clips scored by word error rate.
  • Tier 2: gold < 70% gives a warning and a hold; ≥70% clears the flags. This exit ramp is the fairness guarantee for fast-but-accurate workers.
  • Tier 3: gold failed twice after a warning, then block and reverse unpaid flagged earnings.
Behavior flags -></span> gold clips convict -> honest fast workers always have an exit ramp
Behavior flags -> gold clips convict -> honest fast workers always have an exit ramp
Design principle: the system judges tasks before people, because 99.98% of workers here have exactly one task to their name. A block means only one thing: you repeatedly transcribed known audio incorrectly after a warning. The decision is objective, evidenced, and appealable.

Reflection & learnings

  • "Best" depends on the criterion. The most photorealistic model was the least usable. An overall ranking would have hidden that tradeoff.
  • Humans catch what metrics cannot. Template leaks,invented content, and uninvited English were invisible to spelling checks but obvious to native readers.
  • Design the safeguards, not just the detection. When a quality system can affect someone’s income, the verification process is as important as the detection logic.
  • Local expertise is a competitive advantage. Because my raters use Urdu in everyday commerce, I could evaluate real business use cases that broad multilingual benchmarks often miss.

This project changed how I think about AI quality: the hardest part isn’t generating outputs. It’s defining quality, measuring it fairly, and designing systems people can trust.