Case study · AI Evaluation & Product Thinking
Urdu Creative Bench: judging what global AI benchmarks structurally cannot.
A hiring challenge for an AI data lab, taken as a product problem: I designed and ran a native-judged text-to-image evaluation for Urdu commercial creatives, then designed a data-driven quality system for human transcription work. Two ways of answering one question: what does "good" mean, and how do you measure it?
Role
Product & eval designer (solo)
Context
Josh Talks AI product challenge
Timeline
5 days
Method
3 models · 18 images · 8 native raters
Overview
The challenge
Josh Talks AI works on evaluations and gold-standard datasets for AI in Indian contexts. Their hiring challenge was deliberately under-defined: no fixed steps, no fixed scoring. Part 1 asked me to design and run one useful text-to-image evaluation for India. Part 2 flipped the lens: judge the quality of human work using data signals instead of human raters. Both come down to the same skill, so I treated the whole thing as one product question.
What I delivered
Attached with this study: 2-minute Q1 walkthrough (video) · product-concept slide deck · workbooks, ratings & proofs.
My role
Sole author. I chose the evaluation, wrote the prompts, generated and captured every image, recruited and briefed raters, built the rating instrument, analysed results, designed the product concept, and did the full Part 2 data analysis and system design.
Why it's here
It shows the part of design that isn't screens: defining "good," designing how something gets measured, and turning messy evidence into a system a team can act on.
The eval & why it matters
What I chose to test
Can text-to-image models produce the Urdu-text commercial creatives that Kashmiri small businesses currently pay local designers to make: signboards, festival greetings, posters, wedding cards? Each of six prompts is a real job someone in Srinagar pays for today.
Why a lab building for India should care
- Revenue: design tools, WhatsApp Business marketing and print shops are paying use cases no model could serve unsupervised for Urdu users until very recently.
- Invisible risk: the most dangerous output is not an ugly image, it is a beautiful image with wrong text: shippable by a non-reader, embarrassing to a customer. Only native-reader evaluation catches it.
- Progress meter: labs need per-release, per-script measurement. This eval produced exactly that number (1.19 to 4.94).
How the evaluation works
- Fair setup: six prompts, one use case, identical text to every model, each ending with the same "Square 1:1 format" line. Fresh chat per prompt, first image only, no rerolls or cherry-picking.
- Exact-model discipline: GPT Image 1 and Gemini 2.5 Flash Image via LMArena, Gemini 3.1 Flash Image Preview via Adobe Firefly, with the model ID captured in every proof screenshot. Model access became a finding in itself: models were deprecated and swapped mid-assignment, so exact-name verification mattered.
- Blind judging: images labelled only A / B / C, mapping never revealed, simplified task lines instead of full prompts to avoid spec-checking bias.
- The instrument: four 1 to 5 criteria per image (Urdu text accuracy, legibility, cultural fit, visual appeal), one forced "if this were your shop, which would you use?" pick per set, and an optional "type what the Urdu says" field that forces genuine reading and collects verbatim evidence.
- Participants: 9 responded, 1 excluded for junk identity fields, leaving 8 valid consenting raters aged 23 to 32: 5 fluent readers, 2 slow, and 1 non-reader kept as a "customer" control whose scores test Finding 2.
To view the raw participant responses (the form was shared on WhatsApp), see the Google Form export →
Results & findings
| Rank | Model | Text | Legibility | Cultural fit | Overall / Preferred |
|---|---|---|---|---|---|
| 1 (C) | Gemini 3.1 Flash Image Preview | 4.94 | 4.98 | 4.83 | 4.86 / 46 of 48 |
| 2 (A) | GPT Image 1 | 2.35 | 2.94 | 2.56 | 2.51 / 2 of 48 |
| 3 (B) | Gemini 2.5 Flash Image | 1.19 | 1.58 | 2.10 | 1.66 / 0 of 48 |
All 144 ratings and 48 picks live in the ratings workbook (Participant_Ratings, Leaderboard), with names, emails, ages and consent captured as mandatory fields.
Google fixed Urdu in one generation
On identical prompts, Gemini 2.5 scored 1.19/5 on text accuracy (fluent readers alone: 1.13); Gemini 3.1 scored 4.94. Raters read the same Eid headline as "inaayirik" (A), "andainaaki" (B), and "Eid Mubarak, Kashmir Crafts" (C).
"Looks right, reads wrong" is the dangerous failure
Gemini 2.5 produced the most photorealistic scenes yet wrote decorative pseudo-Urdu or escaped to English, including a fully hallucinated English wedding invitation. A non-reading owner could ship the pretty fakes; a reading customer would mock them. This is exactly what automated metrics and non-native eyes miss.
Even the winner isn't deployable unsupervised
Gemini 3.1 spelled essentially everything correctly but over-generates: uninvited English taglines, an unrequested Bismillah header, invented names, and a literal template leak printed into final art: "Visit us at: [Shop Name & Address]". A spelling-only metric would have scored it flawless.
Cultural fit needs local judges
Only Gemini 3.1 showed the vertical tandoor a Kashmiri kandur uses, papier-mâché-style paisley, and bazaar scenes raters recognised as Srinagar. A quirk also surfaced: location phrases ("for a shop window") made models paint the scene around the poster, breaking the 1:1 artifact.
Every generation was captured with the exact model ID on screen. View all generation screenshots & proof files →
Product concept: Urdu Creative Bench
I designed the eval as a product in a five-part concept, so it reads as something a lab could actually run, not a one-off spreadsheet:
- Frame: a standing, per-script benchmark for commercial creative generation.
- Collect: the blind rating UI, whose typed-reading field doubles as evidence collection.
- Rank: a leaderboard with criterion filters, including a "fluent readers only" view.
- Diagnose: a quality-flag feed that turns rater comments into a fix-list for labs.
- Scale: from 8 raters to a per-script benchmark at Josh Talks' proven size (Bulbul V3 used 50 to 70 native annotators and ~2,000 votes per language).
View the product-concept slide deck →
Question 2: a data-driven quality system
Part 2 asked me to catch low-quality transcribers from data alone. Transcribers are paid per accurate hour of audio; their real cost is time, and listening is the slow part, so all bad work reduces to one behaviour: submitting without truly hearing the audio. That behaviour leaves three fingerprints in the 51,783-task dataset.
Three warning signs
- Impossible listening time:
time_taken < durationmeans the clip physically cannot have been heard. Fires on 5,985 tasks (11.6%); a fifth used under half the audio's length. - Blindly accepting the AI draft: unedited alone is weak, but unedited and faster-than-audio together means no listening and no value added at human wages. This pair is the sharpest single-task signal (3,568 tasks).
- Effort-free submissions: 10 seconds of continuous Hindi holds 25 to 35 words, so a 5-second clip with under 5 characters is speech dismissed, not transcribed.
For each sign I documented what it tells us, how to measure it, red-flag thresholds, and the honest false alarms (autoplay, tiny silent clips, a skilled worker verifying a correct draft). I also caught a data-dictionary error: segment_character_per_second is documented as chars ÷ time_taken but actually equals chars ÷ duration across all rows.
The blocking system: behaviour flags, only verified bad work blocks
Because these accounts get blocked, the logic has to be one I'm sure of. Speed and edit stats are circumstantial, so they only flag; conviction comes from gold-standard clips: audio with a known-correct transcription injected invisibly into a suspect's queue. No honest worker fails gold repeatedly; every cheater does.
- Tier 0: listen ratio < 1 AND unedited, or over-5s audio with under-5 characters, then withhold pay, requeue, 1 strike.
- Tier 1: 3 strikes or ≥20% of the last 20 tasks flagged, then inject 5 unannounced gold clips scored by word error rate.
- Tier 2: gold < 70% gives a warning and a hold; ≥70% clears the flags. This exit ramp is the fairness guarantee for fast-but-accurate workers.
- Tier 3: gold failed twice after a warning, then block and reverse unpaid flagged earnings.
Reflection & learnings
- "Best" depends on the criterion. The most photorealistic model was the least usable; a single score would have hidden that.
- Humans catch what metrics can't. Template leaks and uninvited English never show up in a spelling check, only in native eyes.
- Design the safeguards, not just the signal. When a metric can block someone's income, the fairness exit ramp is as much a part of the design as the detection.
- Proximity is an advantage. My rater pool reads Urdu and my city's commerce runs on it, so I could evaluate something global benchmarks structurally cannot.