This case study is private

It’s under review with Josh Talks. Enter the password to view it.

← Back to work

Case study · AI Evaluation & Product Thinking

Urdu Creative Bench: judging what global AI benchmarks structurally cannot.

A hiring challenge for an AI data lab, taken as a product problem: I designed and ran a native-judged text-to-image evaluation for Urdu commercial creatives, then designed a data-driven quality system for human transcription work. Two ways of answering one question: what does "good" mean, and how do you measure it?

Hero · the Urdu Creative Bench leaderboard or product-concept cover · drop image

Role

Product & eval designer (solo)

Context

Josh Talks AI product challenge

Timeline

5 days

Method

3 models · 18 images · 8 native raters

Overview

The challenge

Josh Talks AI works on evaluations and gold-standard datasets for AI in Indian contexts. Their hiring challenge was deliberately under-defined: no fixed steps, no fixed scoring. Part 1 asked me to design and run one useful text-to-image evaluation for India. Part 2 flipped the lens: judge the quality of human work using data signals instead of human raters. Both come down to the same skill, so I treated the whole thing as one product question.

What I delivered

3 modelsGPT Image 1, Gemini 2.5, Gemini 3.1, judged blind
8 native raters144 ratings, 48 forced picks, consented
51,783 tasksanalysed for the transcription quality system

Attached with this study: 2-minute Q1 walkthrough (video) · product-concept slide deck · workbooks, ratings & proofs.

My role

Sole author. I chose the evaluation, wrote the prompts, generated and captured every image, recruited and briefed raters, built the rating instrument, analysed results, designed the product concept, and did the full Part 2 data analysis and system design.

Why it's here

It shows the part of design that isn't screens: defining "good," designing how something gets measured, and turning messy evidence into a system a team can act on.

The eval & why it matters

What I chose to test

Can text-to-image models produce the Urdu-text commercial creatives that Kashmiri small businesses currently pay local designers to make: signboards, festival greetings, posters, wedding cards? Each of six prompts is a real job someone in Srinagar pays for today.

Urdu is a stress test by construction: right-to-left, cursive Nastaliq where letters change shape by position, and users type it in Roman transliteration ("Eid Mubarak") because that is how real people search. 50M+ Indians read Urdu, yet it ranks worst of 14 languages in Josh Talks' own Voice of India speech benchmark. If a model gets Urdu right, the hard case is covered.

Why a lab building for India should care

  • Revenue: design tools, WhatsApp Business marketing and print shops are paying use cases no model could serve unsupervised for Urdu users until very recently.
  • Invisible risk: the most dangerous output is not an ugly image, it is a beautiful image with wrong text: shippable by a non-reader, embarrassing to a customer. Only native-reader evaluation catches it.
  • Progress meter: labs need per-release, per-script measurement. This eval produced exactly that number (1.19 to 4.94).

How the evaluation works

  • Fair setup: six prompts, one use case, identical text to every model, each ending with the same "Square 1:1 format" line. Fresh chat per prompt, first image only, no rerolls or cherry-picking.
  • Exact-model discipline: GPT Image 1 and Gemini 2.5 Flash Image via LMArena, Gemini 3.1 Flash Image Preview via Adobe Firefly, with the model ID captured in every proof screenshot. Model access became a finding in itself: models were deprecated and swapped mid-assignment, so exact-name verification mattered.
  • Blind judging: images labelled only A / B / C, mapping never revealed, simplified task lines instead of full prompts to avoid spec-checking bias.
  • The instrument: four 1 to 5 criteria per image (Urdu text accuracy, legibility, cultural fit, visual appeal), one forced "if this were your shop, which would you use?" pick per set, and an optional "type what the Urdu says" field that forces genuine reading and collects verbatim evidence.
  • Participants: 9 responded, 1 excluded for junk identity fields, leaving 8 valid consenting raters aged 23 to 32: 5 fluent readers, 2 slow, and 1 non-reader kept as a "customer" control whose scores test Finding 2.

To view the raw participant responses (the form was shared on WhatsApp), see the Google Form export →

The blind rating form · A / B / C with the typed-reading field · drop image
The rating instrument: four criteria, a forced pick, and a "type what it says" field that doubles as evidence.

Results & findings

Rank Model Text Legibility Cultural fit Overall / Preferred
1 (C) Gemini 3.1 Flash Image Preview 4.94 4.98 4.83 4.86 / 46 of 48
2 (A) GPT Image 1 2.35 2.94 2.56 2.51 / 2 of 48
3 (B) Gemini 2.5 Flash Image 1.19 1.58 2.10 1.66 / 0 of 48

All 144 ratings and 48 picks live in the ratings workbook (Participant_Ratings, Leaderboard), with names, emails, ages and consent captured as mandatory fields.

Finding 01

Google fixed Urdu in one generation

On identical prompts, Gemini 2.5 scored 1.19/5 on text accuracy (fluent readers alone: 1.13); Gemini 3.1 scored 4.94. Raters read the same Eid headline as "inaayirik" (A), "andainaaki" (B), and "Eid Mubarak, Kashmir Crafts" (C).

Finding 02

"Looks right, reads wrong" is the dangerous failure

Gemini 2.5 produced the most photorealistic scenes yet wrote decorative pseudo-Urdu or escaped to English, including a fully hallucinated English wedding invitation. A non-reading owner could ship the pretty fakes; a reading customer would mock them. This is exactly what automated metrics and non-native eyes miss.

Finding 03

Even the winner isn't deployable unsupervised

Gemini 3.1 spelled essentially everything correctly but over-generates: uninvited English taglines, an unrequested Bismillah header, invented names, and a literal template leak printed into final art: "Visit us at: [Shop Name & Address]". A spelling-only metric would have scored it flawless.

Finding 04

Cultural fit needs local judges

Only Gemini 3.1 showed the vertical tandoor a Kashmiri kandur uses, papier-mâché-style paisley, and bazaar scenes raters recognised as Srinagar. A quirk also surfaced: location phrases ("for a shop window") made models paint the scene around the poster, breaking the 1:1 artifact.

The 18 generated images, gridded A / B / C across the 6 prompts · drop image
All 18 outputs, same prompt across models. Beautiful and unreadable sits right next to correct.

Every generation was captured with the exact model ID on screen. View all generation screenshots & proof files →

Product concept: Urdu Creative Bench

I designed the eval as a product in a five-part concept, so it reads as something a lab could actually run, not a one-off spreadsheet:

  • Frame: a standing, per-script benchmark for commercial creative generation.
  • Collect: the blind rating UI, whose typed-reading field doubles as evidence collection.
  • Rank: a leaderboard with criterion filters, including a "fluent readers only" view.
  • Diagnose: a quality-flag feed that turns rater comments into a fix-list for labs.
  • Scale: from 8 raters to a per-script benchmark at Josh Talks' proven size (Bulbul V3 used 50 to 70 native annotators and ~2,000 votes per language).
Collect · the rating UI · drop image
Rank · the filterable leaderboard · drop image
Diagnose · comments-into-fixes feed · drop image
Scale · standing benchmark screen · drop image
The five-part concept: rating UI, filterable leaderboard, and a diagnose feed that turns comments into fixes.

View the product-concept slide deck →

Question 2: a data-driven quality system

Part 2 asked me to catch low-quality transcribers from data alone. Transcribers are paid per accurate hour of audio; their real cost is time, and listening is the slow part, so all bad work reduces to one behaviour: submitting without truly hearing the audio. That behaviour leaves three fingerprints in the 51,783-task dataset.

The structural fact that shapes everything: 48,875 unique users, of whom 48,865 appear exactly once. Detection has to work on a single task first; per-user history rules apply only to the tiny minority with a track record.

Three warning signs

  • Impossible listening time: time_taken < duration means the clip physically cannot have been heard. Fires on 5,985 tasks (11.6%); a fifth used under half the audio's length.
  • Blindly accepting the AI draft: unedited alone is weak, but unedited and faster-than-audio together means no listening and no value added at human wages. This pair is the sharpest single-task signal (3,568 tasks).
  • Effort-free submissions: 10 seconds of continuous Hindi holds 25 to 35 words, so a 5-second clip with under 5 characters is speech dismissed, not transcribed.

For each sign I documented what it tells us, how to measure it, red-flag thresholds, and the honest false alarms (autoplay, tiny silent clips, a skilled worker verifying a correct draft). I also caught a data-dictionary error: segment_character_per_second is documented as chars ÷ time_taken but actually equals chars ÷ duration across all rows.

The blocking system: behaviour flags, only verified bad work blocks

Because these accounts get blocked, the logic has to be one I'm sure of. Speed and edit stats are circumstantial, so they only flag; conviction comes from gold-standard clips: audio with a known-correct transcription injected invisibly into a suspect's queue. No honest worker fails gold repeatedly; every cheater does.

  • Tier 0: listen ratio < 1 AND unedited, or over-5s audio with under-5 characters, then withhold pay, requeue, 1 strike.
  • Tier 1: 3 strikes or ≥20% of the last 20 tasks flagged, then inject 5 unannounced gold clips scored by word error rate.
  • Tier 2: gold < 70% gives a warning and a hold; ≥70% clears the flags. This exit ramp is the fairness guarantee for fast-but-accurate workers.
  • Tier 3: gold failed twice after a warning, then block and reverse unpaid flagged earnings.
Design principle: the system judges tasks before people, because 99.98% of workers here have exactly one task to their name. A block always means "you transcribed known audio wrong, repeatedly, after a warning": objective, evidenced, and appealable.
Warning-sign charts + the tiered blocking flow · drop image
Three behavioural fingerprints feeding a tiered system where only gold-clip failure blocks an account.

Reflection & learnings

  • "Best" depends on the criterion. The most photorealistic model was the least usable; a single score would have hidden that.
  • Humans catch what metrics can't. Template leaks and uninvited English never show up in a spelling check, only in native eyes.
  • Design the safeguards, not just the signal. When a metric can block someone's income, the fairness exit ramp is as much a part of the design as the detection.
  • Proximity is an advantage. My rater pool reads Urdu and my city's commerce runs on it, so I could evaluate something global benchmarks structurally cannot.