Case study · AI Evaluation & Product Thinking
Urdu Creative Bench: judging what global AI benchmarks structurally cannot.
Global benchmarks can tell us whether a model performs well. They cannot always tell us whether its outputs are usable, culturally appropriate, or trustworthy in real-world contexts. Through native human evaluation and data-driven quality systems, this project explores one question. Two ways of answering one question: what does "good" mean, and how do you measure it?
Role
Product & eval designer (solo)
Context
AI Evaluation & Product Thinking
Timeline
5 days
Method
3 models · 18 images · 8 native raters
Overview
The challenge
This project began as an open-ended product thinking exercise with no fixed steps or scoring criteria.Part 1 asked me to design and run one useful text-to-image evaluation for India. Part 2 flipped the lens: judge the quality of human work using data signals instead of human raters. Both come down to the same skill, so I treated the whole thing as one product question.
What I delivered
Supporting materials: workbooks, ratings & proofs ·
My role
Sole author. I chose the evaluation, wrote the prompts, generated and captured every image, recruited and briefed raters, built the rating instrument, analysed results, designed the product concept and the quality system, and did the full Part 2 data analysis and system design.
Why it's here
It shows the part of design that isn't screens: defining "good," designing how something gets measured, and turning messy evidence into a system a team can act on.
The eval & why it matters
What I chose to test
Can text to image models produce commercially usable Urdu creatives that Kashmiri small businesses currently pay local designers to make: signboards, festival greetings, posters, and wedding cards? Each of six prompts is a real job someone in Srinagar pays for today.
Why this evaluation matters
- Revenue: design tools, WhatsApp Business marketing and print shops are paying use cases no model could serve unsupervised for Urdu users until very recently.
- Invisible risk: the most dangerous output is not an ugly image, it is a beautiful image with wrong text: shippable by a non-reader, embarrassing to a customer. Only native-reader evaluation catches it.
- Progress meter: labs need per-release, per-script measurement. This eval produced exactly that number (1.19 to 4.94).
How the evaluation works
- Fair setup: six prompts, one use case, identical text to every model, each ending with the same "Square 1:1 format" line. Fresh chat per prompt, first image only, no rerolls or cherry-picking.
- Exact-model discipline: GPT Image 1, Gemini 2.5 Flash Image, and Gemini 3.1 Flash Image Preview were captured with exact model IDs and verified through proof screenshots. No model substitutions or rerolls were allowed.
- Blind judging: images labelled only A / B / C, mapping never revealed, simplified task lines instead of full prompts to avoid spec-checking bias.
- The instrument: Each image was rated on four criteria: Urdu text accuracy, legibility, cultural fit, and visual appeal. Participants also made one forced choice selection per set and could optionally transcribe the Urdu text, providing verbatim evidence of what they actually read.
- Participants:Nine participants responded. One response was excluded, leaving eight consenting native readers aged 23 to 32. One non reader was intentionally retained as a customer control to test Finding 2.
All eighteen generations used identical prompts across models and were captured with their exact model IDs.
Supporting materials: Generation screenshots & Proofs · Google Form export
Results & findings
| Rank | Model | Text | Legibility | Cultural fit | Overall / Preferred |
|---|---|---|---|---|---|
| 1 (C) | Gemini 3.1 Flash Image Preview | 4.94 | 4.98 | 4.83 | 4.86 / 46 of 48 |
| 2 (A) | GPT Image 1 | 2.35 | 2.94 | 2.56 | 2.51 / 2 of 48 |
| 3 (B) | Gemini 2.5 Flash Image | 1.19 | 1.58 | 2.10 | 1.66 / 0 of 48 |
All 144 ratings and 48 picks live in the ratings workbook (Participant_Ratings, Leaderboard), with participant information removed for privacy.
Google fixed Urdu in one generation
On identical prompts, Gemini 2.5 scored 1.19/5 on text accuracy (fluent readers alone: 1.13); Gemini 3.1 scored 4.94. Raters read the same Eid headline as "inaayirik" (A), "andainaaki" (B), and "Eid Mubarak, Kashmir Crafts" (C).
"Looks right, reads wrong" is the dangerous failure
Gemini 2.5 produced the most photorealistic scenes yet wrote decorative pseudo-Urdu or escaped to English, including a fully hallucinated English wedding invitation. A non-reading owner could ship the pretty fakes; a reading customer would mock them. This is exactly what automated metrics and non-native eyes miss.
Even the winner isn't deployable unsupervised
Gemini 3.1 spelled almost everything correctly but still over generated: uninvited English taglines, an unrequested Bismillah header, invented names, and even a template leak printed into the final artwork. A spelling-only metric would have scored it flawless.
Cultural fit needs local judges
Only native readers could tell whether the outputs felt authentically Kashmiri. They recognised local details such as the vertical tandoor used by kandurs and papier mâché motifs, while also catching failures that text accuracy alone missed. Location phrases such as “for a shop window” caused models to paint the scene around the poster rather than the poster itself, breaking the intended 1:1 artifact.
From evaluation to product concept
The evaluation was designed as more than a one time study. I translated it into a reusable product concept that a lab could run repeatedly across models, releases, and Indian scripts.
- Collect: the blind rating UI, whose typed-reading field doubles as evidence collection.
- Rank: a leaderboard with criterion filters, including a "fluent readers only" view.
- Diagnose: a quality-flag feed that turns rater comments into a fix-list for labs.
- Scale:from an 8 rater pilot to a standing benchmark for Indian scripts. Josh Talks AI already operates native language evaluations at much larger scales, making this approach reusable across model releases and languages.
Question 2: a data-driven quality system
Part 2 asked me to catch low-quality transcribers from data alone. Transcribers are paid per accurate hour of audio; their real cost is time, and listening is the slow part, so all bad work reduces to one behaviour: submitting without truly hearing the audio. That behaviour leaves three fingerprints in the 51,783-task dataset.
Three warning signs
- Impossible listening time:
time_taken < durationmeans the clip physically cannot have been heard. Fires on 5,985 tasks (11.6%); a fifth used under half the audio's length. - Blindly accepting the AI draft: unedited text alone is weak evidence, but unedited and faster than the audio strongly suggests no listening and no human value added (3,568 tasks).
- Effort-free submissions: 10 seconds of continuous Hindi holds 25 to 35 words, so a 5-second clip with under 5 characters is speech dismissed, not transcribed.
| Signal | Tasks | Share | What it means |
|---|---|---|---|
| Impossible listening time | 5,985 | 11.6% | Submitted faster than the audio duration itself. |
| Impossible speed + unedited | 3,568 | 6.9% | Accepted the AI draft without edits and submitted faster than the audio could be heard. |
| Effort-free submissions | 1,091 | 2.1% | Long audio clips (>5 seconds) submitted with fewer than 5 characters. |
Supporting documentation: Detailed analyses for all behavioral signals are available here.
The blocking system: behaviour flags, verified evidence blocks
Because these accounts get blocked, the logic has to be one I'm sure of. Speed and edit stats are circumstantial, so they only flag; conviction comes from gold-standard clips: audio with a known-correct transcription injected invisibly into a suspect's queue. No honest worker fails gold repeatedly; every cheater does.
- Tier 0: listen ratio < 1 AND unedited, or over-5s audio with under-5 characters, then withhold pay, requeue, 1 strike.
- Tier 1: 3 strikes or ≥20% of the last 20 tasks flagged, then inject 5 unannounced gold clips scored by word error rate.
- Tier 2: gold < 70% gives a warning and a hold; ≥70% clears the flags. This exit ramp is the fairness guarantee for fast-but-accurate workers.
- Tier 3: gold failed twice after a warning, then block and reverse unpaid flagged earnings.
Reflection & learnings
- "Best" depends on the criterion. The most photorealistic model was the least usable. An overall ranking would have hidden that tradeoff.
- Humans catch what metrics cannot. Template leaks,invented content, and uninvited English were invisible to spelling checks but obvious to native readers.
- Design the safeguards, not just the detection. When a quality system can affect someone’s income, the verification process is as important as the detection logic.
- Local expertise is a competitive advantage. Because my raters use Urdu in everyday commerce, I could evaluate real business use cases that broad multilingual benchmarks often miss.
This project changed how I think about AI quality: the hardest part isn’t generating outputs. It’s defining quality, measuring it fairly, and designing systems people can trust.