Building Real-World AI Evaluation with Samiksha

Samiksha is a rigorous, community and expert-driven evaluation framework for benchmarking AI models in real Indic-language contexts, developed by Karya in collaboration with the Collective Intelligence Project and Microsoft.

23,000+

real-world queries

11

Indian languages

150,000+

human evaluations

1.6M+

automated evaluations

Most existing evaluations rely on synthetic prompts, translated benchmarks, or narrow accuracy metrics—approaches that break down when models are deployed across real users, languages, and social settings. Samiksha addresses this by grounding evaluation in lived realities, combining expert-designed queries with large-scale participation from diverse language communities. The first phase evaluated over 23,000 real-world queries across four critical domains: health, education, finance, and legal.

The evaluations found gaps in model responses that standard benchmarks do not typically surface. For example, when asked in Bengali how to pay an electricity bill online, a leading commercial model directed the user to bKash and Nagad, mobile wallets used in Bangladesh rather than in India. For a user in West Bengal reaching for UPI, PhonePe, or a Jan Dhan-linked account, this answer routed them toward systems they cannot use.


Learn more here

Related