
Evaluating AI Across Agriculture and Law Using Community and Expert Input
Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.
Samiksha is a rigorous, community and expert-driven evaluation framework for benchmarking AI models in real Indic-language contexts, developed by Karya in collaboration with the Collective Intelligence Project and Microsoft.
23,000+
real-world queries
11
Indian languages
150,000+
human evaluations
1.6M+
automated evaluations
Most existing evaluations rely on synthetic prompts, translated benchmarks, or narrow accuracy metrics—approaches that break down when models are deployed across real users, languages, and social settings. Samiksha addresses this by grounding evaluation in lived realities, combining expert-designed queries with large-scale participation from diverse language communities. The first phase evaluated over 23,000 real-world queries across four critical domains: health, education, finance, and legal.
The evaluations found gaps in model responses that standard benchmarks do not typically surface. For example, when asked in Bengali how to pay an electricity bill online, a leading commercial model directed the user to bKash and Nagad, mobile wallets used in Bangladesh rather than in India. For a user in West Bengal reaching for UPI, PhonePe, or a Jan Dhan-linked account, this answer routed them toward systems they cannot use.

Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.



Until recently, Bhili, a tribal language spoken by millions of people across western and central India, had limited representation in the country's digital infrastructure. Karya, working with the Bhil community, has helped build a Bhili language system that now lives on two national-scale platforms.