
Evaluating AI Across Agriculture and Law Using Community and Expert Input
Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.
90K
human evaluations
30
models
10
Indian languages
3
weeks
Evaluation of multilingual LLMs is challenging due to insufficient linguistic diversity, benchmark contamination and the lack of local, cultural nuances in translated benchmarks. Karya’s data experts can evaluate models based on an array of benchmarks, including testing for linguistic acceptability, hallucinations, reasoning, and creativity.

Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.



Until recently, Bhili, a tribal language spoken by millions of people across western and central India, had limited representation in the country's digital infrastructure. Karya, working with the Bhil community, has helped build a Bhili language system that now lives on two national-scale platforms.