Evaluating AI Across Agriculture and Law Using Community and Expert Input

Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.

10,000

community-sourced queries

30,000+

evaluation tasks

7

Indian languages

Hybrid human and expert evaluation

The work focuses on how models perform in high-impact domains such as agriculture, law, and education, using 10,000 community-sourced queries across seven Indian languages. Rather than relying on abstract benchmarks, evaluations are designed to reflect real usage conditions and decision-making contexts.

The approach combines large-scale public evaluation with expert review. Over 30,000 evaluation tasks assess model outputs for correctness, fluency, and trustworthiness, alongside detailed input from domain experts, including agronomists and legal professionals. This enables a deeper understanding of how AI systems behave in practice, where accuracy alone is not sufficient.

Related