
Evaluating AI Across Agriculture and Law Using Community and Expert Input
Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.
Project Vaani, a collaboration between Google and Indian Institute of Science, aims to map India's diverse linguistic landscape by collecting audio speech data from approximately 1 million people across 773 districts.
6000
hours of voice data
30
districts
24K+
unique speakers
15
minutes/speaker
With Karya's expertise, 6000 hours of voice data will be gathered from 30 districts, contributing to one of the largest datasets of Indian dialects, totaling over 150,000 hours of audio upon completion. Karya plays a pivotal role in this initiative by mobilising local communities, training field coordinators, and ensuring fair compensation for data collectors, thus empowering local voices and fostering inclusivity.

Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.



Until recently, Bhili, a tribal language spoken by millions of people across western and central India, had limited representation in the country's digital infrastructure. Karya, working with the Bhil community, has helped build a Bhili language system that now lives on two national-scale platforms.