
Evaluating AI Across Agriculture and Law Using Community and Expert Input
Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.
In Q1 2026, Karya launched Project Apollo, its largest investment in off-the-shelf AI data infrastructure to date. The project created 100,000 hours of transcribed conversational speech across India's 22 Scheduled Languages—the languages recognised in the Eighth Schedule of the Constitution of India and used across government, education, media, and public life.
100,000
hours of transcribed conversational speech
22
Scheduled Languages in India
$2M
disbursed to date through Project Apollo
60,000
workers
Beginning with 15 languages already in production, Apollo is building one of India's largest publicly available multilingual conversational speech corpora for AI research and development. Unlike scripted recordings, conversational speech captures how people naturally communicate across accents, dialects, code-switching, and everyday interactions. These datasets can support the development of speech recognition, translation, voice interfaces, accessibility technologies, and multilingual foundation models.
Project Apollo engages with 60,000 workers across India, with a target of USD 2.5 million in direct worker wages.

Karya is evaluating conversational AI systems in real-world Indian contexts, in partnership with Anthropic.



Until recently, Bhili, a tribal language spoken by millions of people across western and central India, had limited representation in the country's digital infrastructure. Karya, working with the Bhil community, has helped build a Bhili language system that now lives on two national-scale platforms.